munotes®

Theory of Computation Notes | B.Sc. (Computer Science) Semester 3 | Mumbai University | munotes

Get access to whole semester resourcesSemester Pass

Official Notes munotes.in

Theory of Computation

B.SC. (COMPUTER SCIENCE) · SEMESTER 3

Strictly as per the University of Mumbai NEP syllabus in force for B.Sc. (Computer Science)

For B.Sc. (Computer Science) students of the University of Mumbai and all its affiliated colleges

Open the book ↓

munotes.in Second Year

Theory of Computation

Copyright © 2026 munotes.in. All rights reserved.

Written and first published by munotes.in, 2026.

This book is free for individual students to read at munotes.in. No part of it may be reproduced, distributed, stored, translated or used for institutional or classroom purposes in any form without a prior written licence from munotes.in.

Licensing and permissions: contact@munotes.in

The text of statutes and of judgments reproduced in this book is in the public domain under section 52(1)(q) of the Copyright Act 1957. The commentary, arrangement, examples and questions are the original work of munotes.in.

munotes.in is an independent study resource for MU students. It is not affiliated with, endorsed by, or officially connected to the University of Mumbai. Course names and university references describe the students and syllabus the material relates to.

munotes.in

Contents

Module I Introduction to Theory of Computation, Automata Theory, Formal Languages, Regular Languages and Context Free Languages

  1. What This Subject Is About, and the Four Questions It Answers 1
  2. Sets: The Vocabulary Everything Else Is Written In 6
  3. Alphabets, Strings and Languages 11
  4. Relations, and the Equivalence Relation That Runs Through This Subject 16
  5. Functions: What the Transition Function Actually Is 21
  6. Proof Techniques: Five Ways to Be Sure 25
  7. Counting the Uncountable: Diagonalisation, and Why Some Languages Have No Machine 30
  8. What an Automaton Is 35
  9. The Deterministic Finite Automaton, Formally 40
  10. Transitions and Their Properties: the Function, the Table and the Diagram 44
  11. The Extended Transition Function: What a Machine Does to a Whole String 49
  12. Acceptability: When a Machine Accepts a String, and What Language It Accepts 53
  13. Designing a DFA: the Method, and Eight Machines Built With It 58
  14. Nondeterminism, and the Nondeterministic Finite Automaton 64
  15. Empty Moves, and the Epsilon Closure 69
  16. DFA and NDFA Equivalence: the Subset Construction 73
  17. Removing the Empty Moves 79
  18. Mealy and Moore Machines: a Machine That Writes 83
  19. Converting a Moore Machine to a Mealy Machine, and Back 87
  20. Minimizing Automata: the Partition Method 92
  21. The Myhill Nerode Theorem, and a Second Way to Prove a Language Not Regular 98
  22. What a Grammar Is 104
  23. Derivations: How a Grammar Makes a String 109
  24. The Language Generated by a Grammar, and Proving It Is the One You Claim 114
  25. Writing a Grammar for a Language You Are Given 119
  26. The Chomsky Classification of Grammars and Languages 124
  27. Recursive and Recursively Enumerable Sets 130
  28. Operations on Languages 135
  29. Languages and Automata: Which Machine Goes With Which Grammar 141
  30. Regular Grammar: the Right Linear and Left Linear Forms 146
  31. Regular Expressions 151
  32. The Identities of Regular Expressions 155
  33. Writing a Regular Expression for a Language You Are Given 160
  34. From a Regular Expression to a Finite Automaton 166
  35. From a Finite Automaton to a Regular Expression: Arden's Theorem and State Elimination 171
  36. The Pumping Lemma for Regular Languages 177
  37. Applications of the Pumping Lemma 183
  38. Closure Properties of the Regular Languages 189
  39. The Decision Problems of the Regular Languages 195
  40. Regular Sets and Regular Grammar: Kleene's Theorem Assembled 200
  41. Context Free Grammars and Context Free Languages 204
  42. The Derivation Tree 208
  43. Leftmost and Rightmost Derivations 213
  44. Ambiguity of a Grammar 217
  45. Simplifying a Grammar, One: Useless Symbols 222
  46. Simplifying a Grammar, Two: Null Productions 227
  47. Simplifying a Grammar, Three: Unit Productions, and the Reduced Grammar 232
  48. Chomsky Normal Form 237
  49. Greibach Normal Form 242
  50. The Pumping Lemma for Context Free Languages 247
  51. Closure Properties and Decision Problems of the Context Free Languages 254

Module II Pushdown Automata, Linear Bound Automata, Turing Machines, and Computability and Complexity

  1. The Pushdown Automaton 261
  2. Instantaneous Descriptions and Moves 266
  3. Acceptance by a PDA: by Final State and by Empty Stack 271
  4. Designing a PDA: the Balanced Languages 276
  5. Designing a PDA: Palindromes, and Counting Two Things at Once 281
  6. The Deterministic Pushdown Automaton 287
  7. From a Context Free Grammar to a Pushdown Automaton 292
  8. From a Pushdown Automaton to a Context Free Grammar 297
  9. The Linear Bounded Automaton Model 302
  10. Linear Bounded Automata and the Context Sensitive Languages 307
  11. The Turing Machine 312
  12. Representations of a Turing Machine 317
  13. Acceptability by a Turing Machine: Accept, Reject and Loop 322
  14. Designing and Describing a Turing Machine 326
  15. Turing Machine Construction: Machines That Recognise 331
  16. Turing Machine Construction: Machines That Compute 336
  17. Variants of the Turing Machine: More Tapes, More Tracks, More Heads 341
  18. Variants of the Turing Machine: Nondeterministic, Offline, and the Enumerator 346
  19. Recursive and Recursively Enumerable Languages 351
  20. Decidable and Undecidable, and What the Complement Tells You 355
  21. The Church Turing Thesis 359
  22. The Universal Turing Machine 364
  23. The Halting Problem 369
  24. Reduction: Proving a Second Problem Unsolvable 374
  25. Rice's Theorem 379
  26. The Post Correspondence Problem, and the Undecidable Problems of Grammars 384
  27. Time Complexity 389
  28. Space Complexity 393
  29. Big O Notation, and the Family Around It 397
  30. Class P 402
  31. Class NP, and the Certificate 406
  32. Polynomial Reductions 411
  33. NP Complete and NP Hard, and Cook's Theorem 417
  34. Proving a Problem NP Complete: Three Reductions Worked 423
  35. Complexity Hierarchies, and the P Against NP Question 429
  36. The Machines and the Grammars, Side by Side 434
munotes.in

Module I

Introduction to Theory of Computation, Automata Theory, Formal Languages, Regular Languages and Context Free Languages

munotes.in

Chapter One

What This Subject Is About, and the Four Questions It Answers

Syllabus topic Module 1, "Introduction to Theory of Computation: Basics of Computation, Importance of Theory of Computation in Computer Science"

In one line

The theory of computation is the study of what problems a machine can solve, and how much time and space solving them takes.

In the wording a student can write in an examination: the theory of computation is the branch of computer science that builds precise mathematical models of a computing machine, and uses them to classify problems according to whether they can be solved at all, and if so with what resources.

Why a computer science degree teaches this at all

Every other paper on this degree teaches you to make a computer do something. This one asks a different question, and it is the only paper that asks it: what can a computer do at all?

That sounds like a question with an obvious answer, and it is not. There are problems no computer will ever solve, not because the machine is too slow or the programmer is not clever enough, but because no method exists. There are other problems that can be solved in principle and cannot be solved in practice, because every known method would run past the age of the universe on an input the size of a phone book. Both of those facts were proved, by hand, before the first working computer was built.

The four questions

Everything in this paper is an answer to one of four questions. It is worth having them on one page before starting, because each of the five blocks of the syllabus is an attack on one of them.

One. What can be computed at all? This is the question Module 2 ends on. Some problems have no algorithm, and we can prove which ones.

Two. What can be computed quickly? A problem with an algorithm that takes a thousand years is not solved in any sense a working programmer would accept. The last block of Module 2 is about drawing that line precisely.

Three. How much machine does a job need? Not every task needs a full computer. Checking that a password has a digit in it needs almost nothing. Checking that brackets balance needs a little memory. Deciding whether a program halts needs more than any machine has. Module 1 and Module 2 build a ladder of four machines, each stronger than the last, and place jobs on the rungs.

Four. How can we be sure? Every claim in this paper is proved. That is not academic manners; it is the only way to know that a machine is right for every input rather than for the inputs you happened to try.

The four questions, with a problem you already have

A list of questions is easy to forget. Each of these is a problem you have met on this degree already, and each one is one of the four questions in disguise.

munotes.in1

What This Subject Is About, and the Four Questions It Answers

A search box that accepts patterns. You want to find every line in a file that begins with a capital letter and ends in a semicolon. The pattern language that does this is called a regular expression, and the whole of the fourth block of Module 1 is about it. The interesting part is not that it works. It is that there is a precise boundary to what patterns of this kind can express, and "lines with the same number of opening and closing brackets" is on the far side of it. You will prove that in chapter 37.

A compiler. Before your Java compiler can object to a missing bracket, it has to know what a correctly bracketed program looks like. The description it uses is a grammar, which is the third block of Module 1, and the machine that reads it is a pushdown automaton, which opens Module 2. When your compiler reports a syntax error, the thing that found it is one of the machines in this book.

A lock screen. A four digit PIN entry is a machine with a small number of states: how many digits have been typed, and whether they were right. Nothing else needs to be remembered. That is a finite automaton, and it is the second block of Module 1.

A marking scheme. Suppose a lecturer wants a program that reads a student's submitted program and reports whether it will ever loop for ever. Every student wants this to exist. It cannot. That is the halting problem, it is proved in chapter 74, and the proof takes half a page.

What the word computation is standing for

The word is doing a lot of work, so it is worth being careful about it early.

A problem here always means a question about an input, and we shall almost always arrange for the answer to be yes or no. "Is this number prime?" is a problem. "Is this bracket sequence balanced?" is a problem. "Sort this list" looks different, but it can be turned into a yes or no question, and reducing everything to yes or no is what makes a mathematical treatment possible.

An instance of a problem is one particular input. The number 91 is an instance of the primality problem, and the answer for that instance is no, because 91 is 7 times 13.

An algorithm is a finite list of unambiguous instructions that, followed exactly, produces the answer for every instance, and stops. Each of those words matters. Finite, so it can be written down. Unambiguous, so no judgement is needed at any step. Every instance, so it is not a method that works for the cases we tried. And stops, because a procedure that runs for ever has not answered anything.

munotes.in2

What This Subject Is About, and the Four Questions It Answers

Computation is what happens when a machine carries out an algorithm on an instance. This subject makes "machine" precise, four times over, each time a little stronger.

Why the models are so simple

The first thing a student notices about this paper is that the machines are absurdly primitive. A Turing machine has one tape, one head, and no arithmetic. That is not a historical accident or a simplification for teaching. It is the point, and there are two reasons for it.

The first is that a simple model can be reasoned about. You cannot prove anything about a machine whose description takes ten thousand pages. You can prove a great deal about a machine whose description takes five lines, and a proof about the simple machine is worth having only if the simple machine is as strong as the complicated one.

The second is that it turns out to be as strong. This is the single most surprising result in the subject, and chapter 72 is about it: every attempt anybody has made to define a more powerful machine, by adding tapes, heads, dimensions, randomness or guessing, has produced a machine that solves exactly the same set of problems. That is the Church Turing thesis, and it is what licenses the whole enterprise. When we prove that no Turing machine can solve a problem, we are entitled to say that no computer can.

Where the subject came from

Three dates are worth remembering, because they explain why the subject looks the way it does.

In 1936 Alan Turing published "On Computable Numbers, with an Application to the Entscheidungsproblem" in the Proceedings of the London Mathematical Society, pages 230 to 265. The paper's own first page records that it was received on 28 May 1936 and read on 12 November 1936. Turing was not designing a computer. He was answering a question David Hilbert had asked about mathematics, and the machine was a tool invented for the proof. The tool turned out to be more important than the answer.

In 1951 Stephen Kleene, working at RAND, wrote up finite automata and the expressions that describe them, in a research memorandum numbered RM-704 and titled "Representation of Events in Nerve Nets and Finite Automata". Its own summary dates the investigations it reports to August 1951. And in 1956 Noam Chomsky, studying human language rather than machines, wrote "Three Models for the Description of Language", which his own archive dates to September 1956. The classification of grammars in chapter 26 is from that paper.

munotes.in3

What This Subject Is About, and the Four Questions It Answers

So the four machines in this book were invented by a logician, an electrical engineer and a linguist, none of whom was trying to build a computer. That is worth knowing when the subject feels disconnected from programming. It was not derived from programming; programming arrived later and turned out to fit.

How this book is arranged, and why the order is MU's

The chapters follow the University's printed syllabus in her own order, under her own headings. That order is also a good teaching order, which is not always true of a syllabus, and it is worth seeing why before starting.

Module 1 builds the weakest machine first, the finite automaton, and pushes it until it breaks. Then it introduces grammars, which describe languages from the other direction, and shows that the weakest grammar describes exactly what the weakest machine accepts. Then it takes the next grammar up, the context free one, and Module 2 builds the machine that matches it. Then it does the same thing twice more. By the end there are four machines, four kinds of grammar, and a proof that they pair off.

What it does NOT mean

This is not a paper about programming languages. The word "language" in this subject means something much simpler: a set of strings. The language of a machine is just the collection of inputs it says yes to. It has no syntax, no semantics and no compiler.

A machine here is not a computer. It is a mathematical object, defined by a small number of sets and a function. Nothing is ever built. Asking how fast a Turing machine runs in megahertz is asking the wrong sort of question about the wrong sort of object.

"Cannot be computed" is not a statement about today's computers. It is not "no current machine can do this", nor "we have not found a method yet". It means no method exists, and never will, and that this has been proved.

Quick revision

  • The theory of computation studies which problems a machine can solve, and at what cost in time and

space.

  • Four questions: what is computable at all; what is computable quickly; how much machine a job

needs; and how we can be sure.

  • A problem is a question about an input; an instance is one input; an algorithm is a finite,

unambiguous method that answers every instance and stops.

  • A language, in this subject, is a set of strings, and the language of a machine is the set of

inputs it accepts.

  • The models are deliberately primitive, because simple models can be proved things about, and

because they turn out to be as strong as complicated ones.

munotes.in4

What This Subject Is About, and the Four Questions It Answers

  • Turing 1936, pages 230 to 265; Kleene's RAND memorandum RM-704 of 1951; Chomsky's

three models paper of 1956.

  • MU examines this paper in one hour for 30 marks, in three questions of two answers each, so every

topic in this book has to be answerable in about ten minutes.

Test yourself

1. Define an algorithm, and say what each part of the definition rules out. A finite list of unambiguous instructions that, followed exactly, produces the correct answer for every instance of the problem and then stops. Finite rules out a method that cannot be written down; unambiguous rules out a step requiring judgement; every instance rules out a method that works only on the cases tried; and stopping rules out a procedure that runs for ever without answering.

2. Give one reason the machines in this subject are defined so simply. Because a proof can be carried out about a five line definition and cannot be carried out about a real processor, and because the simple machines turn out to solve exactly the same problems as any richer model anyone has proposed.

3. What does the word language mean in this paper? A set of strings over a fixed alphabet. Nothing more: no syntax rules beyond membership, and no meaning attached to the strings.

4. State the four questions the subject answers. What can be computed at all; what can be computed quickly; how powerful a machine a given job needs; and how we can be certain of the answers.

5. A student says a problem is uncomputable because no fast enough computer exists yet. Uncomputable means no algorithm exists at all, which is a proved mathematical statement about the problem and not about any machine. A problem that needs a faster computer is computable but expensive, which is a question for the complexity block of Module 2, not the computability block.

Contents This chapter on its own page

munotes.in5

Chapter Two

Sets: The Vocabulary Everything Else Is Written In

Syllabus topic Module 1, "Mathematical Foundations (Sets, Relations, Functions, Proof Techniques)"

In one line

A set is a collection of distinct things, and the only question you may ask about a set is whether a particular thing is in it.

In the wording a student can write in an examination: a set is a well defined, unordered collection of distinct objects, called its elements or members. Well defined means that for any object it is determinate whether the object belongs to the set or not.

Why this subject needs sets at all

Every single object in this paper is a set, or is built out of sets. A machine's states are a set. Its alphabet is a set. The strings it accepts are a set. The collection of languages a whole class of machines accepts is a set of sets. The definitions in chapters 9, 22, 52 and 62 are all of the form "a five part object, of which three parts are sets and one is a function between them".

So sets are not preliminary material to be got through. They are the notation, and a student who is vague about the difference between an element and a one element set will be defeated by the subset construction in chapter 16, where the states of the new machine are sets of states of the old one.

Writing a set down

There are three ways, and all three appear in this book.

By listing. Write the elements between braces, separated by commas.

A = {0, 1}

B = {q0, q1, q2}

Order does not matter and repetition does not count, so {0, 1} and {1, 0} and {0, 1, 1} are all the same set. That last point matters: a set has no notion of how many times something is in it.

By a property. Write a description of what qualifies.

C = { n : n is a whole number and n is even and n is less than 10 }

The colon is read "such that", so C is "the set of n such that n is a whole number and n is even and n is less than 10", which is {0, 2, 4, 6, 8}. Some books write a vertical bar instead of the colon and mean exactly the same thing.

By a rule that generates. This is the form this subject uses most, and chapter 22 makes it precise. A grammar is a set of rules for writing strings, and the set it defines is everything the rules can produce.

The words, one at a time

Element or member. A thing in the set. We write "0 is in A" and "2 is not in A".

The empty set. The set with no elements at all, written { }. It is a perfectly good set, and it will matter: a finite automaton with no final states accepts the empty language, and chapter 39 is partly about that machine.

munotes.in6

Sets: The Vocabulary Everything Else Is Written In

Careful. The empty set and the set containing the empty set are different. { } has no elements. { { } } has one element, which happens to itself be a set with no elements. The same distinction, with strings instead of sets, catches students in chapter 3.

Subset. A is a subset of B when every element of A is also an element of B. The empty set is a subset of every set, and every set is a subset of itself, both of which follow straight from the definition rather than being special cases. A is a proper subset of B when A is a subset of B and A is not equal to B.

Equality. Two sets are equal when each is a subset of the other. That sounds like a roundabout way to say the obvious, and it is in fact the most useful sentence in this chapter, because it is the shape of almost every proof in the book: to prove two languages are the same, prove each is contained in the other. Chapter 24 does it for a grammar, chapter 16 for two machines.

Cardinality. The number of elements, written with vertical bars. So the size of {0, 1} is 2. A set is finite when its cardinality is a whole number, and infinite otherwise. Chapter 7 is about the fact that there are two different sizes of infinite, which is the most consequential fact in the whole paper.

Disjoint. Two sets are disjoint when they have no element in common.

The operations

Six, and each one has a use in this book.

OperationWrittenMeans
UnionA union Beverything in A, in B, or in both
IntersectionA intersect Beverything in both A and B
DifferenceA minus Beverything in A that is not in B
ComplementA'everything in the universe that is not in A
Cartesian productA times Bevery ordered pair with a first part from A and a second from B
Power set2 to the Athe set of all subsets of A

The complement needs a universe. "Everything not in A" is meaningless until we say what everything is. In this book the universe is almost always the set of all strings over a fixed alphabet, which chapter 3 calls Sigma star, and the complement of a language is every string the language does not contain. Forgetting to fix the universe is the commonest error in a closure proof.

The Cartesian product is how two machines are run at once. In chapter 38 we prove the regular languages are closed under intersection by building a machine whose states are pairs: one component tracks where the first machine would be, the other where the second would be. The set of those pairs is exactly the Cartesian product of the two state sets. The same construction is what the checker behind this book uses to decide whether two machines are equal.

munotes.in7

Sets: The Vocabulary Everything Else Is Written In

The power set is how a nondeterministic machine is made deterministic. Chapter 16 builds, out of a machine that may be in several states at once, a machine whose single state records which set of old states we could be in. Its states are the subsets of the old state set, which is the power set, and that is where the construction gets its name.

Worked example

Let A = {0, 1} and B = {1, 2}, with the universe U = {0, 1, 2, 3}.

ExpressionValue
A union B{0, 1, 2}
A intersect B{1}
A minus B{0}
B minus A{2}
A'{2, 3}
A times B{(0,1), (0,2), (1,1), (1,2)}

The size of A is 2 and the size of B is 2, so the size of A times B is 4, and in general the size of a Cartesian product is the product of the sizes. That is why the product machine of chapter 38 can have as many states as the two machines multiplied together.

Now the power set of A. A has two elements, so it has four subsets:

2 to the A = { { }, {0}, {1}, {0, 1} }

A set with n elements has exactly 2 to the n subsets, and this is worth proving rather than accepting, because the argument is used again in chapter 16. To build a subset, go through the n elements one at a time and make an independent yes or no decision about each. That is n decisions with 2 choices each, so 2 times 2 times ... n times, which is 2 to the n. The empty set comes from saying no n times and the whole set from saying yes n times, which is why both are subsets.

That number is also the reason the subset construction is expensive: a nondeterministic machine with n states becomes a deterministic machine with up to 2 to the n states, and chapter 14 shows a family of machines where it really is that bad.

The laws

These get used, mostly without being named, in the closure proofs of chapters 38 and 51. Each one holds for sets exactly as it holds for the logical connectives.

munotes.in8

Sets: The Vocabulary Everything Else Is Written In

LawStatement
CommutativeA union B = B union A
Associative(A union B) union C = A union (B union C)
DistributiveA intersect (B union C) = (A intersect B) union (A intersect C)
IdempotentA union A = A
De Morgan(A union B)' = A' intersect B'
Double complement(A')' = A

De Morgan's law is the one that earns marks, because it converts a union into an intersection and back. Chapter 38 uses it exactly once, and it is the whole of the proof: the regular languages are closed under complement and under union, and De Morgan turns those two facts into closure under intersection with no new construction at all.

What it does NOT mean

A set is not a list. It has no order and no repeats. If order matters you want a sequence or a tuple, written with round brackets: (0, 1) is not the same as (1, 0), and a string, which chapter 3 defines, is a sequence and not a set.

A subset is not an element. {0} is a subset of {0, 1}. It is not an element of {0, 1}. It IS an element of the power set of {0, 1}. Getting this wrong makes chapter 16 unreadable, because there the elements of the new machine's state set are subsets of the old machine's state set.

The empty set is not nothing. It is a set, it has a size, and a machine can accept the language it denotes.

Quick revision

  • A set is an unordered collection of distinct elements, and the only question is membership.
  • Written by listing, by a property with a colon read "such that", or by a generating rule.
  • Subset: every element of one is an element of the other. Equal: each is a subset of the other, and

this is the shape of most proofs in the book.

  • Cardinality is the number of elements; the empty set { } has none and is a subset of everything.
  • Six operations: union, intersection, difference, complement, Cartesian product, power set.
  • A set of n elements has 2 to the n subsets, which is why the subset construction of chapter 16 can

be exponential.

  • The complement is meaningless without a stated universe; here the universe is Sigma star.
  • De Morgan's law is what turns closure under complement and union into closure under intersection.

Test yourself

1. Write, by a property, the set of all strings of 0s and 1s of even length. { w : w is a string over {0, 1} and the length of w is even }. Note this set is infinite, and it contains the empty string, whose length is 0, which is even.

munotes.in9

Sets: The Vocabulary Everything Else Is Written In

2. How many subsets does a set of 5 elements have, and why?

  1. Each of the 5 elements is independently either in or out of a subset, giving 2 to the 5 choices.

3. Are { } and { { } } the same set? Give their sizes. No. { } has size 0. { { } } has size 1, its single element being the empty set.

4. A = {a, b}, B = {b, c}, universe {a, b, c, d}. Give A union B, A intersect B, A minus B and B'. {a, b, c}; {b}; {a}; {a, d}.

5. Using De Morgan's law, express A intersect B using only union and complement. A intersect B = (A' union B')'. This is the step chapter 38 uses to prove the regular languages are closed under intersection without building a new machine.

6. Why does this subject care about the Cartesian product? Because running two machines side by side on the same input is done by building one machine whose states are pairs of states, and the set of those pairs is the Cartesian product of the two state sets.

Contents This chapter on its own page

munotes.in10

Chapter Three

Alphabets, Strings and Languages

Syllabus topic Module 1, "Mathematical Foundations (Sets, Relations, Functions, Proof Techniques)"

In one line

An alphabet is a finite set of symbols, a string is a finite sequence of symbols from an alphabet, and a language is a set of strings.

In the wording a student can write in an examination: an alphabet, written Sigma, is a finite non empty set of symbols. A string or word over Sigma is a finite sequence of symbols drawn from Sigma. A language over Sigma is any subset of the set of all strings over Sigma.

Why these three and not others

Every problem in this subject has to be presented to a machine as an input, and the machine has to give back one bit of information: yes or no. So the whole of computation, for the purposes of this paper, reduces to this: given a set of strings that we have decided are the good ones, can a machine of a given kind tell a good string from a bad one?

That reduction is what makes the subject possible. It looks as though it throws away too much, and it does not: a number can be written as a string, a graph can be written as a string, and a whole program can be written as a string, which is exactly what chapter 73 does when it feeds one machine to another.

The alphabet

An alphabet is a finite, non empty set of symbols. We write it Sigma. The three used most in this book are:

Sigma = {0, 1}

Sigma = {a, b}

Sigma = {a, b, c}

Two words in that definition are load bearing. Finite, because a machine must be able to look at a symbol and decide what to do, and with infinitely many symbols no finite table of instructions could cover them. Non empty, because an alphabet with no symbols permits exactly one string and makes every question trivial.

A symbol is not required to be a single character. The instruction set of a processor is an alphabet whose symbols are whole instructions. It only has to be an indivisible unit as far as the machine is concerned.

The string

A string over Sigma is a finite sequence of symbols from Sigma. We write strings with no separators and no brackets, so over {a, b} the string is written abba rather than (a, b, b, a).

The length of a string w, written with vertical bars around it, is how many symbols it has, with repeats counted. The length of abba is 4.

The empty string is the string with no symbols at all. It is written with the Greek letter epsilon, so:

ε is the string of length 0

Some textbooks write it as a capital lambda instead. This book uses epsilon everywhere, including inside transition tables and grammars.

munotes.in11

Alphabets, Strings and Languages

The empty string is the hardest of the three definitions for a beginner, because it feels like nothing. It is not nothing. It is a perfectly good string that a machine can be handed, and whether a machine accepts it is a separate question from everything else, which is exactly why it is the string students forget to test. In this book every set of test strings printed under a machine includes epsilon when the machine accepts it, and the checker runs it.

Operations on strings

Concatenation. Writing one string after another. If x is ab and y is ba then the concatenation of x and y is abba. It is written by juxtaposition, xy, and it is associative but not commutative: ab followed by ba is abba, and ba followed by ab is baab.

The empty string is the identity for concatenation, which is the formal reason it has to exist:

εw = w

wε = w

Powers. w to the k means w concatenated with itself k times, and w to the 0 is defined to be epsilon. So a to the 3 is aaa, and (ab) to the 2 is abab.

The definition of w to the 0 as epsilon is a convention, but not an arbitrary one: it is forced if the law w to the m followed by w to the n equals w to the (m plus n) is to hold for every m and n including zero.

Reversal. The string written backwards, written w to the R. The reversal of abb is bba, and the reversal of epsilon is epsilon. A string equal to its own reversal is a palindrome, and the language of palindromes is the witness language in chapters 56 and 57.

Prefix, suffix, substring. A prefix is any number of symbols from the front, a suffix any number from the back, and a substring any run of consecutive symbols from anywhere. For abc the prefixes are epsilon, a, ab, abc; the suffixes are epsilon, c, bc, abc; and the substrings add b and bc and so on. Note that epsilon and the whole string are always a prefix, a suffix and a substring of any string.

A proper prefix is a prefix other than the string itself.

Careful with substring and subsequence. A substring must be consecutive. ac is not a substring of abc, because b is in the way. It is a subsequence, which allows gaps. This book uses substring in the strict sense throughout.

Sigma star and Sigma plus

These two are used on almost every later page.

Sigma star = the set of ALL strings over Sigma, including ε

Sigma plus = the set of all NON EMPTY strings over Sigma

munotes.in12

Alphabets, Strings and Languages

So Sigma plus is Sigma star with epsilon removed, and Sigma star is Sigma plus with epsilon put back.

Sigma star is infinite for any alphabet, because for each length there is at least one string and there is no longest length. But it is infinite in the tamest possible way: the strings can be listed in one endless sequence with none left out, first by length and then alphabetically within a length. Over {a, b} that listing begins

ε, a, b, aa, ab, ba, bb, aaa, aab, aba, abb, baa, bab, bba, bbb, aaaa, ...

This ordering is called canonical order, and it matters twice. Chapter 7 uses it to show Sigma star is countable, which is the first half of the argument that some languages have no machine. And it is the order the checker behind this book uses when it compares two machines over every string up to a given length: it is exhaustive, so nothing can hide in a gap.

The number of strings of length exactly k over an alphabet of size n is n to the k. Over {a, b} there are 2 strings of length 1, 4 of length 2, 8 of length 3. So the number of strings of length at most k is 1 plus n plus n squared and so on, and for {a, b} with k equal to 3 that is 1 plus 2 plus 4 plus 8, which is 15.

The language

A language over Sigma is any subset of Sigma star. That is the whole definition, and it is worth pausing on how permissive it is. A language need not be describable, need not be interesting, and need not have a machine. It is just a set of strings.

Some examples over {a, b}:

LanguageWrittenFinite?
nothing at all{ }finite, size 0
just the empty string{ε}finite, size 1
everythingSigma starinfinite
strings of even length{ w : the length of w is even }infinite
strings with equal a and b counts{ w : w has as many a as b }infinite
palindromes{ w : w equals w reversed }infinite

The two that look the same and are not. { } is the language with no strings in it. {ε} is the language with exactly one string in it, that string being the empty one. A machine that accepts { } accepts nothing and says no to every input including the empty input. A machine that accepts {ε} says yes to the empty input and no to everything else. They are different machines, they differ on exactly one input, and mixing them up is the single commonest slip in the first two weeks of this paper. It is the same distinction as the empty set against the set containing the empty set, from chapter 2.

munotes.in13

Alphabets, Strings and Languages

The operations on languages, briefly

Chapter 28 treats these properly. They are named here because the next few chapters use them.

Since a language is a set, all the set operations apply: union, intersection, difference, complement. Two more come from strings rather than sets. The concatenation of languages L and M is every string made by taking one string from L and one from M in that order. The Kleene closure of L, written L star, is every string made by concatenating any number of strings from L, including none at all, so L star always contains epsilon.

Worked example

Let Sigma = {0, 1}, x = 01 and y = 110.

ExpressionValue
the length of x2
xy01110
yx11001
x to the 3010101
y reversed011
x to the 0ε
proper prefixes of xε, 0

Now a language. Let L be the set of strings over {0, 1} that begin with 0 and have length 3. Listing it is quick, because the first symbol is fixed and the other two are free, so there are 2 times 2 which is 4 strings:

L = {000, 001, 010, 011}

And let M be the set of strings over {0, 1} of length at most 1, which is {ε, 0, 1}. Then the concatenation of L and M has 4 times 3 which is 12 strings if they are all different, and they are, because the pieces have different lengths or different last symbols. The three shortest are 000, 001 and 010 themselves, obtained by taking epsilon from M.

What it does NOT mean

A language is not a programming language. It has no grammar rules of its own, no meaning and no compiler. The word is borrowed from linguistics because Chomsky's work arrived through linguistics.

A string is not a set. Order matters and repeats count, so ab is not ba and aa is not a.

The empty string is not the empty language. One is a string, the other is a set of strings. See the paragraph above on { } against {ε}.

Sigma star is not the same as "all languages". Sigma star is a set of strings. The set of all languages over Sigma is the power set of Sigma star, which is a much bigger object, and chapter 7 shows exactly how much bigger.

Quick revision

  • Alphabet Sigma: a finite non empty set of symbols.
  • String: a finite sequence of symbols from Sigma. Length counts repeats. The empty string is
munotes.in14

Alphabets, Strings and Languages

epsilon, of length 0.

  • Concatenation is associative, not commutative, with epsilon as its identity. w to the 0 is epsilon.
  • Reversal, prefix, suffix, substring (consecutive) and subsequence (gaps allowed).
  • Sigma star: all strings. Sigma plus: all non empty strings. Both infinite, and Sigma star is

countable in canonical order.

  • Strings of length exactly k over an alphabet of size n: n to the k.
  • Language: any subset of Sigma star.
  • { } and {ε} are different languages, and the machines that accept them differ on one input.

Test yourself

1. Over {a, b}, how many strings have length exactly 4, and how many have length at most 4? 2 to the 4, which is 16, of length exactly 4. At most 4 gives 1 plus 2 plus 4 plus 8 plus 16, which is 31.

2. Give all the prefixes and all the suffixes of abb. Prefixes: ε, a, ab, abb. Suffixes: ε, b, bb, abb.

3. Is ac a substring of abc? Is it a subsequence? Not a substring, because a substring must be consecutive and b lies between them. It is a subsequence.

4. Write the language of strings over {0, 1} of length 2 that contain at least one 1. {01, 10, 11}. Note 00 is excluded and the language is finite with three members.

5. What is the difference between { } and {ε}, and how would a machine show it? { } contains no strings, so its machine rejects every input including the empty one; it has no final state reachable. {ε} contains exactly the empty string, so its machine accepts the empty input and rejects everything else; its start state is final and no other reachable state is.

6. Simplify x to the 0 followed by x to the 2, where x is ab. x to the 0 is epsilon, and epsilon is the identity for concatenation, so the result is x to the 2, which is abab.

Contents This chapter on its own page

munotes.in15

Chapter Four

Relations, and the Equivalence Relation That Runs Through This Subject

Syllabus topic Module 1, "Mathematical Foundations (Sets, Relations, Functions, Proof Techniques)"

In one line

A relation says which things are connected to which; an equivalence relation is a relation that behaves like equality, and it always cuts a set into blocks.

In the wording a student can write in an examination: a binary relation R on a set A is any subset of A times A, and we write a R b when the pair (a, b) is in R. A relation is an equivalence relation when it is reflexive, symmetric and transitive, and every equivalence relation on A partitions A into disjoint equivalence classes.

Why this subject needs relations

Three reasons, all of which arrive within twenty chapters.

The first is that the transition of a machine is a relation before it is a function. A nondeterministic machine, in chapter 14, can go from one state to several on the same symbol, so what connects states to states is a relation and not a function. Making it a function again is precisely what the subset construction does.

The second is that the whole idea of minimising a machine is an equivalence relation. Two states are the same for practical purposes when no string tells them apart, and "no string tells them apart" is reflexive, symmetric and transitive. So it partitions the states into blocks, and the minimal machine has one state per block. Chapter 20 is that sentence carried out.

The third is that the derivation relation of a grammar, in chapter 23, is a relation, and the notation used for it is the reflexive transitive closure defined at the end of this chapter.

What a relation is

A binary relation R on a set A is a subset of A times A, the Cartesian product of chapter 2. That is the whole definition: a relation is just a set of ordered pairs. If the pair (a, b) is in R we write a R b and say a is related to b.

Take A = {1, 2, 3} and let R be "is less than". Then

R = {(1,2), (1,3), (2,3)}

and it is a relation on A because it is a set of pairs from A times A. Nothing more is required.

A relation may hold between two different sets, in which case it is a subset of A times B. The transition of a machine is of that kind: it relates a pair (state, symbol) to a state.

The four properties, one at a time

Each is a property a relation may or may not have, and each is checked by asking one question.

Reflexive. Every element is related to itself: a R a for every a in A. "Is less than" is not reflexive, because 1 is not less than 1. "Is less than or equal to" is.

munotes.in16

Relations, and the Equivalence Relation That Runs Through This Subject

Symmetric. Whenever a is related to b, b is related to a. "Has the same length as" is symmetric. "Is less than" is not.

Antisymmetric. If a is related to b and b to a, then a and b are the same element. "Is less than or equal to" is antisymmetric. This one is here for completeness; it belongs to orderings rather than to equivalences and this book does not use it again.

Transitive. Whenever a is related to b and b to c, a is related to c. "Is less than" is transitive. "Is the parent of" is not.

A useful check on all four is that a relation can fail a property because of a single pair. To show a relation is not transitive it is enough to produce one a, b, c where the first two links hold and the third does not. To show it is transitive you must argue about all of them, which is why proofs go one way and counterexamples the other.

The equivalence relation

A relation is an equivalence relation when it is reflexive, symmetric and transitive, all three.

The three together say, in ordinary words, that the relation behaves the way equality behaves. Every thing is like itself. If one thing is like a second the second is like the first. And likeness passes along a chain.

Examples over strings, which is where this book uses them:

Relation on Sigma starEquivalence?Why
has the same lengthyesall three hold
has the same first symbolnofails reflexive for ε, which has no first symbol
has the same number of ayesall three hold
is a prefix ofnonot symmetric: a is a prefix of ab, ab is not a prefix of a
is shorter thannonot reflexive and not symmetric

Equivalence classes and the partition

This is the part that does the work in chapters 20 and 21.

Given an equivalence relation R on A and an element a of A, the equivalence class of a is the set of everything related to a. It is written with square brackets:

[a] = { x in A : x R a }

Three facts follow, and each is a one line proof from one of the three properties.

Every element is in its own class, because R is reflexive, so no element is left out.

Two classes are either identical or disjoint. Suppose some element x lies in both [a] and [b]. Then x R a and x R b. By symmetry a R x, and then by transitivity a R b, and from there every element of [a] is in [b] and conversely. So they are the same class. There is no such thing as two classes that partly overlap.

munotes.in17

Relations, and the Equivalence Relation That Runs Through This Subject

So the classes cut A into pieces with nothing left over and nothing counted twice. A family of non empty, pairwise disjoint sets whose union is the whole of A is called a partition of A. So every equivalence relation gives a partition, and going the other way, every partition gives an equivalence relation, namely "is in the same block as".

The number of classes is called the index of the relation. It may be finite even when A is infinite, and that single sentence is the whole content of the Myhill Nerode theorem in chapter 21: a language is regular exactly when a certain equivalence relation on the infinite set Sigma star has finitely many classes, and the number of classes is the number of states in the smallest machine.

Worked example

Let A be the set of strings over {a, b}, and let x R y mean "x and y have the same number of a, counted modulo 3", that is, the two counts leave the same remainder on division by 3.

Is it an equivalence relation? Reflexive, because any count leaves the same remainder as itself. Symmetric, because sameness of remainder does not depend on order. Transitive, because if the first and second agree, and the second and third agree, the first and third agree. So yes.

Its classes. The remainder can be 0, 1 or 2, so there are exactly three classes, whatever the strings look like:

ClassContainsSmallest members
[ε]strings whose a count leaves 0ε, b, bb, aaa, aab
[a]strings whose a count leaves 1a, ab, ba, aaaa
[aa]strings whose a count leaves 2aa, aab, aba, baa

So an infinite set has been cut into three blocks. That is an index of 3.

Now notice what the table is. It is a machine. Reading a b keeps you in your block. Reading an a moves you on one block, from [ε] to [a] to [aa] and back to [ε]. Three blocks, three states, and the block you are in is everything you need to remember. Chapter 13 builds exactly this machine, and chapter 21 proves that this is not a coincidence but a theorem.

Closures: reflexive, transitive, and the star

One more piece of notation, needed in chapter 23 for derivations.

Given a relation R that is missing a property, the closure of R with respect to that property is the smallest relation containing R that has it.

The reflexive closure adds the pair (a, a) for every a and nothing else. The transitive closure, written R plus, adds a pair (a, c) whenever R has a chain from a to c of one or more steps. The reflexive transitive closure, written R star, is the transitive closure with the reflexive pairs added, so it holds (a, c) whenever there is a chain of zero or more steps from a to c.

munotes.in18

Relations, and the Equivalence Relation That Runs Through This Subject

The zero step case is why R star relates every element to itself. This is exactly the notation a grammar uses: one arrow means one step of derivation, and the starred arrow means any number of steps including none, so a sentential form always derives itself.

Worked, on A = {1, 2, 3} with R = {(1,2), (2,3)}:

ClosureValue
R{(1,2), (2,3)}
reflexive closure{(1,1), (1,2), (2,2), (2,3), (3,3)}
R plus{(1,2), (2,3), (1,3)}
R star{(1,1), (1,2), (1,3), (2,2), (2,3), (3,3)}

The pair (1,3) appears in R plus because there is a two step chain from 1 to 3. It is not in R itself.

What it does NOT mean

An equivalence relation is not equality. It behaves like equality but it identifies things that are genuinely different: two states of a machine can be equivalent without being the same state, and that is the whole point of minimising.

Transitive does not mean reflexive. "Is less than" is transitive and not reflexive. The three properties are independent, and an examination answer that says a relation is an equivalence because it is transitive has proved a third of what was asked.

A partition's blocks are never empty and never overlap. A "partition" with an empty block, or with two blocks sharing an element, is not a partition, and the corresponding relation is not an equivalence.

The index can be finite when the set is infinite. This is not a paradox and it is the single most useful fact in the chapter.

Quick revision

  • A relation on A is any subset of A times A, that is, any set of ordered pairs. a R b means (a, b)

is in it.

  • Reflexive: a R a always. Symmetric: a R b gives b R a. Transitive: a R b and b R c give a R c.
  • Equivalence relation: all three of reflexive, symmetric and transitive.
  • The class [a] is everything related to a. Classes are non empty, pairwise disjoint, and cover A, so

they partition A. The number of them is the index.

  • Every partition gives an equivalence relation and every equivalence relation gives a partition.
  • An infinite set can have an equivalence relation of finite index, and that is what makes a finite

machine possible for an infinite language.

  • R star, the reflexive transitive closure, relates a to c when there is a chain of zero or more
munotes.in19

Relations, and the Equivalence Relation That Runs Through This Subject

steps, and is the notation a derivation uses.

Test yourself

1. Is "has the same last symbol as" an equivalence relation on Sigma star? Justify. No. It fails reflexivity for the empty string, which has no last symbol, so the pair (ε, ε) is not in the relation. On Sigma plus, where every string has a last symbol, it is an equivalence relation.

2. Give the equivalence classes of "same length modulo 2" on strings over {a}. Two classes: {ε, aa, aaaa, ...} and {a, aaa, aaaaa, ...}. The index is 2.

3. State the three properties and give a relation with exactly two of them. Reflexive, symmetric, transitive. "Is less than or equal to" on numbers is reflexive and transitive but not symmetric.

4. What is the index of a relation, and why does this subject care? The number of its equivalence classes. It matters because chapter 21 proves a language is regular exactly when a certain relation on Sigma star has finite index, and the index is then the number of states in the minimal machine.

5. For R = {(1,2), (2,1)} on {1, 2, 3}, give R star. {(1,1), (2,2), (3,3), (1,2), (2,1)}. The reflexive pairs are added for all three elements, and the two step chains from 1 to 1 and from 2 to 2 are already covered by those.

6. Why can two equivalence classes never partly overlap? If x lies in both [a] and [b] then x R a and x R b, so by symmetry and transitivity a R b, and the two classes contain exactly the same elements. Overlap therefore forces identity.

Contents This chapter on its own page

munotes.in20

Chapter Five

Functions: What the Transition Function Actually Is

Syllabus topic Module 1, "Mathematical Foundations (Sets, Relations, Functions, Proof Techniques)"

In one line

A function is a relation that gives exactly one answer for each input it is given.

In the wording a student can write in an examination: a function f from a set A to a set B is a relation from A to B in which every element of A appears as the first component of exactly one pair. A is the domain, B is the codomain, and we write f(a) for the unique element of B paired with a.

Why this subject needs functions

Because the definition of every machine in this book has a function in it, and the machine's whole behaviour is that function.

A finite automaton's transition function takes a state and a symbol and returns a state. A Mealy machine has a second function that takes a state and a symbol and returns an output symbol. A Turing machine's takes a state and a tape symbol and returns a state, a symbol and a direction. Learn the vocabulary here and those three definitions become three readings of the same sentence.

The definition, carefully

A function f from A to B is written f from A to B. For it to be a function, two things must hold for every element a of A:

At least one. There is some b in B with f(a) equal to b. Nothing in the domain is left without an answer.

At most one. There is only one such b. Nothing in the domain gets two answers.

Domain is A, the set the inputs come from. Codomain is B, the set the answers live in. Range or image is the part of B that is actually reached: the set of f(a) for all a in A. The range is a subset of the codomain and need not be all of it, and mixing the two up is the commonest error in this chapter.

Total and partial, which is the distinction that matters here

A function as defined above is called total: every element of the domain has an answer.

A partial function from A to B relaxes the first condition. Some elements of A may have no answer at all. It keeps the second condition: no element has two.

This is not a technicality in this subject; it is the difference between two ways of writing the same machine. A finite automaton is usually defined with a total transition function, so every state has a move on every symbol. In practice a machine is often drawn with some moves missing, meaning "if that happens, reject". Those two are the same machine in different clothes, and chapter 9 says how to turn one into the other: add a single extra state that is not final and whose every move returns to itself, then send all the missing moves there. It is called a dead state or a trap state, and the checker behind this book adds one automatically whenever it has to compare two machines.

munotes.in21

Functions: What the Transition Function Actually Is

One to one, onto, and both

Three properties a function may have. All three are asked directly in examinations, and two of them are used in this book's proofs.

One to one, also called injective: different inputs always give different answers. Equivalently, if f(x) equals f(y) then x equals y. The function that doubles a whole number is one to one. The function that gives the length of a string is not, because ab and ba both give 2.

Onto, also called surjective: every element of the codomain is reached, so the range is the whole codomain. Length, from strings over {a} to whole numbers, is onto, because every number is the length of some string of a's. Doubling, from whole numbers to whole numbers, is not onto, because 3 is not double anything.

Bijective: both one to one and onto. A bijection pairs the two sets off exactly, one element of A for one element of B with nothing left over on either side.

Bijections are how the two sizes of infinity in chapter 7 are defined. A set is countable when there is a bijection between it and the whole numbers, that is, when its elements can be listed in one endless sequence with none repeated and none missed.

PropertyMeansFails when
one to oneno two inputs share an answertwo different inputs give the same answer
ontoevery possible answer is usedsome element of the codomain is never produced
bijectivebotheither of the above fails

Composition

If f goes from A to B and g goes from B to C, then applying f and then g is a function from A to C, written g after f, with (g after f)(a) equal to g(f(a)).

Note the order: the one written on the right is applied first. It is written that way so that the notation matches the nesting of the brackets, and it catches students every time.

Composition is how the extended transition function of chapter 11 works. Running a machine on a string of length 5 is the transition function composed with itself five times, and the reason a machine's behaviour on a whole string is well defined is that composing functions gives a function.

Worked example

Let f be the function from strings over {a, b} to whole numbers that returns the number of a in the string.

Is it a function? Yes. Every string has a definite number of a in it, so there is at least one answer, and a string has only one count, so there is at most one.

munotes.in22

Functions: What the Transition Function Actually Is

Domain, codomain, range. The domain is all strings over {a, b}. The codomain is the whole numbers. The range is also the whole numbers, because for any n the string a to the n has exactly n of them.

One to one? No. f(ab) and f(ba) are both 1, and ab is not ba.

Onto? Yes, by the reason just given.

Now a transition function, which is the one that matters. Take the three state machine that counts a modulo 3 from chapter 4. Its transition function delta has domain {q0, q1, q2} times {a, b} and codomain {q0, q1, q2}:

C1Stateab
start finalq0q1q0
q1q2q1
q2q0q2

Accepts: ε, b, aaa, bbb, aaab

Rejects: a, aa, ab, aab

Every cell of C1 is filled, so delta is total, and every cell holds exactly one state, so it is a function rather than a relation. Six pairs in the domain, six entries in the table. That correspondence is worth noticing: a transition table is nothing but a function written out in full, and its number of cells is the size of the domain, which is the number of states multiplied by the number of symbols.

Is delta one to one? No, and it does not need to be: delta(q0, b) and delta(q1, a) are different inputs, and in this machine delta(q0, b) is q0 while delta(q1, a) is q2, but delta(q0, b) and delta(q0, b) aside, plenty of pairs collide in other machines. Nothing in the definition of an automaton requires the transition function to be one to one, and a machine whose transitions are one to one is a special and rather rare object.

What it does NOT mean

The codomain is not the range. The codomain is declared when the function is declared; the range is discovered by looking at what comes out. A function is onto exactly when they coincide.

A function is not required to be one to one. Most functions in this book are not.

A partial function is not a broken function. It is a perfectly respectable object and it is what a machine with missing moves has. What is not allowed, ever, is two answers for one input: that is a relation, and a machine whose transitions do that is nondeterministic, which is chapter 14.

f(a) is an element, f is a function. Writing "the function f(x)" is sloppy and it leads to confusion later when the same letter is used for a function on strings and a function on symbols.

munotes.in23

Functions: What the Transition Function Actually Is

Quick revision

  • A function from A to B pairs each element of A with exactly one element of B: at least one, at most

one.

  • Domain A; codomain B; range is the part of B actually reached, and is a subset of the codomain.
  • Total: every element of the domain has an answer. Partial: some may not, which is a machine with

missing moves, completed by adding a dead state.

  • One to one: different inputs, different answers. Onto: every element of the codomain is reached.

Bijective: both, and bijections define countability.

  • Composition g after f applies f first. The extended transition function is composition repeated.
  • A transition table is a function written out in full; it has one cell per pair of state and symbol.

Test yourself

1. Give the domain, codomain and range of the transition function of a DFA with 4 states over {0, 1}. Domain: the 4 states times {0, 1}, so 8 pairs. Codomain: the 4 states. The range is whichever states actually appear in the table, which may be fewer than 4.

2. Is the reversal function on strings one to one? Onto? Bijective? All three. Two different strings have different reversals, every string is the reversal of something, namely its own reversal, and so reversal is a bijection from Sigma star to itself.

3. A machine is drawn with no move from state q on symbol b. Is its transition function total? What do you do about it? It is partial. Add a dead state that is not final, send the missing move there, and give the dead state a move to itself on every symbol. The language accepted does not change.

4. Explain why a transition function may not be a relation that gives two answers. Because a function must give at most one answer per input. A transition that offers two next states is a relation, and a machine built on one is nondeterministic, which is a different definition with a set as its codomain.

5. If f has 8 pairs in its domain and 4 elements in its codomain, can f be one to one? No. Eight inputs cannot have eight distinct answers among four possible answers, so at least two inputs must share one. This is the pigeonhole principle of chapter 6.

Contents This chapter on its own page

munotes.in24

Chapter Six

Proof Techniques: Five Ways to Be Sure

Syllabus topic Module 1, "Mathematical Foundations (Sets, Relations, Functions, Proof Techniques)"

In one line

A proof is an argument that leaves no case unexamined, and there are five shapes it can take.

In the wording a student can write in an examination: a proof of a statement is a finite chain of reasoning from accepted premises to the statement, in which every step is justified. The standard techniques are direct proof, proof by contraposition, proof by contradiction, proof by induction and the pigeonhole principle.

Why this subject needs proof at all

Because a machine has to be right for every input, and there are infinitely many inputs.

You can test a program. You cannot test a machine on infinitely many strings, and the strings that break a machine are almost never short. Chapter 13 builds eight machines, and for each one the question "is this right?" is a question about an infinite set. Testing answers it for a few members. A proof answers it for all of them.

The other reason is that the negative results of this paper cannot be obtained any other way. "No machine of this kind accepts this language" is a statement about every machine there could ever be. There is no way to check them one at a time, because there are infinitely many of those too.

One: direct proof

Assume the hypothesis and reason forward to the conclusion.

Statement. If a string w has even length, then w reversed has even length.

Proof. Let w have even length, say length 2k for some whole number k. Reversal does not add or remove symbols, so w reversed has the same number of symbols as w, namely 2k. A number of the form 2k is even, so w reversed has even length.

That is a complete proof and it is three sentences long. The shape is always the same: name the assumption, name what it gives you, and arrive.

Two: proof by contraposition

To prove "if P then Q", prove "if not Q then not P" instead. The two statements are logically the same, and sometimes the second is much easier to argue.

Statement. If the number of a in w is odd then w is not the empty string.

Contrapositive. If w is the empty string then the number of a in w is not odd.

Proof of the contrapositive. If w is the empty string it has no symbols, so it has 0 occurrences of a, and 0 is even, so the count is not odd.

The original statement would have required reasoning about every odd count. The contrapositive has exactly one case.

The trap. The contrapositive of "if P then Q" is "if not Q then not P". It is NOT "if not P then not Q", which is a different statement called the inverse and is not equivalent to the original. This is the single commonest logical error students make in this paper, and chapter 36 is where it costs marks: the pumping lemma says "if regular then pumpable", so the useful contrapositive is "if not pumpable then not regular", and a student who writes "if not regular then not pumpable" has said something both different and false.

munotes.in25

Proof Techniques: Five Ways to Be Sure

Three: proof by contradiction

Assume the statement is false, derive something impossible, and conclude the statement is true.

This is the technique the two most important results in the paper use, so it deserves the most attention. The halting problem in chapter 74 is a proof by contradiction, and so is every application of the pumping lemma in chapter 37.

Statement. There is no largest whole number.

Proof. Suppose there were. Call it N. Then N plus 1 is a whole number and is larger than N, which contradicts N being the largest. So there is no largest whole number.

The shape, in four lines, which is worth learning as a template:

  1. Suppose, for contradiction, that the statement is false.
  2. Write out exactly what that assumption gives you.
  3. Derive two things that cannot both hold.
  4. Conclude that the assumption was wrong, so the statement is true.

Step 2 is where students lose marks. "Suppose the language is regular" is not enough; the next line has to say what being regular gives you, namely a machine with some definite finite number of states, and it is that number the rest of the argument attacks.

Four: proof by induction

Used to prove a statement about every whole number, or about every string, by proving it for the smallest case and then proving that each case carries the next one.

The two parts.

Base case. Prove the statement for the starting value, usually 0 or 1, or for the empty string.

Inductive step. Assume the statement holds for some value n, which is called the inductive hypothesis, and prove it holds for n plus 1.

Once both are done, the statement holds for every value from the base upwards, because the base gives the first, the step carries the first to the second, the second to the third, and so on without end.

Statement. The number of strings of length exactly k over an alphabet of size 2 is 2 to the k.

Base case. For k equal to 0 there is exactly one string of length 0, the empty string, and 2 to the 0 is 1. True.

Inductive step. Assume there are 2 to the k strings of length k. A string of length k plus 1 is a string of length k with one more symbol on the end, and there are 2 choices for that symbol. Different choices give different strings, and every string of length k plus 1 arises this way exactly once. So there are 2 to the k multiplied by 2, which is 2 to the (k plus 1). True.

munotes.in26

Proof Techniques: Five Ways to Be Sure

Therefore the statement holds for every k.

Induction on strings. Very often in this book the induction is not over numbers but over the length of a string. The base case is the empty string and the step assumes the claim for w and proves it for w with one more symbol. Chapter 11 defines the extended transition function exactly this way, and chapter 16 proves the subset construction correct exactly this way. It is the same principle: the strings of each length are built from the strings one shorter.

Five: the pigeonhole principle

If n items are put into m boxes and n is greater than m, then some box holds at least two items.

It is obvious, it needs no proof, and it is the engine of the most examined theorem in Module 1.

Why it matters here. A machine with m states, run on a string of length m or more, visits at least m plus 1 states along the way, counting the one it starts in. There are only m states to be visited. So some state is visited twice. That is the entire content of the pumping lemma, and chapter 36 does nothing but write it out carefully and see what follows.

Worked, on a machine. A machine has 3 states. Feed it the string aaaa, of length 4. The run passes through 5 states in total: where it starts, and where it is after each of the 4 symbols. Five visits, three states, so some state occurs twice. Between the two occurrences the machine has read a non empty piece of the string and come back to where it was, which means that piece can be repeated any number of times without the machine noticing. Hence "pumping".

Putting two together: proving a language is not regular

Here is the combination that earns marks in the examination, using three of the five techniques at once. The full treatment is chapter 37; this is the skeleton.

To prove that the language of strings a to the n followed by b to the n is not regular:

  1. Contradiction: suppose it is regular.
  2. Then some machine accepts it, with some definite number of states, call it m.
  3. Pigeonhole: feed the machine a to the m followed by b to the m. While reading the a part it
munotes.in27

Proof Techniques: Five Ways to Be Sure

passes through more states than it has, so it repeats one.

  1. So a non empty block of a can be repeated without the machine noticing, giving a string with more a

than b that the machine still accepts.

  1. That string is not in the language, so the machine does not accept the language after all.
  2. Contradiction, so no such machine exists, so the language is not regular.

Notice that step 2 says "some definite number", not "a large number". The proof must work whatever m is, which is why the string chosen in step 3 is written in terms of m rather than as a particular string.

Distinctions

TechniqueProve whatTypical use in this book
Directthe statement itselfsimple closure properties
Contrapositionif not Q then not Pusing the pumping lemma the right way round
Contradictionassume false, hit an impossibilitythe halting problem, the pumping lemma
Inductionbase case plus stepcorrectness of a construction, over string length
Pigeonholemore items than boxesthe pumping lemma, and state count bounds

What it does NOT mean

A proof is not a worked example. Showing a machine accepts three strings proves that it accepts three strings. It does not prove the machine is correct, and an examination answer that offers examples where a proof was asked for has answered a different question.

Contraposition is not the inverse. "If not P then not Q" is not equivalent to "if P then Q". See the trap above.

A counterexample is not a proof, except of a negative. One counterexample disproves a universal claim completely. No number of confirming instances proves one.

Induction does not prove the base case by assuming it. The base case must be checked outright. An induction whose base case is skipped can prove false statements, which is why examiners mark it separately.

Quick revision

  • Direct: assume the hypothesis, reason to the conclusion.
  • Contraposition: prove "if not Q then not P", which is equivalent. The inverse is not.
  • Contradiction: assume the negation, write out what it gives, derive an impossibility. Used for the

halting problem and every pumping lemma argument.

  • Induction: base case, then assume for n and prove for n plus 1. Often over string length, with the

empty string as the base.

  • Pigeonhole: more items than boxes forces a repeat. A machine with m states run on a string of length

m or more must revisit a state, and that is the pumping lemma.

  • A proof is needed because a machine must be right on infinitely many inputs, and because negative

results are claims about every possible machine.

Test yourself

1. Give the contrapositive of "if L is regular then L satisfies the pumping lemma". If L does not satisfy the pumping lemma then L is not regular. This is the form in which the lemma is actually used.

munotes.in28

Proof Techniques: Five Ways to Be Sure

2. Why must the base case of an induction be proved separately? Because the inductive step only shows that each case carries the next. Without a first case that is known outright, the chain has nothing to start from, and a step can be perfectly valid while the statement is false for every value.

3. A machine has 5 states. What does the pigeonhole principle tell you about its run on a string of length 6? The run visits 7 states counting the start, among only 5 distinct states, so at least one state is visited twice, and the portion of input read between the two visits can be repeated without changing where the machine ends up.

4. Prove by contradiction that the empty string is in every language of the form L star. Suppose some L star did not contain the empty string. L star is the set of concatenations of zero or more strings from L, and the concatenation of zero strings is the empty string by definition. So the empty string is in L star, contradicting the supposition.

5. Name the technique each of these needs: showing a construction is correct for all inputs; showing a language is not context free; showing two regular expressions denote the same language. Induction, usually on string length; contradiction with the pumping lemma for context free languages; either a direct two way containment argument or, in this book, an exact machine comparison.

6. State the pigeonhole principle and say what the boxes are when it is applied to a finite automaton. If more items than boxes are distributed among the boxes, some box holds two or more. Applied to an automaton, the boxes are the states and the items are the points in time during a run.

Contents This chapter on its own page

munotes.in29

Chapter Seven

Counting the Uncountable: Diagonalisation, and Why Some Languages Have No Machine

Syllabus topic Module 1, "Mathematical Foundations (Sets, Relations, Functions, Proof Techniques)"

In one line

There are more languages than there are machines, so most languages have no machine, and this is provable by counting.

In the wording a student can write in an examination: a set is countable if its elements can be arranged in a single infinite list, that is, if there is a one to one correspondence between the set and the natural numbers. The set of all programs over a finite alphabet is countable, the set of all languages over a finite alphabet is uncountable, and therefore there exist languages for which no program exists.

Why a chapter on counting belongs in this paper

Because of the order in which the two halves of this result arrive.

The rest of this book proves specific negative results. Chapter 37 proves one particular language is not regular. Chapter 74 proves one particular problem has no algorithm. Each of those takes a page of careful work, and each one is about one language.

This chapter proves something much larger in half a page, and it proves it before any machine has been examined: almost every language is beyond every machine. The specific results that follow are then not surprises. They are examples of a situation that was already known to be the normal one.

Two sizes of infinity

The word infinite hides a distinction, and Georg Cantor found it.

A set is countable when its members can be written out in one endless list, a first, a second, a third, with nothing repeated and nothing left out for ever. Equivalently, using the vocabulary of chapter 5, a set is countable when there is a bijection between it and the natural numbers.

A set is uncountable when no such list exists. Not "when nobody has found one", but when it can be proved that none exists.

The whole numbers are countable, trivially. So are the even numbers, by listing 0, 2, 4, 6 and so on, even though they are a proper subset of the whole numbers. That is the first thing about infinite sets that feels wrong and is not: a proper subset can be the same size.

Sigma star is countable

This is the easy half, and it was already done in chapter 3.

Take any alphabet Sigma, with n symbols. List the strings in canonical order: first by length, and within each length alphabetically. There is exactly one string of length 0, n of length 1, n squared of length 2, and so on, and each of these groups is finite, so the listing never gets stuck.

ε, a, b, aa, ab, ba, bb, aaa, aab, ...

Every string appears, because a string of length k appears in the k-th group, and no string appears twice. So Sigma star is countable.

munotes.in30

Counting the Uncountable: Diagonalisation, and Why Some Languages Have No Machine

The set of all machines is countable

This is the step that makes the argument bite, and it is worth doing carefully.

A machine of any of the kinds in this book is a finite object: a finite set of states, a finite alphabet, and a finite table of transitions. A finite object can be written down as a string over some fixed finite alphabet. That is not a trick; it is what chapter 73 does explicitly, encoding a whole Turing machine as a single string so that another machine can read it.

Once each machine is a string, the set of machines is a subset of Sigma star for that encoding alphabet. And a subset of a countable set is countable: list the whole set and cross out what is not in the subset, and what remains is still a list.

So the machines can be listed: a first machine, a second machine, a third, with every machine somewhere in the list.

The same argument applies to programs in any programming language, since a program is a finite string of characters. There are countably many programs. Every program that has ever been written or ever will be written, in every language, is somewhere in one list.

The set of all languages is uncountable

Now the other side, and this is Cantor's diagonal argument.

A language over Sigma is any subset of Sigma star. So the set of all languages is the power set of Sigma star. Chapter 2 counted the subsets of a finite set; this is the infinite case, and the answer is different in kind.

Claim. The set of all languages over Sigma is not countable.

Proof by contradiction. Suppose it were countable. Then the languages could be listed, L1, L2, L3 and so on, with every language somewhere in the list. Also list the strings in canonical order, w1, w2, w3 and so on, which chapter 3 says we can do.

Now build a table. The row for language Li and the column for string wj holds 1 if wj is in Li and 0 if it is not. Every language is a row, because every language is in the list, and every string is a column.

w1w2w3w4
L11011
L20010
L31100
L40111

Now define a new language D by walking down the diagonal and flipping every entry: put wj into D exactly when the entry for Lj and wj is 0. For the table above the diagonal reads 1, 0, 0, 1, so D contains w2 and w3 and not w1 and not w4.

munotes.in31

Counting the Uncountable: Diagonalisation, and Why Some Languages Have No Machine

D is a language over Sigma, because it is a set of strings over Sigma. So D must be somewhere in the list, say D is Lk.

Ask whether wk is in D. By the definition of D, wk is in D exactly when the table entry for Lk and wk is 0, that is, exactly when wk is not in Lk. But D is Lk, so that reads: wk is in Lk exactly when wk is not in Lk. That cannot be.

So the assumption was wrong, and the languages cannot be listed. The set of all languages is uncountable.

The conclusion, in one sentence

Countably many machines. Uncountably many languages. Each machine accepts exactly one language. So the machines cannot cover the languages, and almost every language has no machine at all.

"Almost every" is not loose here. The languages with machines form a countable subset of an uncountable set, which is as small as a subset can be while still being infinite.

What this does and does not tell you

It is a counting argument, and counting arguments have a characteristic strength and a characteristic weakness.

The strength. It settles the general question at once, and with almost no work. There is no need to examine what machines can do in order to know that they cannot do everything.

The weakness. It is not constructive. It proves that undecidable languages exist without naming one. If you ask "which language has no machine?", this argument answers "one of the uncountably many that are not in the list", which is no answer at all.

That is exactly why chapters 74 to 77 are needed. The halting problem is a named problem, with a plain English statement anyone can understand, and it is proved undecidable by the same diagonal technique applied to a particular construction rather than to a table. Naming one is much harder than counting them, and it is the naming that made Turing's paper famous.

Worked example: a small diagonal, in full

Work over the alphabet {a}, so the strings in canonical order are

w1 = ε, w2 = a, w3 = aa, w4 = aaa, w5 = aaaa

Suppose somebody offers this as the beginning of a complete list of all languages over {a}:

εaaaaaa
L11111
L20101
L31100
L40001

The diagonal entries are the ones where the row number and the column number agree: 1 for L1 and ε, 1 for L2 and a, 0 for L3 and aa, 1 for L4 and aaa. Flipping gives 0, 0, 1, 0, so

munotes.in32

Counting the Uncountable: Diagonalisation, and Why Some Languages Have No Machine

D = {aa, ...}

and D is not L1, because they disagree on ε; not L2, because they disagree on a; not L3, because they disagree on aa; and not L4, because they disagree on aaa. Whatever the rest of the list holds, D disagrees with the k-th language on the k-th string, so it is nowhere in the list.

That is the whole argument, and the reason it works is that the diagonal is guaranteed to meet every row exactly once.

Distinctions

CountableUncountable
Definitioncan be listed with nothing missedno list is possible
Example hereSigma star, the machines, the programsthe set of all languages
Proved byexhibiting a listingCantor's diagonal argument
Consequencea machine can be encoded as a stringmost languages have no machine

What it does NOT mean

It does not say the languages we care about have no machine. Almost all languages are undescribable, but the languages anybody wants to describe are, by definition, describable, so they are in the small countable part. This is not a reason to despair.

It does not name a single undecidable language. That needs chapter 74.

It does not depend on the machine model. It works for finite automata, for Turing machines, for Java programs and for anything else that can be written down finitely. That robustness is the point, and it is what connects this chapter to the Church Turing thesis in chapter 72.

A proper subset of an infinite set can be the same size. The even numbers are countable, exactly as the whole numbers are. Size for infinite sets means "there is a bijection", not "is contained in".

Quick revision

  • Countable: the elements can be put in one endless list with nothing missed. Uncountable: no list

exists, and that can be proved.

  • Sigma star is countable, in canonical order, first by length and then alphabetically.
  • Every machine and every program is a finite object, so it can be written as a string, so there are

countably many of them.

  • The set of all languages is the power set of Sigma star and is uncountable, by Cantor's diagonal

argument: from any proposed list build D disagreeing with the k-th language on the k-th string.

  • Countably many machines against uncountably many languages, so almost every language has no machine.
  • The argument is not constructive: it proves such languages exist and names none, which is why the

halting problem is proved separately.

Test yourself

1. Define countable, and show the set of even whole numbers is countable. A set is countable when its elements can be listed in one infinite sequence with none repeated and none omitted, equivalently when there is a bijection with the natural numbers. The evens are listed 0, 2, 4, 6 and so on; the function taking n to 2n is a bijection from the natural numbers to the evens.

munotes.in33

Counting the Uncountable: Diagonalisation, and Why Some Languages Have No Machine

2. Why are there only countably many Turing machines? Each is a finite object and can be written as a finite string over a fixed finite alphabet. The set of all finite strings over a finite alphabet is countable, and a subset of a countable set is countable.

3. Carry out one step of the diagonal argument: the list begins L1 with ε in it, L2 without a in it, L3 with aa in it. Is ε in D? Is a? Is aa? The diagonal entry for L1 and ε is 1, so ε is not in D. For L2 and a it is 0, so a is in D. For L3 and aa it is 1, so aa is not in D.

4. Why does the diagonal argument not name an undecidable language? Because the language it produces is defined in terms of the supposed list, which does not exist. The argument shows a contradiction follows from assuming the list, and a contradiction yields no object.

5. A student concludes from this chapter that most useful problems are unsolvable. Are they right? No. The conclusion is that most languages, taken from the whole uncountable collection, have no machine. A language somebody has described is describable, so it lies in the countable part, and whether it has a machine is a separate question answered by the specific results of Module 2.

6. Which two facts, put together, give the conclusion of this chapter? That there are countably many machines, because each is a finite object encodable as a string, and that there are uncountably many languages, because the power set of a countable infinite set cannot be listed.

Contents This chapter on its own page

munotes.in34

Chapter Eight

What an Automaton Is

Syllabus topic Module 1, "Automata Theory: Defining Automaton"

In one line

An automaton is a machine with a fixed, finite amount of memory, which reads an input one symbol at a time and ends up either accepting or rejecting it.

In the wording a student can write in an examination: an automaton is an abstract computing device consisting of a finite set of states, an input alphabet, a rule that determines the next state from the present state and the present input symbol, a designated initial state, and a set of final or accepting states. It reads the input from left to right, one symbol per move, and accepts the input if it is in a final state when the input is exhausted.

Why the definition has to be this careful

Two machines that behave identically must be the same machine, and two machines that differ must differ in the definition. That is what a formal definition buys, and it is why the drawings come second.

There is a practical reason as well. Chapter 13 asks you to design machines, and design is impossible without knowing exactly what you are allowed to build. The five parts are the parts list.

The five parts, one at a time

One. A finite set of states, Q. The machine has memory, and this set is all of it. Whatever the machine knows at any moment, it knows because of which state it is in, and nothing else. There are finitely many of them, so the machine can only remember finitely many different situations. That single restriction is responsible for every negative result in Module 1.

Two. A finite input alphabet, Sigma. The symbols the machine will be shown, from chapter 3. The machine sees one per move and can never look back.

Three. A transition rule, delta. Given where the machine is and what it is reading, this says where it goes. It is a function from a state and a symbol to a state, using the vocabulary of chapter 5. This is the machine's program, and it is the only part that varies from one interesting machine to another.

Four. An initial state, q0. One particular state from Q, the one the machine is in before it has read anything. Exactly one, never two, and every machine has one.

Five. A set of final states, F. A subset of Q. When the input runs out, the machine accepts if it is in one of these and rejects otherwise. F may be empty, in which case the machine accepts nothing. It may be all of Q, in which case it accepts everything.

Written together, an automaton is the five things in a fixed order, which is why it is called a quintuple:

munotes.in35

What an Automaton Is

M = (Q, Sigma, delta, q0, F)

What each part rules out

The definition is as much about what a machine cannot do as about what it can, and the restrictions are the interesting part.

It cannot count without bound. The memory is the state, and there are finitely many states. A machine can count a to three and remember the answer modulo three, as chapter 4 showed, because that needs three states. It cannot count a to an arbitrary number and compare with the b count, because that would need one state per possible count. Chapter 37 turns this observation into a proof.

It cannot go back. The input is read once, left to right, one symbol per move. There is no rewinding and no looking ahead. Everything about the part already read must have been summarised into the current state, or it is gone.

It cannot write. There is no scratch paper. The pushdown automaton of chapter 52 adds a stack and the Turing machine of chapter 62 adds a writable tape, and each addition buys exactly the extra power you would expect.

It has no clock and no cost. A move takes one move. This model says nothing about time in seconds, and questions about its speed belong to the complexity block of Module 2.

Three machines you have already used

The definition is abstract, so it is worth seeing that it describes ordinary things.

A turnstile. Two states: locked and unlocked. Two inputs: a coin and a push. From locked, a coin gives unlocked and a push gives locked. From unlocked, a push gives locked and a coin gives unlocked. Five parts, all present, and a complete description of the object.

A lift with three floors. States: which floor it is on. Inputs: the buttons. What it does next depends only on where it is and which button was pressed.

A four digit PIN pad. States: how many correct digits have been entered so far, plus one state for having got it wrong. Final state: all four correct. Notice what is NOT remembered: which digits were typed. Once a digit is known to be right, the fact that it was a 7 is not needed again, and a machine that remembered it would have more states for no gain. Deciding what need not be remembered is the whole skill of chapter 13.

The three ways an automaton is written down

All three appear throughout this book and a student must be able to convert between any two of them. Chapter 10 does the conversions properly; they are named here because the next chapter uses all three.

The quintuple, as above, with delta given by a list of equations.

munotes.in36

What an Automaton Is

The transition table, which is delta written out as a grid: one row per state, one column per symbol, the next state in the cell. The start state and the final states are marked in the margin.

The transition diagram, a drawing with a circle per state, an arrow per transition labelled with its symbol, an arrow from nowhere into the start state, and a double circle round each final state.

Worked example

Build the turnstile formally. Call the states L for locked and U for unlocked, and the alphabet {c, p} for coin and push. Let L be the initial state, and let U be the only final state, so the machine accepts exactly those input sequences that leave the turnstile unlocked.

Q = {L, U}

Sigma = {c, p}

q0 = L

F = {U}

and delta given by four equations, one for each combination:

delta(L, c) = U

delta(L, p) = L

delta(U, c) = U

delta(U, p) = L

Four equations, because there are two states and two symbols and the function must be defined on every pair. As a table:

A1Statecp
startLUL
finalUUL

Accepts: c, cc, pc, ppc, cpc

Rejects: ε, p, cp, pp, cpp

Read those lists back against the object: the machine ends unlocked exactly when the last symbol was a coin. The empty input is rejected because the turnstile starts locked, and every input ending in a push is rejected because a push always locks it.

And now the language, which is the point of the whole exercise. The machine accepts a string exactly when its last symbol is c, so

L(A1) = L(R1)

where R1 is the regular expression for that description:

(c + p)* c

That equality is not asserted here; the checker decides it exactly, by building both machines and comparing them over every string there is.

Distinctions

StateInput symbol
What it iswhat the machine rememberswhat the machine is shown
How manyfinitely many, fixed in advancefinitely many, fixed in advance
Changes during a runyes, once per symbolno, the alphabet is fixed
Chosen bythe designerthe problem
Initial stateFinal state
How manyexactly onezero or more
Writtenan arrow from nowherea double circle
Used whenbefore reading beginswhen the input is exhausted

What it does NOT mean

An automaton is not a program. It has no variables, no arithmetic and no subroutines. Its entire memory is which of finitely many states it is in.

Final does not mean last. A final state is one that accepts if the input happens to run out there. The machine may pass through final states in the middle of a run and end up rejecting, and it may pass through non final states and end up accepting. Only the state at the end matters.

munotes.in37

What an Automaton Is

A machine need not stop at a final state. There is no halting here. The machine reads until the input is finished, and then the answer is read off.

An automaton does not produce output. It answers yes or no. The Mealy and Moore machines of chapter 18 are the variant that writes as it goes, and they are a different definition.

Quick revision

  • An automaton is a quintuple: states Q, alphabet Sigma, transition rule delta, initial state q0, final

states F.

  • Q is finite, and it is the machine's entire memory. That restriction causes every negative result in

Module 1.

  • delta takes a state and a symbol and gives a state; it is the machine's program.
  • Exactly one initial state; F may be empty or all of Q.
  • The input is read once, left to right, one symbol per move; no rewinding, no writing, no counting

without bound.

  • Three representations: quintuple, transition table, transition diagram.
  • Accepting depends only on the state when the input runs out.

Test yourself

1. List the five parts of an automaton and say what each is for. A finite set of states, which is the whole memory; an input alphabet, the symbols it will see; a transition rule, which is its program; one initial state, where it starts; and a set of final states, which decide acceptance when the input ends.

2. Why must Q be finite, and what does that stop the machine doing? Because the machine's memory is which state it is in, and a machine with infinitely many states would have unbounded memory. It stops the machine from counting without bound, which is why no finite automaton can check that two counts agree.

3. A machine has 3 states and an alphabet of 2 symbols. How many equations does delta need? Six, one for each pair of state and symbol, because delta must be defined on the whole of Q times Sigma.

4. Can a machine have no final states? What language does it accept? Yes. It accepts the empty language, because acceptance requires ending in a final state and there are none.

5. A student says a machine accepted the string because it visited a final state after the second symbol. Correct them. Only the state reached when the whole input has been read decides acceptance. Visiting a final state part way through means nothing, and the string is accepted only if the final symbol leaves the machine in a final state.

munotes.in38

What an Automaton Is

6. Describe a four digit PIN pad as an automaton, and say what it deliberately does not remember. States: the number of correct digits entered so far, from zero to four, plus a failure state. Initial state: zero correct. Final state: four correct. It does not remember which digits were typed, because once a digit is known to be correct its identity is never needed again.

Contents This chapter on its own page

munotes.in39

Chapter Nine

The Deterministic Finite Automaton, Formally

Syllabus topic Module 1, "Automata Theory: Finite Automaton"

In one line

A deterministic finite automaton is an automaton in which, for every state and every symbol, there is exactly one next state.

In the wording a student can write in an examination: a deterministic finite automaton, or DFA, is a quintuple M = (Q, Sigma, delta, q0, F) where Q is a finite set of states, Sigma is a finite input alphabet, delta is a function from Q times Sigma to Q, q0 in Q is the initial state and F, a subset of Q, is the set of final states. Determinism means that delta is a total function, so at every step the next state is uniquely determined.

Why the word deterministic is in the name

Because it will not always be true.

In the definition above, delta is a function from Q times Sigma to Q, and chapter 5 says a function gives exactly one answer for every input it is defined on. So the machine never has a choice: given where it is and what it reads, there is one place it can go. Run the same machine on the same string twice and the identical sequence of states results.

Chapter 14 replaces delta with a relation, so that a state and a symbol may lead to several states, or to none. That machine is nondeterministic, and the word deterministic exists to mark the difference. Chapter 16 then proves the surprising thing: the two kinds accept exactly the same languages.

The five parts, with the determinism made explicit

M = (Q, Sigma, delta, q0, F)

PartWhat it isConstraint
Qthe finite set of statesfinite, non empty
Sigmathe input alphabetfinite, non empty
deltathe transition functiona total function from Q times Sigma to Q
q0the initial stateexactly one, and a member of Q
Fthe set of final statesany subset of Q, possibly empty

Two consequences of delta being a total function are worth stating as rules, because an examiner will check both.

Every state must have an outgoing move on every symbol. A table with a blank cell does not describe a DFA. Chapter 5 gave the remedy: add a dead state, not final, with a move to itself on every symbol, and send the blanks there. The language does not change and the machine becomes a DFA.

No state may have two moves on one symbol. A cell with two entries in it does not describe a DFA either; it describes a nondeterministic machine.

How a DFA runs

The machine begins in q0 with the whole input unread. Then, over and over, it reads the leftmost unread symbol, and moves to the state delta says. When there are no symbols left, the run stops.

munotes.in40

The Deterministic Finite Automaton, Formally

The input is accepted if the state at that moment is in F, and rejected otherwise. The run always stops, after exactly as many moves as the input has symbols, so a DFA always gives an answer, and always in a predictable number of steps. That is a property no later machine in this book has, and it is worth noticing now so that its loss in chapter 64 is visible.

The language of a machine

The language accepted by M, written L(M), is the set of all strings M accepts:

L(M) = { w in Sigma star : M ends in a final state after reading w }

Three points about this definition catch students.

It is a set of strings, so it is a language in the sense of chapter 3. It is determined entirely by the machine, so a machine has exactly one language. And the same language can be accepted by many different machines, which is why chapter 20 asks which of them is smallest.

A language is called regular when some DFA accepts it. That is the definition, and the whole of Module 1 from here to chapter 40 is an investigation of which languages are regular and which are not.

Worked example: the first machine, built from the definition

The problem. Over the alphabet {0, 1}, accept exactly those strings that end in 1.

What must be remembered? Only whether the symbol just read was a 1. Nothing else about the past can matter, because whether a longer string ends in 1 depends only on its last symbol. So two states suffice: one meaning "the last symbol was not a 1, or there have been no symbols", and one meaning "the last symbol was a 1".

The five parts.

Q = {q0, q1}

Sigma = {0, 1}

q0 is the initial state

F = {q1}

and delta by four equations:

delta(q0, 0) = q0

delta(q0, 1) = q1

delta(q1, 0) = q0

delta(q1, 1) = q1

As a table, with the start and final markings in the margin:

B1State01
startq0q0q1
finalq1q0q1

Every cell is filled and every cell holds one state, so delta is a total function and this is a DFA.

Accepts: 1, 01, 11, 001, 101, 0101

Rejects: ε, 0, 10, 00, 110, 1010

The empty string is rejected because the machine starts in q0, which is not final. That is the case students forget, and it is the case the checker always runs.

And the language, proved rather than described. The regular expression for "ends in 1" is:

munotes.in41

The Deterministic Finite Automaton, Formally

(0 + 1)* 1

L(B1) = L(B1R)

The checker decides that equality exactly: it builds a deterministic machine from each side, forms the product, and searches for any string on which the two disagree. There is none, so the machine is correct on all of the infinitely many strings, not just on the eleven listed above.

A trace, step by step

Running the machine on 0101, written as the state and then what is left to read:

q0 0101
q0 101
q1 01
q0 1
q1 ε

Five configurations for four symbols, because the run includes where it started. The last state is q1, which is final, so 0101 is accepted. The checker re-executes this trace against the table above and fails if any single step is wrong.

Distinctions

DFAThe general automaton of chapter 8
delta isa total functiona rule, which may be a relation
Next stateexactly oneone, several, or none
A blank table cellnot allowedallowed, meaning reject
Runs on a stringone possible runpossibly many
Accepts the stringAccepts the language
Meaningends in a final state on this one inputthe set of all inputs it ends finally on
WrittenM accepts wL(M)
Checked byone runa proof, or an exact machine comparison

What it does NOT mean

Deterministic does not mean simple, small or fast. It means the next state is always uniquely determined. A DFA can have a thousand states.

A DFA does not stop when it reaches a final state. It reads the whole input. Only the state at the end counts.

A missing transition is not a rejection rule in a DFA. It is an incomplete definition. Complete it with a dead state.

Regular is a property of a language, not of a machine. Every DFA's language is regular. Asking whether a machine is regular is a category error; ask whether a language is.

Quick revision

  • A DFA is (Q, Sigma, delta, q0, F) with delta a total function from Q times Sigma to Q.
  • Determinism: exactly one next state for every state and symbol. No blanks, no double entries.
  • It reads the input once, left to right, one symbol per move, and always stops after as many moves as

there are symbols.

  • Accepts if the state when the input runs out is in F.
  • L(M) is the set of strings M accepts. A language is regular when some DFA accepts it.
  • An incomplete table is completed by a dead state, which is not final and returns to itself on every

symbol.

  • The empty string is accepted exactly when q0 is in F.

Test yourself

1. State the definition of a DFA, naming all five parts and the constraint on delta. A quintuple (Q, Sigma, delta, q0, F): Q a finite set of states, Sigma a finite input alphabet, delta a total function from Q times Sigma to Q, q0 in Q the initial state, and F a subset of Q of final states. The constraint is that delta is total and single valued, so the next state is always uniquely determined.

munotes.in42

The Deterministic Finite Automaton, Formally

2. When does a DFA accept the empty string? Exactly when its initial state is a final state, since no moves are made and the run ends where it began.

3. A table for a 3 state machine over {a, b} has 5 filled cells. Is it a DFA? No. It needs 6, one per state and symbol. Add a dead state and send the missing move to it.

4. Build a DFA accepting strings over {a, b} that begin with a. Three states: the start, a state meaning the first symbol was a, and a dead state. From the start, a goes to the accepting state and b to the dead state. The accepting state goes to itself on both symbols; the dead state goes to itself on both. Only the second state is final.

5. How many runs does a DFA have on a given string, and why does that matter? Exactly one, because delta gives exactly one next state at every step. It matters because acceptance is then unambiguous with no search, which is what makes a DFA the machine you implement in practice.

6. Two different DFAs accept the same language. Is that possible, and what does chapter 20 ask about it? Yes, and there are infinitely many such machines for any regular language. Chapter 20 asks which is smallest, and chapter 21 proves the smallest is unique up to the naming of its states.

Contents This chapter on its own page

munotes.in43

Chapter Ten

Transitions and Their Properties: the Function, the Table and the Diagram

Syllabus topic Module 1, "Automata Theory: Transitions and Its properties"

In one line

A transition says where the machine goes next, and it can be written three ways: as an equation, as a row of a table, or as an arrow in a drawing.

In the wording a student can write in an examination: the transition function delta of a finite automaton maps Q times Sigma to Q. It may equivalently be presented as a transition table, whose rows are states, whose columns are input symbols and whose entries are next states, or as a transition diagram, a directed graph whose vertices are states and whose labelled edges are transitions.

Why three representations

Because each is better at a different job, and the examination uses all three.

The equations are what a proof manipulates. When chapter 16 proves the subset construction correct, it argues about delta, not about a drawing.

The table is what a machine is built from and what a program implements. It is also the form that cannot hide a mistake: a table with a missing cell is visibly missing a cell, while a diagram with a missing arrow looks finished. Every machine in this book is printed as a table for that reason, and every one of those tables is read by a program and executed.

The diagram is what a human designs with. You can see a cycle in a drawing and you cannot see it in a table.

Students are asked to convert between them constantly, so the conversions are set out below as procedures.

The transition function

For a DFA, delta is a total function from Q times Sigma to Q, written

delta(q, a) = r

which reads: from state q, on reading symbol a, go to state r.

The size of delta is fixed by the machine and is worth computing before you start filling a table in: the number of equations is the number of states multiplied by the number of symbols. Three states over a two symbol alphabet needs six, and if you have written five you have left a case out.

The transition table, and how this book writes it

The convention used throughout:

D1Stateab
startq0q1q0
q1q1q2
finalq2q1q0

The top left cell names the machine. Here it is D1, and a claim about the machine refers to it by that name.

The second column is the state. One row per state, in a sensible order with the start state first.

The third column is empty in every row, which is what draws the double rule between a state and its moves, and what keeps the table a grid on a phone instead of collapsing into a list.

munotes.in44

Transitions and Their Properties: the Function, the Table and the Diagram

The remaining columns are the input symbols, one each, and the cell holds the next state.

The first column carries the markers: start for the initial state, final for an accepting state, start final for a state that is both, and nothing at all for an ordinary state.

Reading that table off as equations gives six of them, which is the three states times the two symbols:

delta(q0, a) = q1

delta(q0, b) = q0

delta(q1, a) = q1

delta(q1, b) = q2

delta(q2, a) = q1

delta(q2, b) = q0

Accepts: ab, aab, bab, abab

Rejects: ε, a, b, ba, abb

What does D1 do? It accepts a string exactly when the string ends in ab. Reading b from q0 keeps you nowhere; reading a moves you to q1, meaning "an a has just been read"; reading b from q1 reaches q2, the accepting state, meaning "ab has just been read". And from q2, reading a starts a fresh a, while reading b takes you back to the beginning, because bb ends in neither a nor ab.

L(D1) = L(D1R)

(a + b)* a b

The transition diagram

The drawing has four conventions and no more.

A circle for each state, labelled with its name.

An arrow from nowhere into the initial state, so it can be told from the others.

A double circle round each final state.

An arrow from q to r labelled a whenever delta(q, a) is r. Where two symbols lead from the same state to the same state, one arrow carries both labels separated by a comma, so an arrow labelled "a, b" means both.

A loop is an arrow from a state to itself, and it is how a machine ignores a symbol.

The three conversions

Table to equations. Read each cell: the row's state, the column's symbol, the cell's state. Count them at the end: rows times symbol columns.

Equations to table. Draw a grid with one row per state and one column per symbol, and fill each cell from the matching equation. Any cell you cannot fill means an equation is missing, and the machine is not a DFA until it is supplied or a dead state is added.

Table to diagram, and back. One circle per row. One arrow per non empty cell, from the row's state to the cell's state, labelled with the column's symbol. Mark the row labelled start with an incoming arrow from nowhere and every row labelled final with a double circle. Going back, read each arrow as a cell, and check at the end that every state has exactly one outgoing arrow for each symbol.

munotes.in45

Transitions and Their Properties: the Function, the Table and the Diagram

The properties a transition function must have

MU's label is "Transitions and Its properties", and these are the properties.

Totality. delta is defined for every state and every symbol. In a drawing: every circle has one outgoing arrow for each symbol of the alphabet. A machine failing this is incompletely specified, and chapter 5 gives the dead state remedy.

Single valuedness. delta gives one state, never two. In a drawing: no state has two arrows with the same label leaving it. A machine failing this is nondeterministic, which is chapter 14.

Closure within Q. Every value delta returns is a state of the machine. A cell naming a state not in the table is a mistake, and the checker behind this book rejects it outright.

No dependence on anything else. delta depends on the current state and the current symbol, and on nothing else: not on how the machine got there, not on how many symbols have been read, not on what comes later. This is the property that makes the machine finite state, and it is the one a beginner violates when designing: a rule like "if we have seen three a already, then go to q2" is not a transition unless "have seen three a" is itself recorded in a state.

Determinism of the whole run. Because delta is total and single valued, a string determines exactly one sequence of states. That is what chapter 11 formalises.

Worked example: all three representations of one machine

The problem. Over {0, 1}, accept strings containing an even number of 0.

What to remember. Only the parity of the count of 0 so far, which is two possibilities. Two states: even and odd.

Equations.

delta(E, 0) = O

delta(E, 1) = E

delta(O, 0) = E

delta(O, 1) = O

Table.

D2State01
start finalEOE
OEO

Accepts: ε, 1, 00, 11, 001, 0011

Rejects: 0, 01, 10, 000, 0001

Diagram, in words, because the drawing is the same information. Two circles, E and O. E is double circled and has an incoming arrow from nowhere. An arrow from E to O labelled 0, and one from O to E labelled 0. A loop on E labelled 1, and a loop on O labelled 1.

Note that the start state is also final, which is correct: the empty string contains zero 0, and zero is even.

L(D2) = L(D2R)

L(D2) is infinite

D2 is deterministic

1* (0 1* 0 1*)*

That regular expression says: any number of 1, then any number of blocks, each block being a 0, some 1, another 0, some 1. Every block contributes exactly two 0, so the total is even. The equality is decided exactly by the checker rather than argued here.

munotes.in46

Transitions and Their Properties: the Function, the Table and the Diagram

Distinctions

Transition tableTransition diagram
Shows a missing moveas a blank cell, visiblynot at all, it just looks finished
Shows a cyclenot obviouslyat a glance
Good forimplementing, checkingdesigning, explaining
Sizestates times symbols cellsone arrow per transition

What it does NOT mean

A transition is not a step in a program. It has no condition beyond the current symbol and no memory beyond the current state.

A loop is not a delay. An arrow from a state to itself means the symbol is consumed and the state does not change. One symbol is still read.

Two arrows with the same label from one state is not shorthand. It is nondeterminism, and in a DFA it is an error.

An unreachable state is not an error. A machine may have states no input can reach. It is wasteful, and chapter 20 removes them, but the machine is still well defined and still a DFA.

Quick revision

  • delta maps Q times Sigma to Q. The number of equations is states times symbols, and a short count

means a case was left out.

  • Three representations: equations, table, diagram, and a student must convert between any two.
  • This book's table: the corner names the machine, the left column carries start and final, an empty

divider column draws the double rule, and the remaining columns are the symbols.

  • Diagram: circle per state, arrow from nowhere into the start, double circle for final, labelled

arrow per transition, loop for a symbol that changes nothing.

  • The properties: total, single valued, values inside Q, dependent on nothing but the current state and

symbol, and hence determining exactly one run per string.

  • A missing cell means not a DFA; two entries in a cell means nondeterministic.

Test yourself

1. A DFA has 4 states over {a, b, c}. How many cells does its table have, and how many equations does delta need? Twelve of each: four states times three symbols. A table with eleven filled cells is not a DFA.

2. What does a loop labelled b on state q mean? That delta(q, b) is q: the symbol b is read and consumed, and the machine stays in q.

3. Which property fails if a state has two arrows labelled a leaving it? Which fails if it has none? Single valuedness fails in the first case, making the machine nondeterministic. Totality fails in the second, making the machine incompletely specified.

4. Convert these equations into a table: delta(p, 0) = p, delta(p, 1) = q, delta(q, 0) = q, delta(q, 1) = p, with p initial and q final. Two rows. The row for p is marked start and holds p under 0 and q under 1. The row for q is marked final and holds q under 0 and p under 1.

munotes.in47

Transitions and Their Properties: the Function, the Table and the Diagram

5. Why is a table preferred to a diagram for checking a machine? Because an omission shows up as a blank cell, whereas a diagram with a missing arrow looks complete, and because a table can be read and executed by a program while a drawing cannot.

6. Explain why "go to q2 if three a have been seen" is not a legal transition. Because a transition may depend only on the current state and the current symbol. The number of a seen so far is history, and it can only influence the machine if it has already been recorded by being in a particular state.

Contents This chapter on its own page

munotes.in48

Chapter Eleven

The Extended Transition Function: What a Machine Does to a Whole String

Syllabus topic Module 1, "Automata Theory: Transitions and Its properties"

In one line

The extended transition function says where a machine ends up after reading a whole string, and it is built from the one symbol function by induction.

In the wording a student can write in an examination: the extended transition function delta hat maps Q times Sigma star to Q and is defined recursively by delta hat(q, epsilon) = q and delta hat(q, wa) = delta(delta hat(q, w), a) for every state q, string w and symbol a. It is the unique extension of delta from single symbols to strings.

Why it is needed at all

Because delta takes a symbol and a string is not a symbol.

The definition of a DFA in chapter 9 says delta maps Q times Sigma to Q, so delta(q0, 0101) is not a legal expression: 0101 is a string, not a symbol, and the function is not defined on it. Yet every statement worth making about a machine is about a string. "The machine accepts 0101" has to mean something, and what it means is that some function of the whole string lands in a final state.

So a second function is defined, from the first, and it is this second function the rest of the book uses.

The definition, in two lines

The extended function is written delta hat, and it maps Q times Sigma star to Q. It is defined by saying what it does on the empty string, and then how one more symbol changes the answer.

delta hat(q, ε) = q

delta hat(q, w a) = delta( delta hat(q, w), a )

The first line says that reading nothing changes nothing. This is the base case.

The second line says: to read a string that is some string w followed by one symbol a, first read w, which leaves you in some state, and then take one ordinary step from there on a. This is the inductive step.

That is a definition by induction on the length of the string, exactly in the shape chapter 6 gave: a base case for the shortest string, and a rule that gets from every string to the strings one symbol longer. Because every string is either empty or is some shorter string with one symbol added, the two lines together define delta hat for every string, and define it once.

Reading the definition in the other direction

The recursion above peels the symbol off the right, and that is what makes it easy to prove things with. It is not how anybody actually runs a machine, and that is worth saying plainly so the two are never confused.

A machine is run by peeling symbols off the left: start in q0, read the first symbol, move, read the second, move, and so on. The following identity says the two agree, and it is itself proved by induction:

munotes.in49

The Extended Transition Function: What a Machine Does to a Whole String

delta hat(q, a w) = delta hat( delta(q, a), w )

So you may compute delta hat either way and get the same answer. Proofs use the first form; runs use the second.

Worked example

Take D3, which accepts strings over {a, b} in which every a is immediately followed by a b.

D3Stateab
start finalq0q1q0
q1q2q0
q2q2q2

Accepts: ε, b, ab, bab, abab, bb

Rejects: a, aa, aab, aba, baa

State q2 is the dead state: once two a have been seen in a row, no continuation can repair the string, so the machine sits in q2 for ever and q2 is not final. State q1 means "an a has just been read and is still waiting for its b".

Now compute delta hat(q0, abb) by the definition, peeling from the right.

delta hat(q0, abb) = delta( delta hat(q0, ab), b )

= delta( delta( delta hat(q0, a), b ), b )

= delta( delta( delta( delta hat(q0, ε), a ), b ), b )

= delta( delta( delta( q0, a ), b ), b )

= delta( delta( q1, b ), b )

= delta( q0, b )

= q0

Six lines, and the innermost one is the base case. The answer q0 is final, so abb is accepted, which matches the list above.

The same computation peeling from the left is shorter to write and is what a run looks like:

q0 abb
q1 bb
q0 b
q0 ε

The checker re-executes that trace against the table and fails if any single step disagrees.

Why the definition is worth the trouble

Three later results are statements about delta hat, and none of them could be stated without it.

Acceptance. Chapter 12 defines L(M) as the set of w with delta hat(q0, w) in F. Without delta hat there is no way to say it.

Correctness proofs. To prove a machine accepts what you claim, you prove a statement of the form "delta hat(q0, w) is q1 exactly when w has such and such a property", and you prove it by induction on the length of w. The base case is the empty string and the step adds one symbol, which is exactly the shape of the definition above. Chapter 12 works one of these in full.

The subset construction. Chapter 16's whole proof is one identity between the extended transition function of the nondeterministic machine and that of the deterministic one it builds.

munotes.in50

The Extended Transition Function: What a Machine Does to a Whole String

The one property that makes the machine finite state

There is a single fact about delta hat that all of Module 1's negative results rest on, and it is easiest to state here.

delta hat(q, x y) = delta hat( delta hat(q, x), y )

In words: to read x then y, read x, and then read y from wherever that left you. Nothing else about x survives. The state after x is the complete summary of x as far as the rest of the run is concerned.

That is why a machine "cannot count": if two different strings x and z lead to the same state, then every continuation treats them identically, so if xy is accepted then zy must be accepted too. Chapter 21 turns this into the Myhill Nerode theorem, and chapter 36 turns it into the pumping lemma. Both are this one identity with the pigeonhole principle applied.

Distinctions

deltadelta hat
Second argumentone symbola whole string
DomainQ times SigmaQ times Sigma star
Given bythe tablethe definition, from the table
Used forone moveacceptance, and every proof
On the empty stringnot definedreturns the state unchanged

What it does NOT mean

delta hat is not a new machine. It is a function computed from delta, and it adds no power. Two machines with the same delta have the same delta hat.

delta hat(q, epsilon) is not undefined. It is q. Forgetting this is what makes students get the empty string wrong.

The right hand peeling is not the order of a run. It is the order that makes the induction work. The two orders agree, and the identity above says so.

Many books write delta for both. Most textbooks drop the hat once the definition has been given and rely on the reader to see from the argument which is meant. This book keeps them apart wherever a proof depends on it.

Quick revision

  • delta takes one symbol; delta hat takes a string, and maps Q times Sigma star to Q.
  • delta hat(q, epsilon) = q, and delta hat(q, wa) = delta(delta hat(q, w), a). Base case plus one

symbol, which is induction on length.

  • Peeling from the right is for proofs; peeling from the left is a run; the two agree.
  • Acceptance, correctness proofs and the subset construction are all statements about delta hat.
  • The key identity is delta hat(q, xy) = delta hat(delta hat(q, x), y): the state after x is the whole

of what the machine remembers about x.

  • That identity is the source of both the Myhill Nerode theorem and the pumping lemma.

Test yourself

1. Write the two lines that define delta hat. delta hat(q, epsilon) = q, and delta hat(q, wa) = delta(delta hat(q, w), a) for any string w and symbol a.

munotes.in51

The Extended Transition Function: What a Machine Does to a Whole String

2. Why can delta not be applied to a string directly? Because delta is declared as a function on Q times Sigma, so its second argument must be a single symbol. A string of length 2 is not in its domain.

3. For D3 above, compute delta hat(q0, ba) and say whether ba is accepted. delta hat(q0, b) is q0; then delta(q0, a) is q1. So delta hat(q0, ba) is q1, which is not final, so ba is rejected. That is right: the a at the end has no b after it.

4. State the identity that says the state after a prefix summarises that prefix, and say what it implies. delta hat(q, xy) = delta hat(delta hat(q, x), y). It implies that if two strings lead to the same state, then no continuation can distinguish them, which is the source of the pumping lemma and the Myhill Nerode theorem.

5. In a correctness proof by induction on string length, what is the base case? The empty string, where delta hat(q0, epsilon) is q0, so the claim must be checked for the start state and the empty string outright.

6. Does introducing delta hat make the machine more powerful? No. It is defined entirely from delta and the table, and it introduces no new behaviour; it only lets the behaviour already there be spoken about.

Contents This chapter on its own page

munotes.in52

Chapter Twelve

Acceptability: When a Machine Accepts a String, and What Language It Accepts

Syllabus topic Module 1, "Automata Theory: Acceptability by Finite Automaton"

In one line

A machine accepts a string when the state it reaches after reading the whole string is a final state, and the language of the machine is every string it accepts.

In the wording a student can write in an examination: a string w in Sigma star is said to be accepted by the finite automaton M = (Q, Sigma, delta, q0, F) if delta hat(q0, w) is in F. The set T(M), also written L(M), of all strings accepted by M is called the language accepted by M, and a language is regular if it is accepted by some finite automaton.

Why the definition is worded this way

Every word in it is doing something.

delta hat and not delta, because w is a string. Chapter 11 built the function that takes strings.

From q0, because that is where the machine starts. Acceptance is not about paths from anywhere; it is about the one path from the initial state.

In F, because acceptance is decided by the state at the end. Not the states along the way.

The whole string, because a machine reads to the end. A machine that reached a final state after three symbols of a five symbol string has not accepted anything; it must read the other two.

Three things a student confuses

These are worth separating before any work is done, because most errors in this block come from sliding between them.

A string being accepted is a question about one string and has a yes or no answer, found by one run.

A language being accepted is a question about an infinite set, and is answered by the machine's definition rather than by running it. L(M) exists as soon as M does.

A machine being correct is a question about whether L(M) is the language you were asked for. This is the question that needs a proof, and it is the question a table of examples cannot settle.

Finding the acceptability of a string

This is what the papers ask for, and the method is mechanical.

  1. Start in q0.
  2. Read the leftmost unread symbol and look it up in the row for the current state.
  3. Move to the state in that cell.
  4. Repeat until no symbols are left.
  5. Accept if that state is final, otherwise reject.

Set the work out as a run so each step can be marked. Take E1:

E1State01
startq0q1q0
q1q1q2
finalq2q1q0

Accepts: 01, 001, 1101, 1001, 00101

Rejects: ε, 0, 1, 10, 11, 010

E1 accepts a string exactly when it ends in 01. State q1 means "a 0 has just been read", and q2, which is final, means "01 has just been read".

munotes.in53

Acceptability: When a Machine Accepts a String, and What Language It Accepts

Now the three strings, each written as a run.

q0 1001
q0 001
q1 01
q1 1
q2 ε

Ends in q2, which is final, so 1001 is accepted.

q0 010
q1 10
q2 0
q1 ε

Ends in q1, which is not final, so 010 is rejected.

q0 11
q0 1
q0 ε

Ends in q0, which is not final, so 11 is rejected. Note that in the first run the machine passed through the final state q2 and then left it and came back; and in the second it passed through q2 in the middle and ended elsewhere. Only the last state counts.

Every one of those three runs is re-executed by the checker against the table, step by step, so a wrong step in any of them would stop this book being built.

The language accepted, written down

L(E1) = L(E1R)

(0 + 1)* 0 1

and the checker decides that equality exactly, over every string there is, by comparing the two machines rather than by testing examples.

Proving a machine correct, by induction

This is the part the papers do not ask for and the part that makes the difference between a student who can check a machine and one who can trust a machine they built.

Claim. For E1, and for every string w over {0, 1}:

  1. delta hat(q0, w) is q2 exactly when w ends in 01;
  2. delta hat(q0, w) is q1 exactly when w ends in 0;
  3. delta hat(q0, w) is q0 otherwise, that is, when w is empty or ends in 1 but not 01.

Proof, by induction on the length of w.

Base case. w is the empty string. delta hat(q0, epsilon) is q0 by the first line of chapter 11's definition. The empty string does not end in 01 and does not end in 0, so case 3 applies and says the state should be q0. It is.

Inductive step. Assume the claim holds for w, and consider w followed by one symbol a. There are six combinations of the state after w and the symbol a, and the table gives each one.

State after waNew statew a ends inClaim says
q00q10q1
q01q01, not 01q0
q10q10q1
q11q201q2
q20q10q1
q21q01, not 01q0

Take the fourth row as the one that carries the argument. The state after w is q1, so by the inductive hypothesis w ends in 0. Then a is 1, so wa ends in 01, and the claim says the new state should be q2. The table says delta(q1, 1) is q2. They agree.

munotes.in54

Acceptability: When a Machine Accepts a String, and What Language It Accepts

The second row needs one extra word. The state after w is q0, so by hypothesis w is empty or ends in 1 without ending in 01. Adding a 1 gives a string ending in 1, and its last two symbols are either just the single 1, or 11, neither of which is 01. So case 3 applies and the state should be q0, which is what the table gives.

All six rows agree, so the claim holds for wa. By induction it holds for every string.

Therefore delta hat(q0, w) is in F, that is, is q2, exactly when w ends in 01. So L(E1) is the set of strings ending in 01, which is what was asked for.

That is what a correctness proof looks like, and its length is six table rows and two paragraphs. A student who can write this can be sure of a machine they designed; a student who can only test it cannot.

Acceptance by the empty string, and the two smallest languages

Two machines worth having in mind, because chapter 39 uses both and because they are the boundary cases of every construction.

A machine accepting nothing at all. One state, marked start and not final, with a loop on every symbol. It accepts the empty language { }.

A machine accepting only the empty string. Two states: the start state, which is final, and a dead state which is not. Every symbol from the start state goes to the dead state, and the dead state loops to itself.

E2Stateab
start finalsdd
ddd

Accepts: ε

Rejects: a, b, aa, ab, ba, bb, aba

L(E2) is finite

That machine accepts {epsilon}, the language with one string in it. The machine that accepts { } is the same table with the final marking removed, and the two differ on exactly one input, which is the distinction chapter 3 warned about.

Distinctions

Accepting a stringAccepting a language
Objectone stringa set of strings, usually infinite
Decided byone run of the machinethe machine's definition
Answered bytracinga proof, or an exact comparison
A wrong answer meansyou traced badlythe machine is wrong
Passing through a final stateEnding in a final state
Meaningsome prefix is acceptedthe whole string is accepted
Decides acceptancenoyes

What it does NOT mean

A machine does not stop when it accepts. There is no halting in a DFA. It reads to the end of the input and the answer is read off the last state.

munotes.in55

Acceptability: When a Machine Accepts a String, and What Language It Accepts

Visiting a final state proves nothing about the whole string. In the 1001 run above the machine was in q2 only at the end; in the 010 run it was in q2 in the middle and rejected. Both are normal.

A table of accepted strings is not a proof of correctness. It proves the machine handles those strings. The infinitely many others need an argument.

Regular is a property of the language. A language is regular when some DFA accepts it. Every DFA's language is regular by definition, so "is this machine regular" is not a question.

Quick revision

  • M accepts w when delta hat(q0, w) is in F: the state after reading the whole string, from the start

state, must be final.

  • L(M), also written T(M), is the set of all strings M accepts. A language is regular when some DFA

accepts it.

  • To find acceptability: start at q0, follow one cell per symbol, and look at the last state.
  • Only the last state matters. Passing through a final state on the way means nothing.
  • Correctness is proved by induction on string length: state what each state means, check the empty

string, then check every combination of state and symbol against the table.

  • The empty string is accepted exactly when q0 is final. { } and {epsilon} are accepted by machines

differing only in that marking.

Test yourself

1. State the condition for M to accept w, using the right function. delta hat(q0, w) is in F. delta alone will not do, because w is a string.

2. E1 above: is 0101 accepted? Trace it. q0 on 0 gives q1; q1 on 1 gives q2; q2 on 0 gives q1; q1 on 1 gives q2. The last state is q2, which is final, so 0101 is accepted, and indeed it ends in 01.

3. A machine passes through a final state after two of five symbols and ends in a non final state. Is the string accepted? No. Acceptance depends only on the state after the whole string has been read.

4. Describe the machine accepting the empty language, and the one accepting only the empty string. For the empty language: one state, start, not final, looping on every symbol. For {epsilon}: the same machine with the start state marked final, plus a non final dead state that every symbol from the start state leads to and that loops to itself.

5. Why is a proof by induction needed when the machine has been tested on twenty strings? Because the language is infinite and twenty strings say nothing about the rest. The induction covers every string by proving what each state means for every prefix.

munotes.in56

Acceptability: When a Machine Accepts a String, and What Language It Accepts

6. In a correctness proof, how many combinations must the inductive step check? One for each state and each symbol, so the number of states multiplied by the size of the alphabet, which is exactly the number of cells in the transition table.

Contents This chapter on its own page

munotes.in57

Chapter Thirteen

Designing a DFA: the Method, and Eight Machines Built With It

Syllabus topic Module 1, "Automata Theory: Finite Automaton"

In one line

To design a DFA, decide what the machine must remember, give each thing it must remember a state, then fill in every cell of the table.

In the wording a student can write in an examination: the design of a finite automaton proceeds by identifying the finitely many distinguishable conditions in which the machine can find itself after reading a prefix of the input, assigning one state to each, determining the transition for every state and every input symbol, and marking as final those states in which the prefix read so far is itself a string of the language.

Why a method is needed

Because guessing works on the first two exercises and then stops working.

The machines in the examination are small, but they are not obvious, and a student who draws circles and hopes will produce a machine that handles the strings they thought of and fails on one they did not. The method below is not a shortcut; it is the thing that makes the machine right, and its last step, filling in every cell, is the step that catches the omissions.

The method, in five steps

Step 1. Ask what the machine must remember. Not what the string looks like: what has to be carried forward. This is the whole of the design and the other four steps are bookkeeping. The question to ask is: if I have read some prefix, what is the least I need to know about it to handle whatever comes next?

Step 2. Give each answer a state, and write down what it means. Write the meaning next to the state name. It is not decoration: it is what makes step 3 mechanical and what you will use in a correctness proof.

Step 3. Fill every cell. For each state and each symbol, ask: if I am in this condition and I read this symbol, what condition am I in now? Read the answer off the meanings, not off a drawing. The number of cells is the number of states times the size of the alphabet, and a short count means a case was missed.

Step 4. Mark the final states. A state is final when the prefix that put you there is itself in the language. Check the empty string separately, because it is the prefix that puts you in the start state and it is the one everybody forgets.

Step 5. Test the boundaries, then prove it. Run the empty string, the shortest accepted string, the shortest rejected string, and a string that nearly qualifies. Then, if it matters, prove the claim about what each state means by induction, as chapter 12 did.

munotes.in58

Designing a DFA: the Method, and Eight Machines Built With It

The three patterns that cover most exercises

Before the machines, three observations that turn most questions into one of three shapes.

Counting modulo something needs that many states. "An even number of a", "a length divisible by 3", "a number of 1 that leaves remainder 2 on division by 4" all need one state per remainder, and no more. The reason is that the remainder is all you need: adding one symbol changes the remainder in a way that depends only on the old remainder.

Remembering the last few symbols needs one state per possibility. "Ends in ab", "contains aba", "does not contain bb" need a state for each relevant recent history, and the trick is to keep the history as short as possible: for "ends in ab" the states are nothing useful, just seen a, just seen ab.

An irreparable condition needs a dead state. "Every a is followed by a b" can be broken by a prefix, and once broken nothing repairs it. That needs one non final state that loops to itself on everything.

Machine 1: strings over {a, b} ending in ab

MU's own question. What must be remembered: how much of the pattern ab is currently complete at the right hand end. Three conditions, so three states.

G1Stateab
starts0s1s0
s1s1s2
finals2s1s0

s0 means the string so far ends in neither a nor ab. s1 means it ends in a. s2 means it ends in ab.

The two cells students get wrong are in the bottom row. From s2, reading a takes you to s1, not to s0, because that a is itself the beginning of a possible new ab. And from s1, reading a keeps you in s1, because aa still ends in a.

Accepts: ab, aab, bab, abab, bbab

Rejects: ε, a, b, ba, aba, abb

L(G1) = L(G1R)

(a + b)* a b

Machine 2: strings over {a, b} containing aba

What must be remembered: how much of aba is complete at the right hand end, and once the whole pattern has been seen, that it has been seen, which is permanent.

G2Stateab
startp0p1p0
p1p1p2
p2p3p0
finalp3p3p3

p0: no useful suffix. p1: ends in a. p2: ends in ab. p3: aba has been seen somewhere, and nothing can undo that, so p3 loops to itself on both symbols.

The subtle cell is p2 on b: abb ends in neither a nor ab, so it goes back to p0. And p2 on a completes aba and reaches p3.

Accepts: aba, aabab, babaa, ababa, bbabab

munotes.in59

Designing a DFA: the Method, and Eight Machines Built With It

Rejects: ε, ab, abb, aab, bbb, babb

L(G2) = L(G2R)

(a + b)* a b a (a + b)*

Machine 3: an even number of a and an even number of b

What must be remembered: the parity of each count. Two independent facts each with two values, so four states, and the state set is the Cartesian product of chapter 2.

G3Stateab
start finaleeoeeo
oeeeoo
eoooee
ooeooe

The name records the two parities: ee means both even, oe means the a count is odd and the b count even, and so on. Reading an a flips the first letter; reading a b flips the second. The start state is final because zero is even.

Accepts: ε, aa, bb, abab, aabb, baab

Rejects: a, b, ab, ba, aab, abb

L(G3) = L(G3R)

(a a + b b + (a b + b a)(a a + b b)*(a b + b a))*

Machine 4: strings over {0, 1} whose value as a binary number is divisible by 3

What must be remembered: the remainder of the number read so far on division by 3, which is three conditions. This is the machine worth understanding, because it shows that "remember a remainder" is enough for a task that looks arithmetical.

Appending a bit b to a binary number n gives 2n plus b. So if n leaves remainder r, then 2n plus b leaves the remainder of 2r plus b. That rule fills the table with no thought at all.

G4State01
start finalr0r0r1
r1r2r0
r2r1r2

From r0 reading 0 gives remainder 0, and reading 1 gives 1. From r1 reading 0 gives 2 times 1 plus 0 which is 2, and reading 1 gives 3, whose remainder is 0. From r2 reading 0 gives 4, whose remainder is 1, and reading 1 gives 5, whose remainder is 2.

Accepts: ε, 0, 11, 110, 1001, 1100

Rejects: 1, 10, 100, 101, 111

110 is 6, which is divisible by 3. 1001 is 9. 1100 is 12. And 111 is 7, which is not. The empty string is accepted because the machine starts in r0; a book may prefer to exclude it, and doing so needs one extra state, which is left to the exercises.

Machine 5: strings over {a, b} in which every a is immediately followed by a b

What must be remembered: whether an a is currently waiting for its b, and whether the condition has already been broken beyond repair. Three states, the third of them dead.

munotes.in60

Designing a DFA: the Method, and Eight Machines Built With It

G5Stateab
start finalt0t1t0
t1t2t0
t2t2t2

t0: nothing outstanding, and the string so far is legal. t1: an a has just been read and needs a b next. t2: dead, because either two a came in a row or the string ended just after an a.

Note that t1 is not final. That is the whole content of "immediately followed": a string ending in a is not in the language, because the a never got its b.

Accepts: ε, b, ab, bb, abb, abab

Rejects: a, aa, ba, aab, aba

L(G5) = L(G5R)

(b + a b)*

Machine 6: strings over {a, b} of length exactly 3

What must be remembered: how many symbols have been read, up to 4, where 4 means "too many". Five states, and this is the shape every "exactly n" question takes.

G6Stateab
startn0n1n1
n1n2n2
n2n3n3
finaln3n4n4
n4n4n4

Each state counts, and n4 is the dead state for anything longer. Both symbols behave identically, because only the length matters.

Accepts: aaa, aab, abb, bbb, bab

Rejects: ε, a, ab, aaaa, bbbb

L(G6) = L(G6R)

L(G6) is finite

(a + b)(a + b)(a + b)

Machine 7: strings over {a, b} whose third symbol from the left is a

What must be remembered: how many symbols have been read while it still matters, and then the answer for ever. Five states.

G7Stateab
startm0m1m1
m1m2m2
m2yz
finalyyy
zzz

The first two symbols are read and ignored. The third decides: a reaches y, which is final and absorbing, and b reaches z, which is not final and absorbing. Strings shorter than 3 have no third symbol and are rejected, which is why m0, m1 and m2 are not final.

Accepts: aaa, bba, aba, bbab, abaaa

Rejects: ε, a, ab, aab, bbb, abb

L(G7) = L(G7R)

(a + b)(a + b) a (a + b)*

Machine 8: strings over {a, b} that do not contain bb

What must be remembered: whether the previous symbol was a b, and whether bb has already occurred, which is fatal. Three states.

G8Stateab
start finalu0u0u1
finalu1u0u2
u2u2u2

u0: legal so far and the last symbol was not b. u1: legal so far and the last symbol was b. u2: dead. Both u0 and u1 are final, because a legal string may end in b as long as it does not end in bb.

munotes.in61

Designing a DFA: the Method, and Eight Machines Built With It

Accepts: ε, a, b, ab, ba, aba, abab

Rejects: bb, abb, bba, abba, babb

L(G8) = L(G8R)

(a + b a)* (b + ε)

What every one of those eight has in common

The state set is a list of conditions, and the names say which. No machine above has a state whose meaning could not be written in a short phrase, and that is not a coincidence: a state with no meaning is a state you do not need or a design that has not been thought through.

Every cell is filled. Count them: three states times two symbols is six cells in machine 1, four times two is eight in machine 3, five times two is ten in machines 6 and 7.

The empty string is decided on purpose. In machines 3, 4, 5, 6 and 8 the answer for the empty string was a deliberate decision about whether the start state is final, and in each case it followed from the description rather than from habit.

Each one was checked, not trusted. The accepted and rejected strings under each machine are run, and each machine is compared with a regular expression for the same language over every string there is. That is the difference between a machine that looks right and one that is right.

Distinctions

RequirementStates neededShape
count modulo kka cycle
ends in a fixed string of length nn plus 1a chain with back edges
contains a fixed stringn plus 1a chain with an absorbing final state
length exactly nn plus 2a chain with a dead state
two independent conditionsproduct of the twoa grid
a condition that can be broken for everone extra, deada sink

What it does NOT mean

More states is not safer. A machine with spare states is harder to check and chapter 20 will remove them anyway. The discipline of asking "what is the least I must remember" is what produces a correct machine, not a generous state count.

A dead state is not optional in a DFA. If your drawing has a missing arrow, the table has a blank cell, and the machine is not a DFA until the dead state is added.

The start state being final is not a special case. It is the ordinary question of whether the empty string is in the language, answered like any other.

Testing is not designing. The eight machines above were designed from the meanings of their states and then tested. Designing by testing produces a machine that passes the tests you thought of.

munotes.in62

Designing a DFA: the Method, and Eight Machines Built With It

Quick revision

  • Five steps: decide what must be remembered; one state per condition, with its meaning written down;

fill every cell; mark the final states, checking the empty string on purpose; test the boundaries and then prove it.

  • Cells equals states times symbols. A short count means a case was missed.
  • Counting modulo k needs k states. A pattern of length n at the end needs n plus 1. A condition that

can be broken permanently needs a dead state.

  • Two independent conditions need the product of their state counts, and the state names should record

both.

  • A state whose meaning cannot be stated in a phrase is a sign the design is not finished.

Test yourself

1. Design a DFA over {a, b} accepting strings that begin with a and end with b. Four states: the start; a state meaning the first symbol was a and the last was not b; a final state meaning the first was a and the last was b; and a dead state for a string beginning with b. From the start, a goes to the second state and b to the dead state. In the second state, a stays and b goes to the final state. In the final state, a returns to the second and b stays.

2. How many states does a DFA need to accept strings over {0, 1} whose length is divisible by 4? Four, one per remainder, arranged in a cycle, with the remainder zero state both start and final.

3. In machine 1, why does s2 on a go to s1 rather than s0? Because the a just read may be the a of a later ab, so the machine must record that the string now ends in a. Going to s0 would lose that and would reject aba b, which ends in ab.

4. Why are both u0 and u1 final in machine 8? Because a string with no bb may end in a or in a single b. Only two consecutive b are fatal, so ending in one b is legal.

5. A student's machine for "contains aba" sends the state meaning "ends in ab" back to the start on reading b. Is that right? Yes. After ab, reading b gives a string ending in abb, which has no useful suffix towards aba, so the machine must return to the state meaning no progress.

6. What is the first question to ask when designing any DFA? What is the least I need to know about the prefix read so far in order to deal correctly with anything that might come next.

Contents This chapter on its own page

munotes.in63

Chapter Fourteen

Nondeterminism, and the Nondeterministic Finite Automaton

Syllabus topic Module 1, "Automata Theory: Nondeterministic Finite State Machines"

In one line

A nondeterministic finite automaton may have several next states, or none, for a given state and symbol, and it accepts a string if at least one of its possible runs ends in a final state.

In the wording a student can write in an examination: a nondeterministic finite automaton, or NFA, is a quintuple M = (Q, Sigma, delta, q0, F) in which delta is a function from Q times Sigma to the power set of Q, so that delta(q, a) is a set of states, possibly empty. A string w is accepted if delta hat(q0, w) contains at least one final state.

What changes, and it is one line of the definition

Compare the two definitions side by side, because exactly one thing differs.

DFANFA
delta mapsQ times Sigma to QQ times Sigma to the power set of Q
delta(q, a) isone statea set of states
Runs on a stringexactly onepossibly many, possibly none
Accepts w whenthe single run ends finallysome run ends finally

Everything else is unchanged: the states are finite, the alphabet is finite, there is one initial state, and F is a subset of Q.

The codomain being the power set is what makes the empty set a legal answer. delta(q, a) equal to the empty set means there is no move at all, and a run that reaches that point simply dies. It does not reject on behalf of the machine; it is one run that failed, and other runs may still succeed.

Who chooses?

Nobody. This is the point at which most students acquire a wrong picture, so it is worth being blunt about what nondeterminism is not.

It is not guessing. Nothing inside the machine makes a decision. There is no random choice and no oracle.

It is not parallelism. The machine is not running several copies of itself, and the definition says nothing about doing several things at once.

It is not a physical machine at all, any more than a DFA is.

What it is, is a definition of acceptance. The machine's behaviour on a string is the set of all runs the transition relation permits, and the string is in the language when at least one of those runs ends in a final state. That is a statement about a set of paths in a graph, and it needs no agent to carry it out.

If a picture helps, the honest one is the one the mathematics uses: the machine is in a set of states at once. After reading a prefix, the machine is in every state that some run could have reached. That is not a metaphor; it is what the extended transition function computes, and chapter 16 turns exactly that observation into a deterministic machine.

munotes.in64

Nondeterminism, and the Nondeterministic Finite Automaton

The extended transition function for an NFA

The two line definition of chapter 11 needs one change, because the answer is now a set.

delta hat(q, ε) = {q}

delta hat(q, w a) = the union of delta(p, a) over every p in delta hat(q, w)

In words: to read one more symbol, take the set of states you might be in, and collect together everything reachable from any of them on that symbol.

And acceptance:

w is accepted when delta hat(q0, w) contains at least one state of F

If the set ever becomes empty it stays empty, because the union over no states is empty. So a machine whose every run has died rejects everything after that point, which is correct.

Why nondeterminism is worth having

Two reasons, and the second is the one that matters in this paper.

It is easier to design with. Some languages have an obvious nondeterministic machine and an awkward deterministic one. "Contains aba" is the standard example: nondeterministically the machine loops at the start doing nothing, then when it feels like it reads aba and loops at the end. Deterministically it has to track how much of aba is currently complete, which is what machine 2 of chapter 13 does, and that takes real thought.

It is what the constructions produce. Chapter 34 turns a regular expression into a machine, and the machine it produces is nondeterministic, with empty moves, because that is the only way to glue two machines together without knowing anything about them. If nondeterminism were not allowed, that construction would not exist. The same is true of the constructions in chapters 33 and 38.

Machine 1: contains aba, nondeterministically

Three states, and compare it with the four state deterministic machine of chapter 13.

N1Stateab
startn0{n0, n1}n0
n1-n2
n2n3-
finaln3n3n3

Read it as a description rather than as a program. In n0 the machine either stays where it is, ignoring the symbol, or begins the pattern. In n1 it needs a b. In n2 it needs an a, which finishes the pattern and reaches n3, from which nothing can go wrong.

The cell for n0 on a holds two states, which is what makes this an NFA. The cells for n1 on a and n2 on b hold a dash, meaning no move: a run that is looking for a b and gets an a is simply over.

Accepts: aba, aabab, babaa, ababa, bbabab

Rejects: ε, ab, abb, aab, bbb, babb

munotes.in65

Nondeterminism, and the Nondeterministic Finite Automaton

N1 is nondeterministic

L(N1) = L(N1R)

(a + b)* a b a (a + b)*

And it accepts the same language as the deterministic machine of chapter 13, which is what chapter 16 will prove must always be possible.

Machine 2: the exponential case

Here is the family that shows the subset construction of chapter 16 really can be as expensive as its bound.

Consider the language of strings over {a, b} whose fourth symbol from the right hand end is an a. Nondeterministically this is easy: loop at the start, then read an a, then read three more symbols of any kind, and stop.

N2Stateab
startk0{k0, k1}k0
k1k2k2
k2k3k3
k3k4k4
finalk4--

Five states, and k4 has no moves at all, which is correct: if anything follows, that run dies, because then the a was not the fourth from the end.

Accepts: abbb, aaaa, babbb, babab, aabbb

Rejects: ε, a, ab, abb, bbbb, bbbba

N2 is nondeterministic

L(N2) = L(N2R)

(a + b)* a (a + b)(a + b)(a + b)

A deterministic machine for this language has to remember the last four symbols, because until the input ends it cannot know which symbol will turn out to be fourth from the end. Four symbols with two possibilities each is sixteen conditions, so the minimal DFA has sixteen states. The NFA has five, and the general pattern is n plus 1 against 2 to the n. Chapter 16's exponential bound is therefore not pessimism: it is achieved.

Machine 3: an incomplete machine is an NFA

This is worth stating, because students meet it without noticing.

N3Stateab
startc0c1-
finalc1-c1

Every cell holds at most one state, so nothing is being chosen anywhere, and yet this is not a DFA, because two cells are empty and a DFA's delta must be total. It is an NFA in which every set has size zero or one.

Such a machine is sometimes called incompletely specified, and turning it into a DFA needs no subset construction at all: just add the dead state of chapter 5.

Accepts: a, ab, abb, abbb

Rejects: ε, b, ba, aa, aab

L(N3) = L(N3R)

a b*

Worked example: running an NFA by hand

Take N1 and the string abab. Keep track of the set of states, which is the only way to do this without backtracking.

ReadSet of statesHow
nothing{n0}the start
a{n0, n1}n0 on a gives both
ab{n0, n2}n0 on b gives n0; n1 on b gives n2
aba{n0, n1, n3}n0 on a gives n0 and n1; n2 on a gives n3
abab{n0, n2, n3}n0 on b gives n0; n1 on b gives n2; n3 on b gives n3
munotes.in66

Nondeterminism, and the Nondeterministic Finite Automaton

The final set contains n3, which is final, so abab is accepted. And it should be: abab contains aba as its first three symbols.

That table is the whole technique, and it is worth noticing what has just happened. The sets of states in the left hand column behave exactly like the states of a deterministic machine: each row is determined by the row above and the symbol read. That observation, and nothing more, is the subset construction of chapter 16.

Distinctions

DFANFAIncompletely specified
Cells holdone stateany setat most one state
Blank cellsforbiddenallowedpresent, and the reason
Runs per stringonemany or noneat most one
Accepts whenthe run ends finallysome run ends finallythe run exists and ends finally
Fix to make a DFAnone neededsubset constructionadd a dead state

What it does NOT mean

An NFA is not more powerful than a DFA. Chapter 16 proves they accept exactly the same class of languages. Nondeterminism buys convenience and smaller machines, never new languages. This is the result students find hardest to believe, and it is worth holding on to, because the corresponding statement is FALSE for pushdown automata in chapter 57, and false as far as anybody knows for the linear bounded automaton in chapter 61.

A dash in a cell is not a rejection. It kills one run. Other runs continue, and the string is still accepted if one of them succeeds.

An NFA does not backtrack. Backtracking is one way to implement a search for an accepting run. The definition mentions nothing of the kind, and the set of states method above never backtracks.

Nondeterminism is not probability. There are no likelihoods anywhere. A run either exists or does not.

Quick revision

  • An NFA is (Q, Sigma, delta, q0, F) with delta mapping Q times Sigma to the power set of Q, so a cell

may hold several states or none.

  • Accepts w when delta hat(q0, w) contains at least one final state: some run succeeds.
  • delta hat(q, epsilon) is {q}, and one more symbol takes the union of the moves from every state in the

current set.

  • Nondeterminism is a definition of acceptance, not guessing, not parallelism, not probability. The

working picture is that the machine is in a set of states at once.

  • It is easier to design with, and the constructions of chapters 33, 34 and 38 produce it.
  • An incompletely specified machine is an NFA whose sets all have size at most one; add a dead state to
munotes.in67

Nondeterminism, and the Nondeterministic Finite Automaton

make it a DFA.

  • NFA and DFA accept exactly the same languages, but an NFA can be exponentially smaller, and the

fourth from the right example achieves the bound.

Test yourself

1. Give the one difference between the definitions of a DFA and an NFA. The codomain of delta. For a DFA it is Q, so there is exactly one next state. For an NFA it is the power set of Q, so there may be several or none.

2. What does an empty cell mean in an NFA, and what does it mean in a DFA? In an NFA it means that run has no continuation and dies, while other runs are unaffected. In a DFA it means the definition is incomplete, and the machine is not a DFA until a dead state is added.

3. Run N1 on the string bab and say whether it is accepted. Start {n0}. On b: {n0}. On a: {n0, n1}. On b: {n0, n2}. The final set has no final state in it, so bab is rejected, which is right because bab does not contain aba.

4. Why is nondeterminism not the same as guessing? Because nothing in the machine makes any choice. Acceptance is defined as the existence of a successful run among all the runs the transition relation permits, which is a property of a graph and needs no agent.

5. Name a language whose NFA is much smaller than its smallest DFA, and give both sizes. Strings over {a, b} whose fourth symbol from the right is an a. The NFA has 5 states; the minimal DFA has 16, because it must remember the last four symbols.

6. If an NFA's cells all hold exactly one state, what is it? A DFA, written with sets of size one. Determinism is the special case where every set is a singleton and none is empty.

Contents This chapter on its own page

munotes.in68

Chapter Fifteen

Empty Moves, and the Epsilon Closure

Syllabus topic Module 1, "Automata Theory: Nondeterministic Finite State Machines"

In one line

An empty move lets a machine change state without reading anything, and the epsilon closure of a state is every state reachable from it by empty moves alone.

In the wording a student can write in an examination: an NFA with epsilon moves is a quintuple in which delta maps Q times (Sigma union {epsilon}) to the power set of Q. The epsilon closure of a state q, written epsilon closure of q, is the set of all states reachable from q by a sequence of zero or more epsilon transitions, and it always contains q itself.

Why empty moves exist

Because the constructions need them, and for no other reason.

Chapter 34 builds a machine from a regular expression by gluing smaller machines together. To glue the machine for x in front of the machine for y, you need an arrow from x's final state to y's start state that consumes nothing, because the symbol that x finished on has already been used. Without an empty move there is no such arrow, and the construction would have to look inside the two machines and merge their tables, which is far harder and is different for every case.

The same is true of the union construction and the closure construction. Empty moves are the glue, and they are why the machines in chapter 34 look the way they do.

They also buy nothing in power. Chapter 17 shows how to remove every empty move from any machine without changing the language, so a machine with them accepts nothing new.

The definition

An NFA with epsilon moves, sometimes written epsilon-NFA, differs from the NFA of chapter 14 in one place: the alphabet of delta gains the symbol epsilon.

delta maps Q times (Sigma union {ε}) to the power set of Q

So a state may have a cell under the column headed epsilon, and the states in that cell can be moved to without reading a symbol from the input.

Note carefully that epsilon is not a member of Sigma. The input is still a string over Sigma, and epsilon never appears in an input. It is a label on an arrow, not a symbol the machine can read.

The epsilon closure

The closure of a state q is written as the epsilon closure of q and defined as:

ε closure of q = { p : p is reachable from q using ε moves only, in zero or more steps }

Zero or more is the phrase to remember, and it is the source of the one mistake everybody makes: the closure of q always contains q, because zero steps get you from q to q. Chapter 4's reflexive transitive closure is exactly this idea, and the reason the word closure is used.

munotes.in69

Empty Moves, and the Epsilon Closure

The closure of a set of states is the union of the closures of its members.

Computing it: the algorithm

This is asked as a procedure, so it is worth setting out as one.

  1. Start with the answer containing q alone.
  2. Pick any state p in the answer whose epsilon moves have not yet been followed.
  3. Add every state in the epsilon cell of p to the answer.
  4. Repeat from step 2 until nothing new is added.

It terminates because there are finitely many states and the answer only grows.

Worked example

Take P1, a machine with three empty moves.

P1Stateεab
startp{q, r}p-
q-s-
rt-r
s--s
finalt---

Now the closures, one state at a time, each by the algorithm.

State qε closure of qWhy
p{p, q, r, t}p itself; q and r directly; t from r
q{q}itself only, no ε moves
r{r, t}itself, and t directly
s{s}itself only
t{t}itself only

The closure of p is the one that needs the algorithm rather than one glance: t is two empty moves away, through r, and a student who only looks at the epsilon cell of p gets three states instead of four.

Every one of those five sets contains its own state, which is the check to run before writing anything down.

What the closure is for

Two uses, both immediate.

Running the machine. The set of states the machine could be in after reading a prefix must be closed under empty moves, because an empty move can always be taken. So a run keeps the closure at every step:

  1. Begin in the closure of the start state, not in the start state alone.
  2. To read a symbol a, take the current set, follow every a arrow, then take the closure of the result.
  3. Accept if the final set contains a final state.

Step 1 is the step that is forgotten, and it matters whenever the start state can reach a final state by empty moves, which is exactly the case where the machine accepts the empty string.

Removing the moves and building a DFA. Chapters 16 and 17 both use the closure, and in both it appears at exactly the two points named above: the start set, and after each symbol.

Worked example: running P1

Run P1 on the input a, keeping closures throughout.

Start: the closure of p, which is {p, q, r, t}. Note that this set already contains t, which is final, so P1 accepts the empty string without reading anything.

munotes.in70

Empty Moves, and the Epsilon Closure

Read a. Follow a arrows from each of p, q, r, t: p on a gives p, q on a gives s, r and t have none. That gives {p, s}. Now take the closure: the closure of p is {p, q, r, t} and the closure of s is {s}, so the set is {p,q,r,s,t}.

That contains t, so a is accepted.

Accepts: ε, a, aa, b, bb, ab

Rejects: ba, abba, bab

The claim lists are executed by the checker, which computes closures exactly as above. And to see what P1's language actually is:

L(P1) = L(P1R)

a* (b* + a b*)

which is the reading the drawing gives: any number of a at the start, and then either sit in r reading b, or take one a into s and read b there.

The two smallest empty move machines, and why they matter

These are the building blocks of chapter 34, so they are worth meeting on their own.

A machine for the empty string, with one empty move and no symbols read:

P2Stateεa
starty0y1-
finaly1--

Accepts: ε

Rejects: a, aa

A machine for the union of two things, where the start state has an empty move into each of two sub machines. Here the two sub machines accept a and b respectively:

P3Stateεab
startx0{x1, x3}--
x1-x2-
finalx2---
x3--x4
finalx4---

Accepts: a, b

Rejects: ε, ab, ba, aa

L(P3) = L(P3R)

a + b

Notice what the empty moves have bought: the two sub machines were not touched at all. Their tables are unchanged and a new start state was added with two empty arrows. That is the whole technique of chapter 34, and it only works because an arrow can consume nothing.

Distinctions

A symbol moveAn empty move
Consumes inputone symbolnothing
Labelledwith a member of Sigmawith epsilon
Appears in an input stringyesnever
Can be taken at any timeonly when that symbol is nextalways
Adds poweryes, it is how input is readno, chapter 17 removes them all

What it does NOT mean

Epsilon is not a symbol of the alphabet. It labels an arrow. No input string ever contains it, and writing it into Sigma is an error that breaks the definition.

The closure is not just the epsilon cell. It is the transitive closure: follow empty arrows as far as they go. Two steps, three steps, any number.

munotes.in71

Empty Moves, and the Epsilon Closure

The closure of q is never empty. It always contains q. A closure computed as the empty set is an arithmetic mistake.

An empty move is not a free pass. It does not skip a symbol of the input; it moves between states without reading. The input still has to be read in full by symbol moves.

Empty moves add no power. They make constructions possible and machines smaller to write. Chapter 17 removes them all.

Quick revision

  • An epsilon-NFA lets delta take epsilon as well as the symbols of Sigma, so a state may change without

reading input.

  • Epsilon is not in Sigma and never appears in an input string.
  • The epsilon closure of q is every state reachable by zero or more empty moves, and it always contains

q itself.

  • Algorithm: start with q, repeatedly add the epsilon cell of anything in the answer, stop when nothing

new appears.

  • Running the machine: begin at the closure of the start state, and after each symbol take the closure

of the result.

  • The machine accepts the empty string exactly when the closure of the start state contains a final

state.

  • Empty moves exist because the constructions of chapters 33, 34 and 38 need to glue machines together

without looking inside them. They add no power.

Test yourself

1. Define the epsilon closure of a state. The set of all states reachable from it by a sequence of zero or more epsilon transitions. Because zero is allowed, the state itself is always a member.

2. A machine has an epsilon arrow from p to q and from q to r, and none other. Give the closures of p, q and r. Closure of p is {p, q, r}; closure of q is {q, r}; closure of r is {r}.

3. When does an epsilon-NFA accept the empty string? Exactly when the epsilon closure of its start state contains a final state, since no symbols are read.

4. What are the two places a closure is taken when running such a machine? At the start, to get the initial set, and after following the arrows for each input symbol.

5. Why are empty moves useful if they add no power? Because they allow two machines to be joined without examining their tables, which is what makes the construction from a regular expression to an automaton short and uniform.

6. A student computes the closure of a state as the empty set. What has gone wrong? They have forgotten that zero steps count, so the state itself is always in its own closure. No closure is ever empty.

Contents This chapter on its own page

munotes.in72

Chapter Sixteen

DFA and NDFA Equivalence: the Subset Construction

Syllabus topic Module 1, "Automata Theory: DFA and NDFA equivalence"

In one line

For every nondeterministic machine there is a deterministic one accepting the same language, and it is built by taking the sets of states of the first as the states of the second.

In the wording a student can write in an examination: for every NFA M there exists a DFA M' such that L(M') equals L(M). The DFA is constructed by taking as its states the subsets of the state set of M, as its start state the set containing only the start state of M, as its transition on a symbol a the union of the a moves of every state in the subset, and as its final states those subsets containing at least one final state of M.

Why the theorem is surprising, and why it matters

It is surprising because nondeterminism looks like extra power. A machine that can be in several states at once, exploring several possibilities, ought to be able to do more than one following a single path. The theorem says it cannot.

It matters for three separate reasons.

It means the definition of regular is safe. A language is regular when a DFA accepts it. If an NFA could accept more, there would be two notions of regular and the whole of Module 1 would have to pick one. The theorem says there is only one notion.

It is how every practical tool works. A search tool given a pattern builds a nondeterministic machine, because that is easy, and then converts it to a deterministic one, because that is what runs fast. Both halves of that sentence are chapters of this book: chapter 34 builds the NFA, this chapter converts it.

It is a template. The same trick, subsets as states, appears again in chapter 39 for the decision problems and in the product construction of chapter 38.

And it is worth noting immediately that the corresponding statement is false for the machine of chapter 52. A nondeterministic pushdown automaton is strictly stronger than a deterministic one, and chapter 57 gives the witness language. So this theorem is a fact about finite automata and not a general law, which is exactly why it has to be proved.

The construction

Let M be an NFA with states Q, alphabet Sigma, transition delta, start q0 and final states F. Build M' as follows.

States. The subsets of Q. So if M has n states, M' has up to 2 to the n, and chapter 2 counted them.

Start state. The set {q0}, containing the one start state of M and nothing else.

Transition. For a subset S and a symbol a, the new transition is the union of delta(p, a) over every p in S. That set is again a subset of Q, so it is a state of M'.

munotes.in73

DFA and NDFA Equivalence: the Subset Construction

delta'(S, a) = the union of delta(p, a) over all p in S

Final states. Every subset S that contains at least one member of F. Not the subsets contained in F: at least one member, because an NFA accepts when some run succeeds.

That last point is the commonest error in the construction, and it is worth a sentence of its own. A subset like {q1, q3} where only q3 is final IS a final state of M', because the run that ended in q3 is a successful run.

Building only what is needed

The construction as stated builds 2 to the n states, and almost all of them are usually unreachable. The practical version, and the one the examination expects, builds outward from the start.

  1. Write down the start set {q0} as the first row.
  2. Take any row not yet expanded, and for each symbol compute the new set.
  3. Any set that has not appeared before becomes a new row.
  4. Stop when every row has been expanded.
  5. Mark as final every row whose set contains a final state of M.

This is a reachability search, and it terminates because there are finitely many subsets. It usually produces far fewer than 2 to the n rows, and for the family in chapter 14 it produces exactly 2 to the n, which is why the bound cannot be improved.

With empty moves

If M has empty moves, the construction changes in exactly the two places chapter 15 identified.

The start state is the epsilon closure of {q0}, not {q0}.

The transition is the epsilon closure of the union, not the union.

delta'(S, a) = ε closure of ( the union of delta(p, a) over all p in S )

Everything else is unchanged. A machine with no empty moves has closures equal to the states themselves, so the two versions agree, and it is safe to learn only the second.

Worked example 1: the aba machine

Convert N1 from chapter 14. Its table again, so this chapter stands alone:

Q1Stateab
startn0{n0, n1}n0
n1-n2
n2n3-
finaln3n3n3

Accepts: aba, ababa

Rejects: ε, ab, bb

Now build outward. The start set is {n0}.

Seton aon b
{n0}{n0, n1}{n0}
{n0, n1}{n0, n1}{n0, n2}
{n0, n2}{n0, n1, n3}{n0}
{n0, n1, n3}{n0, n1, n3}{n0, n2, n3}
{n0, n2, n3}{n0, n1, n3}{n0, n3}
{n0, n3}{n0, n1, n3}{n0, n3}

Six sets appeared, out of the sixteen subsets of a four state machine. Working through row three as the one that shows the method: from {n0, n2} on a, n0 gives {n0, n1} and n2 gives {n3}, so the union is {n0, n1, n3}; and on b, n0 gives {n0} and n2 gives nothing, so the union is {n0}.

munotes.in74

DFA and NDFA Equivalence: the Subset Construction

The final states are the three sets containing n3: {n0, n1, n3}, {n0, n2, n3} and {n0, n3}.

Relabel the six sets A to F in the order above and the DFA is:

Q2Stateab
startABA
BBC
CDA
finalDDE
finalEDF
finalFDF

Accepts: aba, aabab, babaa, ababa, bbabab

Rejects: ε, ab, abb, aab, bbb, babb

L(Q2) = L(Q1)

Q2 is deterministic

Q1 is nondeterministic

The equality is decided exactly, over every string there is, by comparing the two machines. So the construction has not merely been carried out; the result has been proved right.

Six states, and chapter 13's hand designed machine for the same language had four. The subset construction is not obliged to give the smallest machine, and chapter 20 is how to shrink it.

Worked example 2: with empty moves

Convert P1 from chapter 15, which has three empty moves.

Q3Stateεab
startp{q, r}p-
q-s-
rt-r
s--s
finalt---

Accepts: ε, a, b

Rejects: ba, bab

The closures were computed in chapter 15: the closure of p is {p, q, r, t}, of r is {r, t}, and each of q, s, t is its own closure.

Start state: the closure of {p}, which is {p, q, r, t}. Call it S1.

S1 on a. The a moves: p gives p, q gives s, and r, t give nothing. Union {p, s}. Closure: {p, q, r, t} union {s}, which is {p,q,r,s,t}. Call it S2.

S1 on b. The b moves: r gives r, and the others give nothing. Union {r}. Closure {r, t}. Call it S3.

S2 on a. The a moves from p, q, r, s, t: p gives p, q gives s. Union {p, s}, closure S2 again.

S2 on b. r gives r, s gives s. Union {r, s}, closure {r, s, t}. Call it S4.

S3 on a. Neither r nor t has an a move, so the union is empty, and the closure of the empty set is empty. That is the dead state, which must be included to make the machine total.

S3 on b. r gives r. Closure {r, t}, which is S3.

munotes.in75

DFA and NDFA Equivalence: the Subset Construction

S4 on a. No a moves from r, s, t, so dead.

S4 on b. r gives r and s gives s. Union {r, s}, closure S4.

The final sets are those containing t: S1, S2, S3 and S4, all four of them. So:

Q4Stateab
start finalS1S2S3
finalS2S2S4
finalS3deadS3
finalS4deadS4
deaddeaddead

Accepts: ε, a, aa, b, bb, ab

Rejects: ba, abba, bab

L(Q4) = L(Q3)

Q4 is deterministic

Two things to notice. The empty set turned up as a state and had to be kept, because a DFA needs a total transition function; it is the dead state under another name. And S1 is both start and final, so the machine accepts the empty string, which is right because the closure of p contains the final state t.

The proof

MU asks for the equivalence to be proved, so here it is in full. It is an induction on the length of the string, of exactly the shape chapter 6 described, and the whole content is one identity.

Claim. For every string w over Sigma,

delta hat prime( {q0}, w ) = delta hat( q0, w )

that is, the single state M' reaches on w is precisely the set of states M could reach on w.

Base case. w is the empty string. The left side is delta hat prime({q0}, epsilon), which is {q0} by chapter 11's definition applied to M'. The right side is delta hat(q0, epsilon), which is {q0} by chapter 14's definition applied to M. They agree.

Inductive step. Assume the identity for w, and take one more symbol a. Then

delta hat prime( {q0}, w a ) = delta prime( delta hat prime({q0}, w), a )

= delta prime( delta hat(q0, w), a )

= the union of delta(p, a) over p in delta hat(q0, w)

= delta hat(q0, w a)

The first line is the definition of the extended function for M'. The second uses the inductive hypothesis. The third is the definition of delta prime, which is what the construction set it to be. The fourth is the definition of the extended function for M. So the identity holds for wa, and by induction for every string.

Finishing. M' accepts w exactly when delta hat prime({q0}, w) is a final state of M', which by the construction means exactly when that set contains a member of F. By the claim that set is delta hat(q0, w), so M' accepts w exactly when delta hat(q0, w) contains a member of F, which is exactly when M accepts w. Hence L(M') equals L(M).

munotes.in76

DFA and NDFA Equivalence: the Subset Construction

And the converse direction needs no work at all. Every DFA is an NFA in which every cell holds a set of size one, so a language accepted by a DFA is accepted by an NFA trivially. The two inclusions together give the theorem.

The cost

M with n states gives M' with at most 2 to the n, and chapter 14's family shows the bound is reached: the language whose k-th symbol from the right is an a has an NFA with k plus 1 states and a minimal DFA with 2 to the k.

NFA statesDFA states, worst case
38
532
101024
201048576

That is why the conversion is done once, when a pattern is compiled, and not on every input.

Distinctions

The NFA MThe DFA M'
A state isa statea SET of states of M
Startq0{q0}, or its closure if there are empty moves
A cella set of statesone set, which is one state of M'
Finala state in Fany set meeting F in at least one element
Number of statesnup to 2 to the n

What it does NOT mean

The construction does not give the smallest DFA. Example 1 gave six states where four suffice. Chapter 20 minimises.

A final state of M' is not a subset of F. It is a subset meeting F. Requiring containment would give a machine accepting a different, smaller language.

The empty set is a legitimate state. It is the dead state, and leaving it out makes the transition function partial and the machine not a DFA.

Equivalence of the two models is not a general law. It is a theorem about finite automata. For pushdown automata it is false, and for the linear bounded automaton of chapter 61 it is an open question.

The exponential blow up is not always suffered. Most machines convert to something small. The bound is worst case and is achieved only by families built for the purpose.

Quick revision

  • Theorem: for every NFA there is a DFA accepting the same language, so nondeterminism adds no power to

a finite automaton.

  • Construction: states are subsets of Q; start is {q0}, or its epsilon closure; the cell for S on a is

the union of the a moves of the members of S, closed under epsilon moves; final states are the subsets meeting F in at least one element.

  • Build only the reachable subsets, outward from the start, and include the empty set as the dead state.
  • Proof: by induction on string length, delta hat prime({q0}, w) equals delta hat(q0, w); acceptance

then matches by the definition of the final states.

munotes.in77

DFA and NDFA Equivalence: the Subset Construction

  • The converse is trivial: a DFA is an NFA with singleton cells.
  • Cost: up to 2 to the n states, and the bound is achieved by the k-th from the right family.
  • The subset machine need not be minimal.

Test yourself

1. State the four parts of the subset construction. States are the subsets of Q; the start state is the epsilon closure of {q0}; the transition for S on a is the epsilon closure of the union of delta(p, a) over p in S; a subset is final when it contains at least one final state of the NFA.

2. An NFA has 4 states, one of them final. How many states could the DFA have, and how many of them could be final? Up to 16. Of those, 8 contain the final state, so up to 8 are final.

3. Why must final states of M' be subsets meeting F rather than subsets of F? Because the NFA accepts when at least one run succeeds. A set containing one final state and three non final ones represents a successful run and must accept.

4. What is the empty set doing in the constructed machine? It is the state reached when every run has died, and it is the dead state: not final, and mapping to itself on every symbol. Including it is what makes the transition function total.

5. Give the base case and the inductive step of the equivalence proof. Base: on the empty string both sides are {q0}. Step: assuming the identity for w, the definition of delta prime and of the extended functions turn delta hat prime({q0}, wa) into the union of delta(p, a) over p in delta hat(q0, w), which is delta hat(q0, wa).

6. Does the theorem hold for pushdown automata? No. A nondeterministic pushdown automaton is strictly more powerful than a deterministic one, and the language of even length palindromes is the witness, which chapter 57 works through.

Contents This chapter on its own page

munotes.in78

Chapter Seventeen

Removing the Empty Moves

Syllabus topic Module 1, "Automata Theory: Nondeterministic Finite State Machines"

In one line

Every empty move can be removed by giving each state the moves its empty successors had, and the language does not change.

In the wording a student can write in an examination: for every NFA with epsilon transitions there is an NFA without them accepting the same language. It is obtained by defining, for each state q and each symbol a, the new move from q on a to be the epsilon closure of the set of states reachable on a from the epsilon closure of q, and by making final every state whose epsilon closure contains a final state.

Why remove them at all

Two reasons, one theoretical and one practical.

Theoretically it completes the picture. Chapter 15 added empty moves to make the constructions possible, and a reader is entitled to ask whether that addition bought new languages. It did not, and this chapter is the proof.

Practically it shortens the chain. Chapter 34 produces a machine with many empty moves, and it must end up deterministic. There are two routes: remove the empty moves and then apply the subset construction, or apply the subset construction with closures built in as chapter 16 described. Both work, and knowing the first lets you check the second.

Method 1: the closure method

This is the one MU's papers ask for and the one to learn first.

Let M be an NFA with epsilon moves, states Q, start q0, final F. Build M' with the same states and the same start state, and:

The new transition. For every state q and every symbol a of Sigma,

delta'(q, a) = ε closure of ( the union of delta(p, a) over all p in ε closure of q )

Read it from the inside out, which is also the order to compute it in: take the closure of q; from every state in it, follow the a arrows; then take the closure of the result.

The new final states. F', the new set of final states, is F together with every state whose epsilon closure contains a member of F.

F' = F union { q : ε closure of q contains a state of F }

That second part is the step students forget. If q can reach a final state by empty moves alone, then q must itself be final in the new machine, because in the old machine a run ending at q could have slipped into the final state for free. In particular the start state becomes final whenever the old machine accepted the empty string.

The empty moves are then deleted.

Worked example, method 1

Take P1 from chapter 15 once more, so this chapter stands alone.

munotes.in79

Removing the Empty Moves

R1Stateεab
startp{q, r}p-
q-s-
rt-r
s--s
finalt---

Accepts: ε, a, b, ab, aa

Rejects: ba, bab, abba

The closures, from chapter 15:

Stateε closure
p{p, q, r, t}
q{q}
r{r, t}
s{s}
t{t}

The new a column.

For p: the closure is {p, q, r, t}; a moves from those give p (from p) and s (from q), so {p, s}; the closure of that is {p, q, r, t} union {s}, which is {p,q,r,s,t}.

For q: the closure is {q}; a from q gives s; the closure of {s} is {s}.

For r: the closure is {r, t}; neither has an a move; so the empty set.

For s: the closure is {s}; no a move; empty.

For t: the closure is {t}; no a move; empty.

The new b column.

For p: the closure is {p, q, r, t}; b moves give r (from r); the closure of {r} is {r, t}.

For q: closure {q}; no b move; empty.

For r: closure {r, t}; b from r gives r; closure {r, t}.

For s: closure {s}; b from s gives s; closure {s}.

For t: closure {t}; no b move; empty.

The new final states. t is final already. The closure of p contains t, so p becomes final. The closure of r contains t, so r becomes final. The closures of q and s do not, so they do not.

Putting it together, with the epsilon column gone:

R2Stateab
start finalp{p,q,r,s,t}{r, t}
qs-
finalr-{r, t}
s-s
finalt--

Accepts: ε, a, b, ab, aa, bb

Rejects: ba, bab, abba

L(R2) = L(R1)

The checker decides that equality exactly, so the removal has not merely been carried out, it has been proved to preserve the language.

Notice that p is now marked final, which it was not before. That is the step 2 correction, and without it the new machine would reject the empty string while the old one accepts it.

Method 2: the direct method

Sometimes quicker by hand, and it gives the same answer.

Instead of computing closures, look at each empty arrow in turn and push it through:

  1. For each empty arrow from p to q, copy every outgoing symbol move of q onto p as well.
  2. If q is final, mark p final.
  3. Repeat until nothing changes, because a copied move may itself need pushing through another empty
munotes.in80

Removing the Empty Moves

arrow.

  1. Delete the empty arrows.

Applied to R1: the arrows are p to q, p to r, and r to t.

Pushing r to t first, because t's moves are what r must gain: t has no moves, so nothing is copied, but t is final, so r becomes final.

Pushing p to r: copy r's b move, which is r, onto p, so p gains b to r. And r is now final, so p becomes final.

Pushing p to q: copy q's a move, which is s, onto p, so p gains a to s.

The result has p with a to {p, s} and b to {r}, q with a to s, r with b to r, and p, r, t final. That is the same machine as R2 with some redundant states in the cells collapsed, and the checker confirms the language is the same.

The two methods differ only in bookkeeping. Method 1 computes the whole closure once per state; method 2 propagates one arrow at a time and has to be iterated. Method 1 is safer under examination conditions because it cannot be stopped too early.

The proof that the language is unchanged

Claim. For every non empty string w, delta hat prime(q0, w) equals the epsilon closure of delta hat(q0, w).

Sketch, by induction on the length of w. For a string of length one, the definition of delta prime is exactly the right hand side. For a longer string wa, the inductive hypothesis gives the set after w as a closed set, and delta prime applied to a closed set then follows a moves and closes again, which is what delta hat does on the original machine. So the two agree.

The empty string is the case the claim excludes, and it is handled by the new final states. On the empty string delta hat prime(q0, epsilon) is {q0}, while delta hat(q0, epsilon) is the closure of q0. The two differ whenever q0 has empty moves. But the second part of the construction made q0 final exactly when its closure contains a final state, so the two machines still agree about whether the empty string is accepted.

That is why the final state adjustment is not a tidying up detail: it is the whole of the empty string case of the proof.

Distinctions

Remove empty moves, then subsetSubset with closures built in
Number of stagestwoone
Closures computedonce per stateonce per set, at each step
Resultthe same DFA, up to state namesthe same
When to preferwhen the question asks for an NFA without epsilonwhen the question asks for a DFA
munotes.in81

Removing the Empty Moves

What it does NOT mean

Removing empty moves does not make the machine deterministic. The result is still an NFA: cells may hold several states. Determinism needs chapter 16 as well.

It does not reduce the number of states. The state set is unchanged. Only the arrows and the final markings change.

The final state adjustment is not optional. Skipping it breaks the machine on the empty string, and on every string that ends where an empty move would have led to acceptance.

The new machine may have more arrows. That is normal: the work an empty move used to do is now distributed among the symbol moves.

Quick revision

  • Every epsilon-NFA has an equivalent NFA without empty moves, with the same states and the same start

state.

  • New move: delta'(q, a) is the closure of the union of the a moves from the closure of q. Compute from

the inside out.

  • New final states: the old ones, plus every state whose closure contains an old final state. This is the

step that preserves acceptance of the empty string.

  • Method 2, equivalently: for each empty arrow from p to q, copy q's symbol moves onto p and make p final

if q is, and repeat until nothing changes.

  • The state set does not change; the arrows and the final markings do.
  • The result is an NFA, not a DFA; chapter 16 is still needed for determinism.

Test yourself

1. Give the formula for the new transition. delta'(q, a) is the epsilon closure of the union of delta(p, a) taken over every p in the epsilon closure of q.

2. Which states become final that were not, and why? Every state whose epsilon closure contains an old final state, because in the old machine a run could reach that final state from there for free, and in the new machine there are no free moves to reach it with.

3. A machine has an empty arrow from the start state to a final state. What must the new machine do? Make the start state final, so that the empty string is still accepted.

4. After removing empty moves, is the machine deterministic? No. It is an NFA with no epsilon column; cells may still hold several states. The subset construction is a separate step.

5. Does removing empty moves change the number of states? No. The state set and the start state are unchanged; only the transitions and the final markings differ.

6. Why must method 2 be repeated until nothing changes? Because pushing one empty arrow through can create a move that must itself be pushed through another empty arrow. A single pass can stop before the machine is correct.

Contents This chapter on its own page

munotes.in82

Chapter Eighteen

Mealy and Moore Machines: a Machine That Writes

Syllabus topic Module 1, "Automata Theory: Mealy and Moore Machines"

In one line

A Mealy or Moore machine reads an input string and writes an output string, one symbol at a time, instead of answering yes or no.

In the wording a student can write in an examination: a Moore machine is a sextuple (Q, Sigma, Delta, delta, lambda, q0) in which lambda maps Q to Delta, so the output depends on the state alone. A Mealy machine is a sextuple of the same shape in which lambda maps Q times Sigma to Delta, so the output depends on the state and the current input symbol. Neither has final states, because neither accepts a language: each defines a function from strings to strings.

Why a machine that writes is worth defining

Because most real sequential machines are of this kind, and because MU examines it.

A traffic light controller does not answer yes or no about the sequence of sensor readings it has seen. It emits a signal. A serial adder reads pairs of bits and emits sum bits. A protocol encoder reads a message and emits a coded one. All of these are finite state, and none of them is an acceptor.

So this is a different definition, not a variant of the same one, and the first thing to get straight is what is missing: there are no final states. Asking which strings a Mealy machine accepts is a category error. It transforms; it does not decide.

The two definitions

Both are sextuples, and they differ in one place only: what lambda, the output function, takes.

Moore: lambda maps Q to Delta

Mealy: lambda maps Q times Sigma to Delta

PartMeaning
Qthe finite set of states
Sigmathe input alphabet
Deltathe output alphabet, which may differ from Sigma
deltathe transition function, Q times Sigma to Q, exactly as before
lambdathe output function, and the only difference
q0the initial state

Note that Delta is a new letter and a new set. The machine may read 0 and 1 and write x and y; nothing requires the two alphabets to agree.

The output length trap

This is the thing that decides whether an examination answer is right, and it follows from the definitions with no extra rule.

A Moore machine's output depends on the state. The machine is in a state before it reads anything, so it emits a symbol before reading anything. On an input of length n it therefore emits n plus 1 symbols: one for the start state, and one for each state it moves into.

A Mealy machine's output depends on the state and the symbol. No symbol, no output. On an input of length n it emits exactly n symbols.

munotes.in83

Mealy and Moore Machines: a Machine That Writes

So on the empty input a Moore machine writes one symbol and a Mealy machine writes nothing. That single sentence is worth more marks in this topic than anything else in the chapter, and it is why a converted machine's output is never quite the same string as the original's.

A Moore machine, worked

The task. Read a string over {a, b} and emit 1 whenever the string read so far ends in an even number of a, and 0 otherwise. Two states, one per parity.

This book writes a Moore machine as a transition table with one extra column, headed Output, holding the symbol the state emits.

S1StateabOutput
starteoe1
oeo0

State e means an even number of a so far and emits 1; state o means odd and emits 0.

S1(ε) = 1

S1(a) = 10

S1(aa) = 101

S1(ab) = 100

S1(abab) = 10011

Read the third line: on input aa the machine emits 1 for the start state e, then moves to o and emits 0, then moves back to e and emits 1. Three symbols for two of input, as the rule says. Every one of those five output strings is produced by executing the table, so a mistyped digit would stop this book being built.

A Mealy machine, worked

The task. Read a string over {0, 1} and emit, for each symbol, a 1 if that symbol differs from the one before it and a 0 if it is the same. The first symbol has nothing before it, and we choose to treat it as differing, so it emits 1.

Two states are needed, recording what the last symbol was. This book writes a Mealy machine with a pair of columns per input symbol: one for the next state, one for the output.

S2State00 output11 output
startinitz1u1
zz0u1
uz1u0

State init means nothing has been read; z means the last symbol was 0; u means the last symbol was 1.

S2(ε) = ε

S2(0) = 1

S2(00) = 10

S2(0110) = 1101

S2(0101) = 1111

The first line is the trap made visible: on the empty input a Mealy machine writes nothing at all.

Line four: reading 0110, the first 0 emits 1 and moves to z; the 1 differs so emits 1 and moves to u; the second 1 is the same so emits 0 and stays in u; the 0 differs so emits 1 and moves to z. Output 1101, four symbols for four of input.

munotes.in84

Mealy and Moore Machines: a Machine That Writes

Three differences, and the fourth that is not a difference

MooreMealy
Output depends onthe state alonethe state and the input symbol
lambda mapsQ to DeltaQ times Sigma to Delta
Output length for input of length nn plus 1n
Output on the empty inputone symbolnothing
Written in a diagraminside the circle, after a slashon the arrow, after a slash
Number of states for a given joboften moreoften fewer
Final statesnonenone
Powerthe samethe same

The last row is the one to remember. Chapter 19 converts either into the other, so neither can do anything the other cannot. The choice between them is a matter of convenience and of how many states you are willing to have, exactly as the choice between a DFA and an NFA was.

Why Moore often needs more states. A Moore machine's output is fixed by the state, so if the same condition must sometimes emit one symbol and sometimes another, the condition has to be split into two states. A Mealy machine reads the symbol before emitting and so can decide without splitting. Chapter 19 makes this precise: converting a Mealy machine to a Moore machine can multiply the states by the size of the output alphabet.

Where they are used

In hardware description. A Moore machine's output is stable for a whole clock period because it does not depend on the input, which matters when the output drives something. A Mealy machine reacts within the period, which is faster but can glitch. Every textbook of digital design has this comparison, and it is the practical reason both models exist.

In encoders and protocol machines. Anything that reads a stream and writes a stream.

Not in this book after chapter 19. MU's syllabus places Mealy and Moore machines here and does not return to them, and the rest of Module 1 is about acceptors. So the two conversion procedures of chapter 19 are what the examination actually asks for, and that is where the work is.

What it does NOT mean

These machines do not accept languages. They have no final states. A question asking which strings a Mealy machine accepts has misunderstood the definition.

The output alphabet is not the input alphabet. It is a separate set, Delta, and it may be different.

Moore is not slower or weaker. The two are equally powerful, and chapter 19 proves it by construction in both directions.

A Moore machine's extra output symbol is not an error. It is the symbol of the start state, emitted before any input is read, and it is required by the definition.

Neither is nondeterministic here. Both definitions above use a transition function, so both are deterministic. Nondeterministic transducers exist and are outside this syllabus.

munotes.in85

Mealy and Moore Machines: a Machine That Writes

Quick revision

  • A Moore machine is (Q, Sigma, Delta, delta, lambda, q0) with lambda from Q to Delta: output per state.
  • A Mealy machine is the same with lambda from Q times Sigma to Delta: output per state and symbol.
  • Neither has final states; neither accepts a language; each computes a function from strings to strings.
  • Moore emits n plus 1 symbols on an input of length n, Mealy emits n. On the empty input Moore writes one

symbol and Mealy writes none.

  • In a diagram, a Moore output is written in the state circle and a Mealy output on the arrow.
  • The two are equally powerful, and chapter 19 converts each into the other.
  • Moore often needs more states, because one state cannot emit two different symbols.

Test yourself

1. Give the one difference between the definitions, and the consequence for output length. lambda maps Q to Delta for Moore and Q times Sigma to Delta for Mealy. So Moore emits a symbol for the start state before reading anything, giving n plus 1 symbols on an input of length n, while Mealy emits n.

2. What does a Mealy machine output on the empty input, and what does a Moore machine? Mealy outputs nothing at all. Moore outputs one symbol, the output of its start state.

3. Do these machines have final states? What follows? No. It follows that they do not accept languages; they define functions from input strings to output strings, and asking what they accept is a category error.

4. For S1 above, what is the output on bb, and why? The machine starts in e and emits 1; b keeps it in e and emits 1; b again keeps it in e and emits 1. So 111, three symbols for two of input, and every symbol is 1 because zero a is an even number of a.

5. Why might a Moore machine need more states than a Mealy machine for the same task? Because its output is determined by the state alone, so a condition that must emit different symbols depending on the symbol just read has to be split into several states.

6. Which is more powerful? Neither. Chapter 19 gives constructions in both directions, so each can simulate the other, and the choice between them is about convenience and state count.

Contents This chapter on its own page

munotes.in86

Chapter Nineteen

Converting a Moore Machine to a Mealy Machine, and Back

Syllabus topic Module 1, "Automata Theory: Mealy and Moore Machines"

In one line

A Moore machine becomes a Mealy machine by moving each state's output onto the arrows coming into it, and a Mealy machine becomes a Moore machine by splitting each state into one copy per output symbol.

In the wording a student can write in an examination: for every Moore machine there is an equivalent Mealy machine, and for every Mealy machine there is an equivalent Moore machine, where equivalent means that for every input string the two produce the same output, except that the Moore machine's output carries one extra symbol at the front, being the output of its initial state.

What equivalent has to mean here

The definition of equivalence needs stating before either construction, because the output lengths differ and so the two machines can never produce literally the same string.

Chapter 18 established that on an input of length n a Moore machine writes n plus 1 symbols and a Mealy machine writes n. So:

A Moore machine and a Mealy machine are equivalent when, for every input w, the Moore output equals the output of the initial state followed by the Mealy output.

That is the standard convention and it is what an examiner expects. Written out: if the Moore machine's initial state emits b, then Moore output equals b followed by Mealy output, for every input including the empty one, where the Mealy output on the empty input is the empty string.

Moore to Mealy

This direction is the easy one, because information is being thrown away rather than created.

The procedure.

  1. Keep the states, the start state, the transitions and both alphabets exactly as they are.
  2. For each state p and each symbol a, let the Mealy output on that arrow be the Moore output of the state

the arrow leads to, that is, of delta(p, a).

  1. Delete the state outputs.

The whole idea is that a Moore machine emits a symbol on arriving at a state, and a Mealy machine emits on traversing an arrow. Arriving at a state is the same event as traversing an arrow into it, so the output can simply be moved.

The state count does not change. This direction never needs a new state.

The one symbol that cannot be moved is the output of the initial state, which is emitted on arriving at q0 before any arrow has been traversed. There is no arrow to hang it on, which is exactly why the equivalence is stated with that symbol prefixed.

Worked: converting S1

S1 from chapter 18, reprinted so this chapter stands alone:

T1StateabOutput
starteoe1
oeo0

T1(ε) = 1

T1(abab) = 10011

Now apply the procedure. From e on a we reach o, whose output is 0, so that arrow carries 0. From e on b we reach e, whose output is 1, so that arrow carries 1. From o on a we reach e, output 1. From o on b we reach o, output 0.

munotes.in87

Converting a Moore Machine to a Mealy Machine, and Back

T2Stateaa outputbb output
starteo0e1
oe1o0

T2(ε) = ε

T2(abab) = 0011

Check the equivalence. T1 on abab gave 10011. The output of T1's initial state is 1. T2 on abab gave 0011. And 1 followed by 0011 is 10011. They agree, and on the empty input T1 gives 1 while 1 followed by T2's empty output is also 1.

Both output strings are produced by running the two tables, so the check above is not a claim about what should happen; it is a comparison of two executions.

Mealy to Moore

This direction is the one examinations set, and it is harder, because a Moore machine's output is tied to its state and a Mealy machine's state may need to emit different symbols on different arrows. The remedy is to split.

The procedure.

  1. For each state q of the Mealy machine, look at every output symbol that appears on an arrow into q.

Make one copy of q for each such symbol, named q with that symbol attached, and give that copy the symbol as its Moore output.

  1. If no arrow enters q at all, and q is the initial state, make a single copy of it and give it any output

symbol; the convention is to use the first symbol of Delta, and the examiner will accept a stated choice.

  1. Transitions. From the copy of q labelled with symbol b, on reading a, go to the copy of delta(q, a)

labelled with the Mealy output lambda(q, a). Note which machine each part comes from: the destination state is the Mealy destination, and the label on it is the Mealy output of the arrow just taken.

  1. The initial state of the Moore machine is the copy of q0 made in step 2, or, if arrows do enter q0, the

copy carrying the chosen initial output.

  1. There are no final states in either machine, so nothing is carried over.

The state count can grow. A Mealy machine with n states and an output alphabet of size k becomes a Moore machine with at most n times k states, because each state may need one copy per output symbol. In practice many copies are unreachable and are dropped.

Worked: converting S2

S2 from chapter 18, the machine that emits 1 when a symbol differs from its predecessor:

munotes.in88

Converting a Moore Machine to a Mealy Machine, and Back

T3State00 output11 output
startinitz1u1
zz0u1
uz1u0

T3(ε) = ε

T3(0110) = 1101

T3(0101) = 1111

Step 1: which outputs enter each state?

Arrows into z carry: 1 (from init on 0), 0 (from z on 0), 1 (from u on 0). So z needs copies for 0 and 1.

Arrows into u carry: 1 (from init on 1), 1 (from z on 1), 0 (from u on 1). So u needs copies for 0 and 1.

No arrow enters init, so by step 2 it gets a single copy. Its output is a free choice; take 0 and say so.

That gives five states: init, z0, z1, u0, u1, where the digit is the state's Moore output.

Step 3: the transitions. From any copy of a state, reading a symbol, go to the copy of the Mealy destination labelled with the Mealy output.

From init on 0: Mealy goes to z with output 1, so go to z1. From init on 1: to u with output 1, so u1.

From either copy of z on 0: Mealy goes to z with output 0, so z0. On 1: to u with output 1, so u1.

From either copy of u on 0: Mealy goes to z with output 1, so z1. On 1: to u with output 0, so u0.

T4State01Output
startinitz1u10
z0z0u10
z1z0u11
u0z1u00
u1z1u01

T4(ε) = 0

T4(0110) = 01101

T4(0101) = 01111

Check the equivalence. T3 on 0110 wrote 1101. T4's initial output is 0. And T4 on 0110 wrote 01101, which is 0 followed by 1101. They agree. The same holds for 0101: 1111 against 0 followed by 1111.

Five states where the Mealy machine had three, which is the growth the procedure warns about, and here it is three plus two rather than three times two because init needed only one copy.

Both conversions on one page

Moore to MealyMealy to Moore
Ideamove each state's output onto its incoming arrowssplit each state, one copy per incoming output
Statesunchangedup to n times the size of Delta
Transitionsunchangeddestination is the Mealy destination, labelled with the Mealy output
The initial outputcannot be moved, so it is prefixedchosen freely if no arrow enters q0
Difficultymechanicalneeds care over which label goes where

Why the constructions work

Moore to Mealy. In the Moore machine, the symbol emitted at step i of a run is the output of the state reached after i symbols. In the Mealy machine as constructed, the symbol emitted at step i is the output written on the arrow taken at step i, which the construction set to the output of the state that arrow leads to, which is the same state. So the two agree at every step from step one onwards, and step zero is the initial output, which is the prefixed symbol.

munotes.in89

Converting a Moore Machine to a Mealy Machine, and Back

Mealy to Moore. In the Moore machine as constructed, the copy the machine sits in after i symbols is labelled with the Mealy output of the arrow taken at step i. So the Moore output at step i, for i at least one, is exactly the Mealy output at step i. At step zero the Moore machine emits the label of the initial copy, which is the prefixed symbol. The transitions are faithful because the destination state of each Moore arrow is the Mealy destination, with only the label added, so the run visits the same Mealy states in the same order.

Both arguments are inductions on the length of the input, and both have the same shape as chapter 16's: assume the two runs agree after i symbols, then read one more.

What it does NOT mean

The two outputs are not equal as strings. The Moore output is one symbol longer, and the equivalence is stated with that symbol prefixed. An answer claiming the outputs are identical has not understood the definitions.

Mealy to Moore does not always multiply the states. It multiplies in the worst case. Here three became five, not six, because the initial state needed only one copy, and unreachable copies are always dropped.

The freely chosen initial output is not arbitrary in the answer. It is a free choice, but it must be stated, because the equivalence is only meaningful once the prefixed symbol is named.

Moore to Mealy does not lose information. The state outputs are recoverable, since every arrow into a state carries that state's output.

Quick revision

  • Equivalence for these machines: the Moore output equals the initial state's output followed by the Mealy

output, for every input.

  • Moore to Mealy: keep everything, and label each arrow with the output of the state it leads to. The state

count does not change. The initial output cannot be moved and is prefixed.

  • Mealy to Moore: split each state into one copy per output symbol appearing on an arrow into it; from a

copy of q on a, go to the copy of delta(q, a) labelled lambda(q, a); the initial state's output is a free stated choice when no arrow enters it.

  • Mealy to Moore can need up to n times the size of Delta states; unreachable copies are dropped.
  • Both constructions are proved by induction on input length, the two runs being shown to agree step by
munotes.in90

Converting a Moore Machine to a Mealy Machine, and Back

step.

Test yourself

1. State what equivalence means for a Moore and a Mealy machine. For every input string, the Moore output equals the output of the Moore machine's initial state followed by the Mealy output. The two are never equal as strings, because the Moore output is one symbol longer.

2. Convert this Moore machine to a Mealy machine: states p and q, p start with output 0 and q with output 1, p on a goes to q and on b to p, q on a goes to q and on b to p. Same two states. The arrow p on a leads to q, so it carries 1. The arrow p on b leads to p, so it carries

  1. The arrow q on a leads to q, so 1. The arrow q on b leads to p, so 0. The state outputs are deleted, and

the initial 0 is prefixed.

3. How many states can a Mealy machine with 4 states and an output alphabet of 3 symbols need after conversion? Up to 12, one copy of each state per output symbol, less however many copies turn out to be unreachable.

4. In the Mealy to Moore construction, where does the label on the destination copy come from? From the Mealy output of the arrow being taken, not from the destination state and not from the source.

5. Why is the initial state sometimes given a freely chosen output? Because a Moore machine must emit a symbol for its initial state, and if no arrow enters that state there is no Mealy output to take the symbol from. The choice must be stated, since it is the symbol the equivalence prefixes.

6. Does either conversion change what the machine can compute? No. Both directions exist, so the two models compute exactly the same functions from strings to strings, and the conversions are about convenience and state count.

Contents This chapter on its own page

munotes.in91

Chapter Twenty

Minimizing Automata: the Partition Method

Syllabus topic Module 1, "Automata Theory: Minimizing Automata"

In one line

Two states are the same for practical purposes when no string tells them apart, and the minimal machine has one state per group of states that nothing tells apart.

In the wording a student can write in an examination: two states p and q of a DFA are equivalent, written p is k-equivalent to q for every k, if for every string w the machine accepts from p exactly when it accepts from q. Equivalence is an equivalence relation on the states, and the minimum state automaton is obtained by discarding unreachable states and then taking the equivalence classes as states.

Why minimise

Three reasons, and the third is the one that matters mathematically.

Because the constructions produce waste. The subset construction of chapter 16 gave six states for a language that a hand designed machine did in four. Chapter 34's construction from a regular expression is worse. Anything built by a general procedure has redundancy in it.

Because a smaller machine is cheaper. A machine is a table, and the table is the memory a program uses.

Because the minimal machine is unique. This is the deep reason. For any regular language there is exactly one minimal DFA, up to the names of its states, and chapter 21 proves it. That makes the minimal machine a canonical object belonging to the language rather than to any particular construction, which is what lets two machines be compared: minimise both and see whether the results are the same. It is also what makes the decision problems of chapter 39 decidable.

The two things to discard

A machine can be too big in two different ways, and they are removed by two different procedures. Doing them in the wrong order wastes work, and forgetting the first is the commonest error.

Unreachable states. A state no input string can bring the machine to. It can be deleted outright with no thought at all, because nothing that happens there ever happens.

Indistinguishable states. Two reachable states that behave identically from here on. They cannot be deleted; they must be merged.

Do the unreachable ones first. An unreachable state can be distinguishable from everything, so leaving it in gives it a class of its own and the answer comes out too big. This is the step examiners look for.

Step one: finding the unreachable states

A reachability search from the start state.

  1. Mark the start state.
  2. Repeatedly, for every marked state and every symbol, mark whatever its cell names.
  3. Stop when nothing new gets marked.
  4. Delete every unmarked state and its row.

Step two: finding the indistinguishable states

Two states are distinguishable when some string is accepted from one and not from the other. Such a string is a distinguishing string, and producing one is how a claim of distinguishability is proved.

munotes.in92

Minimizing Automata: the Partition Method

The method builds up by the length of the distinguishing string, which is why it is called k-equivalence.

Definition. p and q are k-equivalent when, for every string w of length at most k, the machine accepts w from p exactly when it accepts w from q.

0-equivalence is the first partition: every string of length at most 0 is the empty string, and the empty string is accepted from p exactly when p is final. So the states split into two blocks, the final ones and the non final ones.

The refinement step. Given the k-equivalent blocks, p and q are (k plus 1)-equivalent when they are k-equivalent and, for every symbol a, the states delta(p, a) and delta(q, a) lie in the same k-equivalent block. In words: they agree now, and every symbol takes them to places that agree.

Stopping. Keep refining. Each round either splits a block or changes nothing, and when a round changes nothing the partition is final and the process stops. It must terminate, because a partition of a finite set can only be refined finitely often.

The answer. The blocks are the states of the minimal machine. A block is final if its members are final, which is well defined because 0-equivalence never puts a final state with a non final one. The transition from a block on a symbol is the block containing delta(p, a) for any p in it, which is well defined because that is exactly what the refinement guaranteed.

Worked example 1: eight states down to five

Start with W1, over {0, 1}.

W1State01
startABF
BGC
finalCAC
DCG
EHF
FCG
GGE
HGC

Accepts: 01, 10, 011, 101, 0111, 1011

Rejects: ε, 0, 1, 00, 11, 010, 111

Step one: reachability. Mark A. From A: B and F. From B: G and C. From F: C and G. From G: G and E. From C: A and C. From E: H and F. From H: G and C. Marked: A, B, C, E, F, G, H. D is never marked, so D is unreachable and its row goes. Seven states remain.

Notice that D's cells name C and G, both of which are reachable by other routes, so deleting D changes nothing about them.

Step two: 0-equivalence. The final states are {C} and the non final ones are {A, B, E, F, G, H}. Two blocks.

munotes.in93

Minimizing Automata: the Partition Method

Refining to 1-equivalence. For each state in the big block, record which block each symbol leads to, writing 1 for the final block and 2 for the non final one.

Stateon 0blockon 1block
AB2F2
BG2C1
EH2F2
FC1G2
GG2E2
HG2C1

Three patterns appear among these six: (2, 2) for A, E and G; (2, 1) for B and H; (1, 2) for F. So the big block splits into {A, E, G}, {B, H} and {F}, and with {C} that is four blocks.

Refining to 2-equivalence. Repeat with the four blocks, numbering them 1 = {C}, 2 = {A, E, G}, 3 = {B, H}, 4 = {F}.

Stateon 0blockon 1block
AB3F4
EH3F4
GG2E2
BG2C1
HG2C1
FC1G2
CA2C1

Within {A, E, G}: A gives (3, 4), E gives (3, 4), G gives (2, 2). So G separates from A and E, and the block splits into {A, E} and {G}. Within {B, H}: both give (2, 1), so they stay together. The blocks are now {C}, {A, E}, {G}, {B, H}, {F}: five of them.

Refining again. With five blocks, numbering 1 = {C}, 2 = {A, E}, 3 = {G}, 4 = {B, H}, 5 = {F}: A gives (4, 5) and E gives (4, 5), so {A, E} holds; B gives (3, 1) and H gives (3, 1), so {B, H} holds. Nothing splits, so the partition is final.

The minimal machine, five states, naming each block after its first member:

W2State01
startABF
BGC
finalCAC
FCG
GGA

Reading the last row: G on 1 goes to E in the original, and E now lives in the block called A, so it goes to A.

Accepts: 01, 10, 011, 101, 0111, 1011

Rejects: ε, 0, 1, 00, 11, 010, 111

W2 = minimal(W1)

That single line is the whole claim of this example, and the checker proves both halves of it: that the two machines accept exactly the same language, decided exactly by product construction, and that five is the minimum, decided by minimising W1 independently and counting.

Eight states down to five: one unreachable, and two merges.

Worked example 2: the subset construction's waste removed

Chapter 16 converted the aba machine and got six states, where the hand designed machine of chapter 13 had four. Here is the six state machine again.

munotes.in94

Minimizing Automata: the Partition Method

W3Stateab
startABA
BBC
CDA
finalDDE
finalEDF
finalFDF

Accepts: aba, aabab, ababa

Rejects: ε, ab, abb, aab

0-equivalence: finals {D, E, F}, non finals {A, B, C}.

1-equivalence. Among the finals: D on a gives D (final block), on b gives E (final block), so (F, F) writing F for the final block. E on a gives D (F), on b gives F (F), so (F, F). F likewise (F, F). No split. Among the non finals: A gives B (N) and A (N), so (N, N). B gives B (N) and C (N), so (N, N). C gives D (F) and A (N), so (F, N). So C separates: {A, B} and {C}.

2-equivalence. Blocks 1 = {D, E, F}, 2 = {A, B}, 3 = {C}. A gives B (2) and A (2). B gives B (2) and C (3). So A and B separate. Blocks: {D, E, F}, {A}, {B}, {C}: four.

3-equivalence. Among {D, E, F}: D gives D (1) and E (1); E gives D (1) and F (1); F gives D (1) and F (1). All (1, 1), no split. Nothing else can split. Four blocks, and the process stops.

W4Stateab
startABA
BBC
CDA
finalDDD

Accepts: aba, aabab, ababa, bbabab

Rejects: ε, ab, abb, aab, bbb

W4 = minimal(W3)

L(W4) = L(W3)

Four states, which is exactly the size of the machine designed by hand in chapter 13, and comparing the two tables shows they are the same machine with the states renamed. That is not luck; chapter 21 proves it had to happen.

Worked example 3: a machine that is already minimal

Worth doing, because a student must be able to say so rather than keep refining.

W5Stateab
startpqp
finalqqp

Accepts: a, aa, ba, aba

Rejects: ε, b, ab, bb

0-equivalence gives {p} and {q}, two blocks, and with two states there is nothing to refine. Both states are reachable. So the machine is minimal.

W5 = minimal(W5)

That claim is not a tautology to the checker: it minimises W5 and requires the answer to have the same number of states, so a machine wrongly called minimal would fail.

Distinctions

Unreachable stateIndistinguishable states
What is wrongnothing reaches ittwo behave identically
Found bya reachability search from the startthe k-equivalence refinement
Remedydelete itmerge them into one
Orderfirstsecond
Effect of skipping itthe answer comes out too bigthe answer comes out too big
munotes.in95

Minimizing Automata: the Partition Method

0-equivalencek plus 1-equivalence
Distinguishing strings of length0, that is, the empty stringat most k plus 1
The partitionfinals against non finalsk-equivalent, and agreeing under every symbol

What it does NOT mean

Minimising does not change the language. That is the whole point, and every claim in this chapter is checked for it.

Two states with the same row are not the only mergeable pair. In example 1, A and E have different rows (B and F against H and F) and merge anyway, because B and H turn out to be equivalent themselves. Comparing rows by eye finds some merges and misses others; the refinement finds all of them.

A final and a non final state are never equivalent. The empty string distinguishes them, which is exactly what 0-equivalence records.

The dead state does not disappear. If the machine needs a state from which nothing is accepted, the minimal machine has exactly one such state. Minimal means fewest states, not no waste of a human kind.

Minimising an NFA is a different and harder problem. Everything here is about DFAs. The minimal NFA for a language need not be unique, and finding it is far harder than this refinement. MU's label is about the DFA case.

Quick revision

  • Two states are equivalent when no string is accepted from one and rejected from the other. Such a string

is a distinguishing string.

  • Minimise in two steps, in this order: delete unreachable states, then merge indistinguishable ones.
  • 0-equivalence splits finals from non finals. Refine: p and q stay together when they are k-equivalent

and every symbol takes them into the same k-equivalent block.

  • Stop when a round changes nothing. Termination is guaranteed because a finite partition refines finitely

often.

  • The blocks are the states of the minimal machine; a block is final if its members are; the transition

from a block is the block containing the transition of any member.

  • The minimal DFA is unique up to state names, which chapter 21 proves, and that is what makes it

canonical.

  • Two states with different rows may still merge. Comparing rows by eye is not the method.

Test yourself

1. Give the two things minimisation removes, and the order. Unreachable states first, by a reachability search, deleted outright. Then indistinguishable states, by k-equivalence refinement, merged. Skipping the first makes the answer too big.

2. What partition is 0-equivalence, and why? The final states against the non final states, because the only string of length 0 is the empty string, and it is accepted from a state exactly when that state is final.

munotes.in96

Minimizing Automata: the Partition Method

3. State the refinement condition. p and q are (k plus 1)-equivalent when they are k-equivalent and, for every symbol a, delta(p, a) and delta(q, a) lie in the same k-equivalent block.

4. How do you know the process stops? Because each round either splits a block or changes nothing, and a partition of a finite set of states can only be refined a finite number of times.

5. Two states have identical rows in the table. Are they necessarily equivalent? And if their rows differ, are they necessarily distinguishable? Identical rows with the same final status do make them equivalent. Different rows do not make them distinguishable: in this chapter's first example A and E have different rows and merge, because the states their rows name are themselves equivalent.

6. Why does the uniqueness of the minimal DFA matter? Because it makes the minimal machine a property of the language rather than of a construction, which is what allows two machines to be compared by minimising both, and what makes the equivalence problem of chapter 39 decidable.

Contents This chapter on its own page

munotes.in97

Chapter Twenty-One

The Myhill Nerode Theorem, and a Second Way to Prove a Language Not Regular

Syllabus topic Module 1, "Automata Theory: Minimizing Automata"

In one line

Two strings are equivalent for a language when no ending distinguishes them, and the language is regular exactly when there are finitely many classes of equivalent strings, one for each state of the minimal machine.

In the wording a student can write in an examination: for strings x and y over Sigma, write x is L-equivalent to y if for every string z, xz is in L exactly when yz is in L. This is an equivalence relation on Sigma star. The Myhill Nerode theorem states that L is regular if and only if the relation has finitely many equivalence classes, and in that case the number of classes equals the number of states of the minimal DFA accepting L.

Why this theorem is the right way round

The pumping lemma of chapter 36 says: if a language is regular then it has a certain property. That is a one way implication, and its only use is the contrapositive of chapter 6: if the property fails, the language is not regular. It can never prove a language regular, and it can never prove a language not regular when the property happens to hold anyway, which does happen.

Myhill Nerode says if and only if. That makes it strictly stronger, and it makes it useful in both directions. It also tells you the number of states, which the pumping lemma has nothing to say about.

The cost is that the relation is on an infinite set and its classes must be identified by an argument rather than computed, which is why examinations tend to set the pumping lemma. Both are worth having, and a student who knows only one is missing the better half.

The relation

Fix a language L over Sigma. For strings x and y, say x and y are L-equivalent when

for every z in Sigma star: x z is in L exactly when y z is in L

In words: whatever you write after them, the two behave the same way. Such a z, when it exists and distinguishes them, is a distinguishing extension.

It is an equivalence relation. Reflexive, because x and x obviously agree. Symmetric, because the condition is symmetric in x and y. Transitive, because if x agrees with y and y with w on every z, then x agrees with w. So chapter 4 applies: it partitions Sigma star into classes, and the number of them is the index of the relation.

The whole force of the theorem is that an infinite set can have a relation of finite index, which is exactly the observation chapter 4 closed on.

munotes.in98

The Myhill Nerode Theorem, and a Second Way to Prove a Language Not Regular

The first half: a regular language has finite index

Suppose L is regular, so some DFA M with n states accepts it.

If delta hat(q0, x) equals delta hat(q0, y), then x and y are L-equivalent. The reason is the identity of chapter 11: delta hat(q0, xz) equals delta hat from the state reached after x on z, and the same for y, so if those states are the same then both end in the same state on every z, and so agree about acceptance.

So the map sending a string x to the state delta hat(q0, x) is constant on L-equivalence classes, and therefore there are at most as many classes as there are states. At most n classes: finite.

The second half: finite index gives a machine

Suppose the relation has finitely many classes. Build a machine whose states are the classes.

States: the L-equivalence classes. Start state: the class of the empty string. Transition: from the class of x on symbol a, go to the class of xa. This is well defined and it is the only thing that needs checking: if x and y are L-equivalent, then xa and ya are L-equivalent, because any z distinguishing xa from ya would make az a distinguishing extension of x and y. Final states: the classes whose members are in L. Well defined because the empty string is a legal z, so two L-equivalent strings are both in L or both out.

Then, by induction on the length of w, the machine reaches the class of w on input w, so it accepts w exactly when the class of w is final, which is exactly when w is in L. So L is regular, and the machine has one state per class.

Putting the halves together gives the theorem, and one more thing: the machine just built has exactly index-many states, and the first half showed no machine can have fewer. So it is minimal, and it was constructed from the language alone with no reference to any machine. That is the uniqueness chapter 20 promised: any minimal DFA for L must have index-many states, and merging its states by the relation gives this same machine, so all minimal DFAs for L are the same up to the names of their states.

Worked example 1: a regular language, and its index

Let L be the strings over {a, b} that end in ab. Chapter 13's machine has three states, so the index should be 3, and here are the three classes with the reasoning.

The only thing that matters about a string, for deciding whether anything appended to it ends in ab, is how much of ab it currently ends with. So the candidate classes are:

Class, named by a memberThe strings in itWhy it is its own class
εends in neither a nor abappending b leaves it out of L
aends in a but not in abappending b puts it in L
abends in abappending nothing leaves it in L
munotes.in99

The Myhill Nerode Theorem, and a Second Way to Prove a Language Not Regular

Three, and no more: every string falls into one of the three, because it either ends in ab, or ends in a without ending in ab, or neither.

Three, and no fewer: each pair is distinguished by an explicit z, which is what a proof requires.

PairDistinguishing zBecause
ε and abb is not in L; ab is in L
ε and abεε is not in L; ab is in L
a and abεa is not in L; ab is in L

So the index is exactly 3, the language is regular, and the minimal DFA has 3 states. Here it is, being the machine of chapter 13 with its states renamed after the classes:

Y1Stateab
startnoneseenAnone
seenAseenAseenAB
finalseenABseenAnone

Accepts: ab, aab, bab, abab

Rejects: ε, a, b, ba, aba, abb

Y1 = minimal(Y1)

The claim is checked, so the machine really does have the smallest possible number of states, which the index argument predicted.

Worked example 2: a language proved NOT regular

Let L be the set of strings a to the n followed by b to the n, for n at least 0. So L holds epsilon, ab, aabb, aaabbb and so on.

Claim. L is not regular.

Proof by Myhill Nerode. Consider the infinitely many strings

ε, a, aa, aaa, aaaa, ...

that is, a to the i for every i at least 0. Take any two of them with different exponents, say a to the i and a to the j with i less than j. Then the extension

z = b to the i

distinguishes them: a to the i followed by b to the i is in L, while a to the j followed by b to the i is not, because j is not i.

So no two of these strings are L-equivalent, and there are infinitely many of them, so there are infinitely many classes. The index is infinite, and by the theorem L is not regular.

That is the whole proof. Compare it with the pumping lemma proof of the same fact in chapter 37, which needs a supposed machine, a pigeonhole argument and a case analysis. This one needs an infinite family and one distinguishing extension.

The shape to remember, because every such proof has it:

munotes.in100

The Myhill Nerode Theorem, and a Second Way to Prove a Language Not Regular

  1. Name an infinite family of strings, indexed by a number.
  2. Take any two of them with different indices.
  3. Produce one z, written in terms of the indices, that is accepted after one and not after the other.
  4. Conclude that there are infinitely many classes, so the language is not regular.

Step 3 is the whole work, and step 1 is where the thinking goes: the family has to be chosen so that a distinguishing extension exists.

Worked example 3: a second language, for the pattern

Let L be the set of palindromes over {a, b}, that is, strings equal to their own reversal.

Family: a to the i followed by b, for i at least 1. So ab, aab, aaab, and so on.

Take two, with i less than j. Distinguishing extension: z equal to a to the i. Then a to the i, then b, then a to the i is a palindrome, because it reads the same both ways. But a to the j, then b, then a to the i is not, because it has j of a at the front and i at the back and j is not i.

Infinitely many classes, so the palindromes are not regular. Chapter 56 builds a pushdown automaton for this language, which is the machine that can do it.

Worked example 4: a language that IS regular, proved so

The theorem works in this direction too, which the pumping lemma cannot do.

Let L be the strings over {0, 1} whose value as a binary number is divisible by 3, as in chapter 13's machine 4.

Claim. The index is at most 3, so L is regular.

Proof. If x and y have the same remainder on division by 3 when read as binary numbers, then for any z the numbers xz and yz have the same remainder as each other, because appending a bit b takes a number n to 2n plus b, and that operation depends only on n's remainder. So xz is in L exactly when yz is, and x and y are L-equivalent. There are only three remainders, so there are at most three classes, and the index is finite. By the theorem, L is regular.

And it is exactly 3, because 0, 1 and 10 have remainders 0, 1 and 2 and so lie in three different classes, each pair being distinguished: 0 and 1 by the empty extension, 1 and 10 by the extension 1, since 11 is 3 and is divisible by 3 while 101 is 5 and is not.

Distinctions

Myhill NerodeThe pumping lemma
Formif and only ifone way only
Can prove a language regularyesno
Can prove it not regularyes, always, in principleonly when pumping fails
Tells you the state countyes, it is the indexno
What you must producean infinite family and a distinguishing extensiona string in terms of the pumping length, and a case analysis
Usual exam settingrarercommoner
munotes.in101

The Myhill Nerode Theorem, and a Second Way to Prove a Language Not Regular

States of a machineClasses of the relation
Belong toone particular machinethe language itself
Countcan be any number at least the indexexactly the index
Depend on the constructionyesno

What it does NOT mean

It is not about equal strings. Two L-equivalent strings can look nothing alike: aab and bbbab are in the same class for the ends-in-ab language.

Finite index is not the same as a finite language. The ends-in-ab language is infinite and has index 3.

The classes are not the language and its complement. They are classes of Sigma star under the relation, and there are as many as the minimal machine has states, some of them classes of strings not in L at all.

Producing one distinguishing extension is not enough on its own. You need infinitely many pairwise distinguishable strings. One pair proves two classes, which proves nothing.

The theorem does not give an algorithm for the index of an arbitrary language. It characterises regularity; identifying the classes still takes an argument.

Quick revision

  • x and y are L-equivalent when, for every z, xz is in L exactly when yz is. Such a z that separates them

is a distinguishing extension.

  • It is an equivalence relation on Sigma star, and the number of classes is its index.
  • Myhill Nerode: L is regular if and only if the index is finite, and then the index is the number of states

of the minimal DFA.

  • Half one: strings reaching the same state are L-equivalent, so classes are at most states.
  • Half two: build a machine whose states are the classes, start at the class of epsilon, move from the class

of x on a to the class of xa, and make final the classes inside L.

  • Hence the minimal DFA is unique up to state names, which is what chapter 20 relied on.
  • To prove a language not regular: name an infinite family, and for any two members give one z that

separates them.

  • Unlike the pumping lemma it is an if and only if, so it can also prove a language regular.

Test yourself

1. Define the Myhill Nerode relation, and say what a distinguishing extension is. x is L-equivalent to y when for every string z, xz is in L exactly when yz is in L. A distinguishing extension is a z for which the two disagree, which proves the strings are in different classes.

munotes.in102

The Myhill Nerode Theorem, and a Second Way to Prove a Language Not Regular

2. State the theorem, both halves, and the extra fact about counting. L is regular if and only if the relation has finitely many classes, and when it does, the number of classes equals the number of states of the minimal DFA for L.

3. Prove that the language of strings over {a, b} with equally many a and b is not regular. Take the family a to the i for i at least 0. For i less than j, the extension b to the i is accepted after a to the i and not after a to the j, since the second has j of a and i of b. So there are infinitely many classes and the language is not regular.

4. The ends-in-ab language has index 3. Name the three classes and one extension separating each pair. The class of strings ending in neither a nor ab, of strings ending in a but not ab, and of strings ending in ab. The extension b separates the first from the second; the empty string separates the first from the third and the second from the third.

5. Why can the pumping lemma not prove a language regular, while this theorem can? Because the pumping lemma is a one way implication: regular languages pump, so failing to pump proves irregularity, but pumping successfully proves nothing. Myhill Nerode is an equivalence, so finite index proves regularity outright.

6. A student gives one pair of distinguishable strings and concludes the language is not regular. What is missing? Infinitely many pairwise distinguishable strings. One pair shows only that the index is at least 2, which every non trivial language satisfies.

Contents This chapter on its own page

munotes.in103

Chapter Twenty-Two

What a Grammar Is

Syllabus topic Module 1, "Formal Languages: Defining Grammar"

In one line

A grammar is a set of rewriting rules that starts from one symbol and produces strings, and the strings it can produce are its language.

In the wording a student can write in an examination: a grammar is a quadruple G = (V, T, P, S) where V is a finite set of variables or non terminals, T is a finite set of terminals disjoint from V, P is a finite set of productions each of the form alpha to beta with alpha containing at least one variable, and S in V is the start symbol.

Why a second way of describing a language

Because the two directions answer different questions, and both are asked.

A machine is a recogniser. You hand it a string and it says yes or no. That is what you want when the string already exists: a password to check, a file to validate, an input to reject.

A grammar is a generator. You hand it nothing and it produces strings. That is what you want when the string does not exist yet: a compiler writer defining what a legal program looks like, a language designer specifying syntax, a test tool producing valid inputs.

And the connection between the two is the whole of the rest of Module 1. Chapter 29 sets out the correspondence, chapter 30 proves the smallest case, and chapter 40 assembles it. For now the point is that they are different machinery aimed at the same objects.

The four parts

G = (V, T, P, S)

V, the variables. Also called non terminals. These are the working symbols, the ones that stand for something yet to be filled in. By universal convention they are written as capital letters.

T, the terminals. The symbols that actually appear in the finished strings. These are the alphabet of the language, the Sigma of chapter 3, and by convention they are lower case letters or digits. V and T must have nothing in common: a symbol is either still to be expanded or it is finished, never both.

P, the productions. The rules. Each says that one string of symbols may be replaced by another, and it is written with an arrow:

S->aSb

which reads: wherever you see S, you may write aSb instead. The left hand side must contain at least one variable, because a rule that rewrote only terminals could never be applied to anything worth rewriting.

S, the start symbol. One particular variable, where every derivation begins. By convention it is called S, and where it is not, the grammar says which variable it is.

Reading and writing productions

Three conventions, all of which appear in MU's papers and all of which this book uses.

munotes.in104

What a Grammar Is

The vertical bar means "or". Several rules with the same left hand side may be written on one line:

S->aSb | ab

is two productions, S to aSb and S to ab. It is an abbreviation and nothing more.

Epsilon is a legal right hand side. A rule whose right hand side is the empty string is written with the epsilon symbol, and it means the variable may be deleted:

S->aSb | ε

Such a rule is called a null production or an epsilon production, and chapter 46 is about removing them where possible.

Spaces are optional. MU writes S->0S1 and this book does too. Writing S -> 0 S 1 means exactly the same. Where a variable has a name longer than one letter the spaces stop being optional, and this book then uses them.

The first grammar, worked

Take:

S->aSb | ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

Its four parts. V is {S}. T is {a, b}. P is the two productions. The start symbol is S.

What it produces. Begin with S, the start symbol. Apply the first rule: aSb. Apply it again: aaSbb. Again: aaaSbbb. At any point apply the second rule instead, which deletes the S and leaves a finished string. So the strings are

ε, ab, aabb, aaabbb, aaaabbbb, ...

that is, a to the n followed by b to the n for every n at least 0. And that is the language chapter 21 proved is not regular. So a grammar this small describes something no finite automaton can accept, which is the first sign that grammars will reach further than machines of chapter 9. Module 2 builds the machine that matches.

The vocabulary of a derivation

Three words are needed and are defined properly in chapter 23. They are named here because they appear in every later definition.

A sentential form is any string of variables and terminals that a derivation can reach. S is one, aSb is one, and aabb is one. A sentential form with no variables left in it is called a sentence or a terminal string, and it is finished: no rule can apply to it, because every left hand side contains a variable.

Derives is what one step of rewriting does, and it is written with a single arrow. Derives in any number of steps, including none, is written with a starred arrow, which is the reflexive transitive closure of chapter 4.

The language generated

L(G) = { w : w is a string of terminals and S derives w in any number of steps }

munotes.in105

What a Grammar Is

Two conditions, and students drop the first. The string must consist of terminals only: a derivation that has reached aSb has not produced a member of the language, because S is still there. And it must be reachable from the start symbol, not from some other variable.

Chapter 24 treats this properly, including how to prove that a grammar generates the language you claim.

Grammars are not restricted to context free ones

This matters, because MU's Chomsky classification in chapter 26 needs the general definition and most books quietly give the context free one.

In the general definition, a production is alpha to beta where alpha is any string containing at least one variable, and beta is any string at all. So a rule may have several symbols on the left:

aSb->bSa

is a perfectly legal production of a general grammar. It says that the three symbol sequence aSb may be replaced by bSa, and it can only be applied where all three occur together in that order.

A grammar in which every left hand side is a single variable is called context free, and that is the kind chapter 41 onwards is about. The name says why: a rule like S to aSb may be applied to any S no matter what surrounds it, so the context does not matter. A rule like aSb to bSa is context sensitive: the S may only be rewritten when it has an a before it and a b after it.

Most of this book's grammars are context free. The general form is what makes the four types of chapter 26 four different things.

A worked grammar with two variables

Something less symmetrical, to show the variables working.

S->0A | 1S | ε
A->0S | 1A

V is {S, A}, T is {0, 1}, and the start symbol is S.

What does it do? Read the variables as conditions, exactly as the states of a machine were read in chapter

  1. S means an even number of 0 has been produced so far; A means an odd number. Producing a 0 switches, so S

sends you to A and A sends you back to S; producing a 1 does not, so each variable keeps itself. Only S can finish, because only S has the epsilon rule.

So the language is the strings over {0, 1} with an even number of 0. And here is the machine for the same language, from chapter 10:

A3State01
start finalEOE
OEO

Accepts: ε, 1, 00, 11, 001, 0011

Rejects: 0, 01, 10, 000, 0001

munotes.in106

What a Grammar Is

L(A3) = L(A2)

The checker decides that equality by running the machine and the grammar against each other over every string up to its bound, so the reading of the variables just given is not a claim to be taken on trust.

Notice how closely the grammar copies the machine: one variable per state, one production per transition, and an epsilon rule on the variable for the final state. That correspondence is exactly what chapter 30 formalises, and it is why the smallest class of grammars matches the smallest class of machines.

Distinctions

A machineA grammar
Directionreads a string, answerswrites a string from nothing
Calleda recogniser or acceptora generator
Its partsstates, alphabet, transitions, start, finalsvariables, terminals, productions, start symbol
What "start" meanswhere reading beginsthe symbol a derivation begins from
What finishing meansthe input ran outno variables remain
VariableTerminal
Writtena capital lettera lower case letter or digit
Appears in a finished stringneveryes
Can be rewrittenyesno
Belongs toVT, which is the alphabet of the language
Context free productionGeneral production
Left hand sideexactly one variableany string with at least one variable
ExampleS->aSbaSb->bSa
Class of grammartype 2type 0 or type 1

What it does NOT mean

The arrow is not an equation. S->aSb does not say S equals aSb. It says S may be replaced by aSb, in that direction only, as many times as you like or not at all.

A grammar does not read anything. It has no input. Asking what a grammar accepts is loose talk for what it generates, and the distinction matters when a question asks you to convert one into a machine.

A variable is not a state. The resemblance in the worked example above is real and useful, and it is a theorem rather than a definition. In the grammars of chapter 41 onwards a variable stands for a whole sub structure, not for a condition.

The empty string is not "nothing". A rule to epsilon deletes a variable and leaves a legitimate string, possibly empty. Chapter 3 said the same thing about strings.

V and T must be disjoint. A symbol that is both a variable and a terminal makes the definition of a finished string meaningless.

Quick revision

  • A grammar is G = (V, T, P, S): variables, terminals, productions, start symbol. V and T are disjoint.
  • A production alpha to beta may be applied wherever alpha occurs; the left hand side must contain at least

one variable.

  • The vertical bar abbreviates several productions with the same left hand side.
  • A right hand side may be epsilon, which deletes the variable. That is a null production.
  • A sentential form is anything a derivation can reach; one with no variables is a finished terminal string.
  • L(G) is the set of terminal strings derivable from S in any number of steps. Both conditions matter.
  • Context free means every left hand side is a single variable. The general definition allows more, and
munotes.in107

What a Grammar Is

chapter 26 needs it.

  • A grammar generates and a machine recognises. S->aSb | ε generates a language no finite automaton can

accept.

Test yourself

1. Give the four parts of a grammar and the condition on a production. V the variables, T the terminals with V and T disjoint, P the productions, and S in V the start symbol. Each production is alpha to beta with alpha containing at least one variable.

2. What language does S->aS | a generate? One or more a. Every derivation applies the first rule some number of times and then must finish with the second, so at least one a is produced and there is no upper limit.

3. What is the difference between a sentential form and a sentence? A sentential form may still contain variables; a sentence contains only terminals and is therefore finished, because every production's left hand side holds a variable.

4. Why must the two conditions in the definition of L(G) both be stated? Because a derivation from S may stop at a form that still contains variables, and that form is not in the language; and because a string derivable from some other variable need not be derivable from S.

5. Is aSb->bSa a legal production? What does allowing it mean? Legal in a general grammar, because its left hand side contains a variable. Allowing it means the grammar is not context free: the S may be rewritten only when an a precedes and a b follows it.

6. Write a grammar for the strings over {a, b} that begin with a. S->aA, A->aA | bA | ε. The first rule forces the initial a and A then generates anything at all, including the empty continuation.

Contents This chapter on its own page

munotes.in108

Chapter Twenty-Three

Derivations: How a Grammar Makes a String

Syllabus topic Module 1, "Formal Languages: Derivations"

In one line

A derivation is the list of sentential forms you pass through on the way from the start symbol to a finished string, each obtained from the one before by applying one production once.

In the wording a student can write in an examination: if alpha to beta is a production of G and gamma alpha delta is a sentential form, then gamma alpha delta directly derives gamma beta delta, written with a single arrow. The relation derives in zero or more steps, written with a starred arrow, is the reflexive transitive closure of that relation, and a derivation of w from S is a finite sequence of sentential forms beginning with S and ending with w in which each step is a direct derivation.

Why it needs its own chapter

Because the definition of the language in chapter 24 is a statement about derivations, and because setting a derivation out properly is what earns the marks when MU asks for one.

Her papers set it both ways. They give a grammar and ask what it generates, which needs derivations to be read. And they give a string and a grammar and ask for a derivation tree, which needs derivations to be written. Both are mechanical once the definition is clear, and both go wrong in the same place: applying a production to the wrong occurrence, or applying two at once.

The one step relation

Take a sentential form and find in it the left hand side of some production. Replace that occurrence, and only that occurrence, with the production's right hand side. What you get is the next form, and the step is written with a single arrow.

Three points, each of which is a rule an examiner checks.

One occurrence per step. If the form holds three copies of S, one step rewrites one of them. Rewriting two is two steps, and writing them as one loses a mark and makes the derivation unverifiable.

One production per step. A step applies a single rule, not a chain of them.

Everything else is untouched. The symbols to the left and right of the occurrence are copied over unchanged. That is what the gamma and delta in the formal statement are doing: they are the surroundings.

The zero or more steps relation

Written with a starred arrow. It holds between two forms when there is a chain of single steps from the first to the second, of any length, including zero.

The zero case is why the star is there and why the relation is called reflexive. Every sentential form derives itself in zero steps. It looks like a technicality and it is needed in the base case of every induction in the next chapter, exactly as delta hat of the empty string was needed in chapter 11.

munotes.in109

Derivations: How a Grammar Makes a String

This is the reflexive transitive closure of chapter 4, and the notation is the same.

How to set a derivation out

One form per line, starting with the start symbol and ending with the finished string. Each line is one step. It is worth writing which production was used at each step where the grammar has several, and any examiner will accept it and most will expect it.

Take the grammar of chapter 22:

S->aSb | ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abab

A derivation of aabb:

S
aSb
aaSbb
aabb

Four lines, so three steps: the first rule twice and then the second rule to delete the S. That is the whole derivation, and the checker re-derives it, so if the second line read aSSb the book would not build.

A grammar with a choice at every step

Here is one where the derivations are more interesting, because there is more than one variable to rewrite.

S->AB
A->aA | a
B->bB | b

Accepts: ab, aab, abb, aabb, aaabbb

Rejects: ε, a, b, ba, bab

This grammar makes one or more a followed by one or more b. Note the difference from B1: there is no matching here, because A and B grow independently, so aab and abb are both in the language while neither is in B1's.

A derivation of aabb, rewriting the leftmost variable every time:

S
AB
aAB
aaB
aabB
aabb

Six lines, five steps. Look at line three: the form is aAB, there are two variables in it, and the leftmost is A, so the leftmost discipline rewrites A. The checker proves not only that each step is legal but that the leftmost occurrence was the one chosen, so a derivation labelled leftmost that is not one fails the build.

And the same string derived rewriting the rightmost variable every time:

S
AB
AbB
Abb
aAbb
aabb

The same six lines in a different order, and the same string at the end. Chapter 43 is about what these two have in common, which turns out to be the derivation tree.

The trap: one occurrence at a time

Take a grammar with a variable appearing twice on a right hand side.

S->SbS | a

Accepts: a, aba, ababa, abababa

Rejects: ε, b, ab, ba, abab, aa

A correct derivation of ababa:

S
SbS
abS
abSbS
ababS
ababa

Six lines, five steps. The second line has two copies of S. Step two rewrites the leftmost of them, and the other is copied over untouched. A student who rewrites both in one line has written something that is not a derivation, and the checker refuses it.

munotes.in110

Derivations: How a Grammar Makes a String

Note also that this grammar is ambiguous, which chapter 44 is about: ababa has more than one derivation tree, and the checker can produce two of them.

B3 is ambiguous

A general grammar, where the left hand side is longer

Chapter 22 said a production's left hand side may be more than one symbol, and a derivation then has one extra thing to look for: the occurrence must match the whole left hand side in order.

S->aSBc | abc
cB->Bc
bB->bb

Accepts: abc, aabbcc, aaabbbccc

Rejects: ε, a, ab, abcc, aabc, aabbc

This is a context sensitive grammar, and it generates a to the n, b to the n, c to the n. A derivation of aabbcc:

S
aSBc
aabcBc
aabBcc
aabbcc

Five lines, four steps. Step one applies the first rule. Step two applies S to abc inside the form aSBc, leaving the B and the c where they were. Step three applies cB to Bc, which needs a c immediately followed by a B, and the form aabcBc has one. Step four applies bB to bb, which needs a b immediately followed by a B.

The B is doing the work of a courier. It is created on the right of the S, and each application of the first rule creates one more, so the number of B always equals the number of a. Then cB to Bc lets a B walk left past a c, and bB to bb turns it into a b when it reaches the b block. So the count of b ends up equal to the count of a, and the count of c was equal from the start. That is the standard way a context sensitive grammar counts three things at once, and it is worth seeing once, because chapter 26 classifies grammars by exactly this sort of rule and chapter 60 builds the machine for them.

Reading a grammar you are given

MU's 2018 paper sets this directly: "If G = ({S}, {0, 1}, {S->0S1, S->epsilon}, S), find L(G)".

The method is to derive the shortest few strings and look for the pattern, then check the pattern by argument rather than by more examples.

S->0S1 | ε

Accepts: ε, 01, 0011, 000111

Rejects: 0, 1, 10, 001, 011, 0101

Shortest first. Applying the second rule immediately gives epsilon. Applying the first rule once and then the second gives 01. Twice then the second gives 0011. So the pattern is 0 to the n followed by 1 to the n.

munotes.in111

Derivations: How a Grammar Makes a String

And the argument: the only way to finish is the epsilon rule, which must be applied to the single S, and every application of the first rule adds exactly one 0 on the left of the S and one 1 on the right. So the counts are equal by construction, and the 0 all precede the 1 because nothing ever moves a symbol. Chapter 24 sets out how to write that argument as a proof.

Distinctions

One stepZero or more steps
Writtena single arrowa starred arrow
Number of productions appliedexactly oneany number, including none
Relates a form to itselfnoyes, in zero steps
Chapter 4's name for itthe relationits reflexive transitive closure
Sentential formSentence
May contain variablesyesno
Can be rewritten furtheryesno
In the languagenot unless it is a sentenceyes, if derived from S
Leftmost derivationRightmost derivation
Which occurrence is rewrittenthe leftmost variablethe rightmost variable
Number for a given treeexactly oneexactly one
Used fordefining ambiguity, parsing top downparsing bottom up

What it does NOT mean

A derivation is not a proof that the grammar is correct. It shows one string is in the language. Chapter 24 is about the other direction.

A step does not rewrite every occurrence. One occurrence, one production, everything else copied.

The starred arrow is not "many steps". It is zero or more, and the zero case is used in every induction.

A derivation does not have to be leftmost. Any order of rewriting is a derivation. Leftmost and rightmost are disciplines imposed for a purpose, which chapter 42 explains.

Two derivations of the same string are not automatically a sign of ambiguity. The grammar B2 above has a leftmost and a rightmost derivation of aabb and is not ambiguous. Ambiguity is about two derivation trees, which chapter 44 defines precisely.

Quick revision

  • One step: find a production's left hand side in the form, replace that one occurrence by the right hand

side, copy everything else. Written with a single arrow.

  • Zero or more steps: written with a starred arrow, the reflexive transitive closure, so every form derives

itself.

  • A derivation is written one form per line, from S to a terminal string, one step per line.
  • Leftmost rewrites the leftmost variable at every step; rightmost the rightmost. Each tree has exactly one

of each.

  • A variable appearing twice in a form is rewritten one copy at a time.
  • In a general grammar the whole left hand side must occur, in order, for the rule to apply.
  • To find what a grammar generates: derive the shortest strings, spot the pattern, then argue it.
munotes.in112

Derivations: How a Grammar Makes a String

Test yourself

1. Give the definition of one step of derivation. If alpha to beta is a production and the current form is gamma alpha delta, then the next form may be gamma beta delta: one occurrence of one left hand side is replaced, and the surroundings are unchanged.

2. Why is the starred arrow reflexive? Because it allows zero steps, and a chain of zero steps leads from any form to itself. That case is what the base case of every induction over derivations uses.

3. For S->SbS | a, give a leftmost derivation of aba. S, then SbS, then abS, then aba. Three steps: the first rule, then the second applied to the leftmost S, then the second applied to the remaining S.

4. A student writes S, then SbS, then aba in two steps. What is wrong? The second step rewrites both copies of S at once, which is two applications of a production and therefore two steps. A derivation must show each one.

5. In the grammar S->aSBc | abc, cB->Bc, bB->bb, why can cB->Bc not be applied to the form aSBc? Because the rule needs a c immediately followed by a B, and in aSBc the B is preceded by S and followed by c. The whole left hand side must occur, in order, for a production to apply.

6. What is the difference between a derivation and a derivation tree? A derivation is an ordered sequence of rewriting steps, so the same tree gives many derivations differing only in the order the variables were taken. A tree records which production was used on which symbol and forgets the order, which is why ambiguity is defined on trees and not on derivations.

Contents This chapter on its own page

munotes.in113

Chapter Twenty-Four

The Language Generated by a Grammar, and Proving It Is the One You Claim

Syllabus topic Module 1, "Formal Languages: Languges generated by Grammar"

In one line

The language of a grammar is every terminal string the start symbol can derive, and proving a grammar correct means proving two containments, not showing examples.

In the wording a student can write in an examination: the language generated by G, written L(G), is the set of all strings w over T such that S derives w in zero or more steps. To prove L(G) equals a claimed set A, one proves that every string derivable from S lies in A, and that every string of A is derivable from S.

Why two containments

Because "the same set" means two things, and chapter 2 already said so: two sets are equal when each is a subset of the other.

The two halves are genuinely different pieces of work and they fail in different ways.

Everything the grammar makes is in A. This is the half that catches a grammar that is too generous: one that also produces strings you did not want. It is proved by induction on the derivation, and the usual method is to find a property that every sentential form has and that forces the conclusion once the variables are gone.

Everything in A is made by the grammar. This is the half that catches a grammar that is too mean: one that misses strings you did want. It is proved by taking an arbitrary member of A and constructing a derivation of it, usually by induction on its length.

A grammar can fail either way, and a table of examples finds neither failure reliably. A generous grammar passes every example you thought to try, because your examples are the strings you wanted. A mean grammar fails on the first example you did not think of.

Finding L(G) when you are given G

This is the examination question, and the method has three steps.

Step 1. Derive the shortest strings. Start with the productions that finish, the ones whose right hand sides hold no variables, and work up. Get four or five strings.

Step 2. Say what the pattern is, in words. Not "strings like 01, 0011" but a description: n zeros followed by n ones, for n at least 0.

Step 3. Check the pattern against the productions, not against more examples. Look at each production and ask what it contributes. A rule that adds a symbol on each side of a variable keeps something balanced. A rule that adds only on the left does not.

Worked, on MU's own grammar

S->0S1 | ε

Accepts: ε, 01, 0011, 000111

Rejects: 0, 1, 10, 001, 011, 0101

Step 1. The finishing rule is the epsilon one, so the shortest string is epsilon. One application of the first rule then the second gives 01. Two then the second gives 0011.

munotes.in114

The Language Generated by a Grammar, and Proving It Is the One You Claim

S
0S1
00S11
0011

Step 2. The pattern is 0 to the n followed by 1 to the n, for n at least 0.

Step 3. The grammar has one variable and two rules. The only way to finish is the epsilon rule, and it applies to the single S. Every application of the first rule adds exactly one 0 immediately to the left of the S and one 1 immediately to the right. So at every stage the form is 0 to the k, then S, then 1 to the k, and nothing can disturb that because there is only ever one S and nothing moves symbols. When the epsilon rule is used the S vanishes and 0 to the k followed by 1 to the k is left.

So:

L(C1) = { 0 to the n then 1 to the n : n at least 0 }

That step 3 argument is the first containment, done informally. Below it is done properly.

The proof, both halves, on C1

Write A for the claimed set, the strings 0 to the n followed by 1 to the n for n at least 0.

Half one: L(C1) is contained in A

The technique is to find an invariant: a property every sentential form reachable from S has.

Claim. Every sentential form derivable from S is either 0 to the k, S, 1 to the k for some k at least 0, or 0 to the k followed by 1 to the k for some k at least 0.

Proof by induction on the number of steps.

Base case, zero steps. The form is S, which is 0 to the 0, S, 1 to the 0. True.

Inductive step. Suppose the form after i steps has the claimed shape. If it has no S in it, no production applies, and there is no i plus 1 step to consider. If it is 0 to the k, S, 1 to the k, then exactly one production can be applied, to the S, and there are two choices:

  • The first rule replaces S by 0S1, giving 0 to the (k plus 1), S, 1 to the (k plus 1). Still the first

shape.

  • The epsilon rule deletes S, giving 0 to the k followed by 1 to the k. The second shape.

Either way the claim holds after i plus 1 steps.

Finishing. A string of L(C1) is a terminal string, so it has no S, so by the claim it is 0 to the k followed by 1 to the k for some k. That is a member of A.

munotes.in115

The Language Generated by a Grammar, and Proving It Is the One You Claim

Half two: A is contained in L(C1)

Proof by induction on n.

Base case, n equal to 0. The string is epsilon, and S derives epsilon in one step by the second rule.

Inductive step. Suppose 0 to the n followed by 1 to the n is derivable from S, by some derivation. Then

S derives 0 S 1 in one step, and S derives 0 to the n then 1 to the n

so putting a 0 in front and a 1 behind every form of that derivation gives a derivation of 0 to the (n plus

  1. followed by 1 to the (n plus 1). This works because a derivation step is unaffected by what surrounds the

occurrence rewritten, which is exactly the gamma and delta of chapter 23's definition.

Therefore every member of A is derivable, and with half one, L(C1) equals A.

That is what a complete answer looks like. It is two inductions, each about half a page, and the pattern is the same in every such proof: half one finds an invariant over forms, half two builds a derivation from the string.

A second worked pair: a grammar that is too generous

Here is the failure mode half one catches, made concrete.

Suppose a student is asked for a grammar for the strings over {a, b} with equally many a and b, and writes:

S->aSb | bSa | ε

Accepts: ε, ab, ba, abab, aabb

Rejects: a, b, aab, abb, baab, abba, aabbb

Is it right? Every rule adds one a and one b, so half one goes through: every string it makes has equally many of each. So the grammar is not too generous.

But it is too mean, and the claim lists above already show it: baab and abba both have two a and two b, so both belong in the language, and this grammar rejects both.

Take baab and see why. It begins with b, so its derivation must begin with the rule b S a, giving the form b S a with the final a already placed. The middle must then derive aa. But every rule of this grammar adds one a and one b, so S can only ever derive a string with equal counts, and aa has two a and no b. The same argument kills abba, whose first rule must be a S b, leaving the middle to derive bb.

S->aSb | bSa | SS | ε

Accepts: ε, ab, ba, baab, abba, abab, aabb, aabbab

munotes.in116

The Language Generated by a Grammar, and Proving It Is the One You Claim

Rejects: a, b, aab, abb, aba, aabbb

Adding the rule S->SS repairs it, because it allows a balanced string to be split into two balanced pieces placed side by side: baab is ba followed by ab, and abba is ab followed by ba. The checker runs both grammars, so the difference between them is demonstrated rather than asserted: C2 rejects baab and abba, and C3 accepts both.

C3 is ambiguous

And the repaired grammar is now ambiguous, which chapter 44 will explain is the price of the SS rule.

The other direction: a grammar that is too generous

S->aS | Sb | ε

Accepts: ε, a, b, ab, aab, abb, aabb

Rejects: ba, bab, aba

A student might offer this for "n a followed by n b". It is not: it accepts a, which has one a and no b, and it accepts aab. The rules add a on the left and b on the right independently, with nothing forcing the counts to agree. Half one of the proof fails immediately, because the invariant "the counts are equal" does not survive the first rule.

Its actual language is a star followed by b star, which is a perfectly good regular language and not the one asked for. That is why the first half of the proof is worth doing: the grammar looks plausible and is wrong.

Distinctions

Too generousToo mean
Faultderives strings outside Afails to derive some string of A
Caught byhalf one, the invarianthalf two, the construction
Symptom on examplespasses every example you triedfails on one you did not think of
Example aboveC4 for balanced countsC2 for equal counts
Finding L(G)Proving L(G) equals A
Giventhe grammarthe grammar and a claim
Methodshortest derivations, then a pattern, then check the rulestwo inductions
Length of answera paragraphabout a page
MU sets itoftenrarely, but it is what makes the rest safe

What it does NOT mean

A list of derivable strings is not L(G). L(G) is usually infinite, and a list is evidence about a pattern rather than the answer.

A sentential form is not in L(G) unless it is terminal. 0S1 is derivable from S and is not in the language.

A string derivable from some variable is not in L(G) unless it is derivable from S. In C3 above, strings derivable from S happen to be all of them because there is one variable, but in a grammar with several variables this catches people.

Half one alone is not a proof. Nor is half two. A grammar can pass either and fail the other, and the two worked failures above are one of each.

munotes.in117

The Language Generated by a Grammar, and Proving It Is the One You Claim

Examples are not a proof. They are how you find the pattern, which is step 1 of a different task.

Quick revision

  • L(G) is every terminal string derivable from S in zero or more steps. Both conditions matter.
  • To find L(G): derive the shortest strings, state the pattern in words, then justify it from the

productions.

  • To prove L(G) equals A: two containments. L(G) inside A by an invariant over sentential forms, proved by

induction on the number of steps. A inside L(G) by constructing a derivation of an arbitrary member, usually by induction on its length.

  • A grammar that is too generous fails half one; a grammar that is too mean fails half two.
  • A derivation step is unaffected by its surroundings, which is what lets half two wrap a derivation in extra

symbols.

Test yourself

1. Give the definition of L(G). The set of strings w over the terminals such that S derives w in zero or more steps. The string must be terminal and must come from S.

2. What are the two halves of a correctness proof, and what does each catch? That everything derivable is in the claimed set, which catches a grammar producing too much; and that everything in the claimed set is derivable, which catches a grammar producing too little.

3. Find L(G) for S->aSa | bSb | a | b. The odd length palindromes over {a, b}. Each of the first two rules adds the same symbol at both ends, and the only ways to finish leave a single a or b in the middle, so the string reads the same both ways and has odd length.

4. Is S->aS | Sb | ε a grammar for n a followed by n b? Why not? No. The two rules act independently, so nothing forces the counts to agree, and it derives a, b and aab. Its language is a star followed by b star.

5. Why is S->aSb | bSa | ε not a grammar for equal counts of a and b? Because it cannot derive baab, or abba. Every derivation places a matching pair at the two ends, so the middle must itself have equal counts, and the middle of baab is aa. Adding S->SS repairs it, since baab splits as ba followed by ab.

6. In half two of the proof for S->0S1 | ε, what justifies wrapping a whole derivation in a 0 and a 1? That a derivation step depends only on the occurrence rewritten and copies its surroundings unchanged, so adding the same symbols to both ends of every form in a derivation leaves every step legal.

Contents This chapter on its own page

munotes.in118

Chapter Twenty-Five

Writing a Grammar for a Language You Are Given

Syllabus topic Module 1, "Formal Languages: Languges generated by Grammar"

In one line

To write a grammar, decide what each variable will stand for, then write one production for each way a string of that kind can be built.

In the wording a student can write in an examination: a grammar for a language L is constructed by identifying the recursive structure of the strings of L, assigning a variable to each kind of substring that recurs, writing productions that express how a longer string of each kind is built from shorter ones, and providing at least one production for each variable whose right hand side contains no variable, so that every derivation can terminate.

Why a method rather than practice

Because the languages in this paper fall into four or five shapes, and once the shape is recognised the grammar writes itself. Guessing works on a star b star and stops working the first time two counts have to match.

The single most useful habit is the one chapter 13 recommended for machines: write down what each variable means before writing any production. A variable with no meaning is a production you cannot check, and it is where wrong grammars come from.

The method, in four steps

Step 1. Say what the strings look like. In words, precisely. Not "strings with a and b" but "any number of a, then the same number of b".

Step 2. Find the recursion. Ask: given a string of the language, how do I get a longer one? And what are the shortest ones? The answers become the productions directly.

Step 3. Give each kind of substring a variable, and write its meaning down. If two parts of the string are independent, they get separate variables.

Step 4. Make sure every derivation can stop. Every variable needs at least one production whose right hand side has no variables in it, or the grammar derives nothing at all. This is the step people forget, and it makes the language empty rather than wrong, which is worse because it looks fine.

The four shapes

Almost every exercise is one of these.

ShapeRecursionPattern of rules
any number of somethingadd one at either endA->xA or A->Ax, and A->ε
two counts that must matchadd one at EACH end at onceS->xSy, and S->ε
something then something else, independentlyone variable eachS->AB
a choice between formsone rule per formS->form1 given as alternatives

The second row is the one a finite automaton cannot do, and the reason is chapter 21's: matching needs memory that grows.

Grammar 1: any number of a, then any number of b

Step 1. Zero or more a followed by zero or more b. So epsilon, a, b, ab, aab, abb are in and ba is not.

munotes.in119

Writing a Grammar for a Language You Are Given

Step 2. The two halves are independent: adding an a does not require adding a b.

Step 3. Two variables. A produces the a block, B produces the b block.

S->AB
A->aA | ε
B->bB | ε

Accepts: ε, a, b, ab, aab, abb, aabbb

Rejects: ba, aba, bab, abab

This language is regular, so there is a machine, and the two can be compared exactly:

L(D1) = L(D1R)

a* b*

Grammar 2: n a followed by n b

Step 1. Equal numbers, all a before all b.

Step 2. The recursion is the difference from grammar 1: a longer string is obtained by adding an a at the front and a b at the back, together. Shortest string: epsilon.

Step 3. One variable does, because the two additions happen in one rule.

S->aSb | ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

Compare D1 and D2 as written: one extra variable in D1 and one fewer in D2, and that is the whole difference between a regular language and one that is not.

Grammar 3: equal numbers of a and b, in any order

Step 1. The counts agree; the order does not matter. So ab, ba, abba, baab, aabb are all in.

Step 2. Two recursions, and this is the case chapter 24 got wrong first. A longer string is obtained either by wrapping a shorter one in a matching pair, which gives aSb and bSa, or by placing two shorter strings of the language side by side, which gives SS. The third is not optional: without it baab cannot be made.

S->aSb | bSa | SS | ε

Accepts: ε, ab, ba, baab, abba, aabb, abab, aabbab

Rejects: a, b, aab, abb, aba, aaabb

D3 is ambiguous

The grammar is ambiguous, and that is a genuine cost of the SS rule which chapter 44 discusses. For a grammar that is asked to generate a language and nothing more, ambiguity is permitted.

Grammar 4: palindromes over {a, b}

Step 1. A string equal to its own reversal.

Step 2. A longer palindrome is a shorter one with the same symbol added at both ends. The shortest ones are epsilon, a and b, and all three are needed: epsilon and the two single symbols, because a palindrome may have odd or even length.

S->aSa | bSb | a | b | ε

Accepts: ε, a, b, aa, aba, abba, ababa

munotes.in120

Writing a Grammar for a Language You Are Given

Rejects: ab, ba, aab, abb, aabb

The three terminating rules are what makes this right. Dropping epsilon loses the even length palindromes; dropping a and b loses the odd length ones; and a student who writes only S->aSa | bSb | ε generates only the even ones.

Grammar 5: strings over {0, 1} with an odd number of 1

Step 1. The count of 1 is odd; the 0 are unrestricted.

Step 2 and 3. This is a counting language, so read it as a machine and copy the machine: one variable per condition. E means an even number of 1 produced so far, O means odd. Producing a 0 does not switch; producing a 1 does. Only O can finish, because only an odd count is wanted.

S->E
E->0E | 1O
O->0O | 1E | ε

Accepts: 1, 01, 10, 111, 0100, 1011

Rejects: ε, 0, 00, 11, 0101, 1111

This language is regular, and the machine for it is:

D6State01
startevenevenodd
finaloddoddeven

Accepts: 1, 01, 10, 111, 0100, 1011

Rejects: ε, 0, 00, 11, 0101, 1111

L(D5) = L(D6)

Two variables, two states, and the productions are the transitions. That correspondence is the subject of chapter 30, and it means that any language you can build a machine for you can build a grammar for mechanically, by copying the table.

Grammar 6: strings over {a, b} containing the substring abb

Step 1. Anything, then abb, then anything.

Step 2 and 3. Three parts, and the middle is fixed. A variable for "anything" serves both ends.

S->XabbX
X->aX | bX | ε

Accepts: abb, aabb, abba, babba, abbabb

Rejects: ε, ab, ba, aab, bab, abab

L(D7) = L(D7R)

(a + b)* a b b (a + b)*

Note how the fixed middle is simply written out as terminals inside a production. There is no need for a variable for something that does not vary.

Grammar 7: a to the n b to the n c to the m, with n and m independent

Step 1. A matched a and b block, then an unrelated number of c.

Step 2 and 3. Two independent parts, so a variable each, and the first part is grammar 2.

S->AC
A->aAb | ε
C->cC | ε

Accepts: ε, ab, c, abc, aabbc, abcc, aabbccc

Rejects: a, b, ba, ac, abbc, cab

The lesson of this one is the word independent in step 3. Putting the c inside the recursive rule, as in A->aAbc, would tie the count of c to the count of a, which is a different language and one no machine in this book's Module 1 or Module 2 can accept without a linear bounded automaton.

munotes.in121

Writing a Grammar for a Language You Are Given

Grammar 8: the strings over {a, b} of even length

Step 1. Length divisible by 2.

Step 2. Add two symbols at a time. There are four ways to add two symbols, and all four are needed.

S->aaS | abS | baS | bbS | ε

Accepts: ε, aa, ab, ba, bb, aabb, abab

Rejects: a, b, aab, aba, abb, ababa

L(D9) = L(D9R)

((a + b)(a + b))*

A neater grammar for the same language is S->TTS | ε with T->a | b, which has two variables and two rules and says the same thing. Both are correct; the second scales better if the alphabet grows.

The three ways a grammar goes wrong

Each of these is worth recognising by its symptom.

It derives nothing at all. Some variable has no terminating production, so every derivation from S gets stuck with a variable it cannot remove. Symptom: the language is empty, and everything you try fails. Check step 4.

It derives too much. A rule adds symbols independently where they should be tied. Symptom: an example you did not intend is accepted. Chapter 24's C4 is this.

It derives too little. A way of building a string of the language has no rule. Symptom: an example you did intend is rejected. Chapter 24's C2 is this, and the SS rule is what repairs it.

Distinctions

S->aSbS->aS, S->Sb
Addsone a and one b in the same stepone at a time, independently
Ties the countsyesno
Language, with an epsilon rulen a then n ba star b star
Regularnoyes
One variableTwo variables
Use whenone thing recurstwo parts vary independently
ExampleS->aSb for matched countsS->AB for a star b star

What it does NOT mean

More variables is not safer. Each one needs a meaning and a terminating rule, and an extra variable is an extra thing to get wrong.

A grammar need not be unambiguous. Unless the question asks for one, ambiguity is allowed. Chapter 44 is where it matters.

A grammar for a regular language need not be regular in form. D9 above is context free in form though its language is regular. Chapter 30 is about when the form can be restricted.

Writing down the recursion is not the same as writing down examples. Examples are step 1 evidence. The recursion is the answer.

munotes.in122

Writing a Grammar for a Language You Are Given

Quick revision

  • Four steps: say what the strings look like; find how a longer one is built from a shorter one and what the

shortest are; give each recurring part a variable with its meaning written down; and give every variable a terminating production.

  • Four shapes: any number of something (A->xA, A->ε); two counts that match (S->xSy, S->ε); independent parts

(S->AB); a choice of forms (alternatives).

  • Matched counts need the additions in ONE rule. Independent parts need separate variables.
  • Equal counts in any order need three recursive rules, including S->SS, or baab is unreachable.
  • Palindromes need all three terminating rules: epsilon for even length, a and b for odd.
  • A counting language can be written by copying a machine: one variable per state, one production per

transition, an epsilon rule on each final state.

  • Three failure modes: derives nothing (a missing terminating rule), too much (counts not tied), too little (a

missing way of building).

Test yourself

1. Write a grammar for the strings over {a, b} that begin and end with the same symbol, of length at least two. S->aAa | bAb, A->aA | bA | ε. The outer rule fixes the two ends to agree and A supplies anything in between, including nothing, so the shortest strings are aa and bb.

2. Write a grammar for a to the n b to the 2n. S->aSbb | ε. Each step adds one a and two b, so the b count is always twice the a count.

3. Why does S->aSb | bSa | ε fail for equal counts of a and b, and what repairs it? Because every derivation wraps a matching pair around a shorter string of the language, so the middle must itself have equal counts; baab would need aa in the middle. Adding S->SS repairs it, since baab splits into ba and ab.

4. Write a grammar for the strings over {a, b} of odd length. S->TTS | T with T->a | b. The T at the end supplies the odd symbol and TT adds symbols in pairs. Equivalently S->aS' | bS' with the pair rule written out, but the two variable form is shorter.

5. What goes wrong with A->aAb, with no other rule for A? Nothing can terminate, because every production of A contains A, so no derivation ever removes the variable and the language is empty.

6. Give a grammar for the strings over {0, 1} whose length is a multiple of 3. S->TTTS | ε with T->0 | 1. Three symbols are added at a time and the epsilon rule finishes, so every string has length divisible by 3.

Contents This chapter on its own page

munotes.in123

Chapter Twenty-Six

The Chomsky Classification of Grammars and Languages

Syllabus topic Module 1, "Formal Languages: Chomsky Classification of Grammar and Languages"

In one line

Grammars fall into four types according to how much the shape of a production is restricted, and each type matches one of the four machines of this book.

In the wording a student can write in an examination: the Chomsky hierarchy classifies phrase structure grammars into four types. Type 0 or unrestricted grammars place no restriction on productions. Type 1 or context sensitive grammars require that no production shortens the string. Type 2 or context free grammars require every left hand side to be a single variable. Type 3 or regular grammars additionally restrict every right hand side to a terminal, or a terminal followed by a single variable. Each type generates a strictly larger class of languages than the next, and each corresponds to a machine model.

Why classify at all

Because "grammar" as defined in chapter 22 is too permissive to be useful on its own. A type 0 grammar can describe anything a computer can compute, which means nothing can be proved about it in general. Restricting the shape of a production buys results: the more restricted the grammar, the more you can decide about the language it generates.

That is the trade and it is the whole point of the hierarchy. Going down the list, each step gives up expressive power and gains decidability. A regular language's emptiness, finiteness and equality are all decidable, as chapter 39 shows. For a context free language, emptiness and membership are decidable and equality is not. For a type 0 language, even membership is undecidable, which chapter 74 proves.

The four types, each as a rule about productions

Write a production as alpha to beta.

Type 0: unrestricted, also called phrase structure

The rule. alpha must contain at least one variable. Nothing else is required. beta may be longer, shorter or empty.

That is the general definition of chapter 22 with no extra restriction, which is why MU names it "phrase structure grammar" in one paper and "type 0" in another; they are the same thing.

Machine: the Turing machine, chapter 62. Languages: the recursively enumerable languages, chapter 27.

Type 1: context sensitive

The rule. Every production alpha to beta has beta at least as long as alpha. Productions are non shortening.

One exception is universally allowed: S to epsilon, where S is the start symbol and S appears on no right hand side. Without it the empty string could never be generated by any type 1 grammar, which is an accident of the definition rather than a fact about languages.

The name comes from an equivalent way of stating the rule: every production may be written in the form alpha A beta to alpha gamma beta, with A a single variable and gamma non empty. Read that way, it says a variable may be rewritten in the context of what stands around it. The two formulations generate the same class, and the non shortening one is easier to check.

munotes.in124

The Chomsky Classification of Grammars and Languages

Machine: the linear bounded automaton, chapter 60. Languages: the context sensitive languages.

S->aSBc | abc
cB->Bc
bB->bb

Accepts: abc, aabbcc, aaabbbccc

Rejects: ε, a, ab, abcc, aabc, aabbc

Check the rule on each production. S to aSBc: one symbol becomes four, so not shortening. S to abc: one becomes three. cB to Bc: two become two. bB to bb: two become two. So E1 is type 1. And it generates a to the n, b to the n, c to the n, which chapter 50 proves is not context free, so this grammar could not have been written as a type 2 one.

Type 2: context free

The rule. Every left hand side is exactly one variable. The right hand side is anything at all, including epsilon.

The name says why: the rule A to gamma applies to any occurrence of A, whatever is around it. The context is irrelevant.

Machine: the pushdown automaton, chapter 52. Languages: the context free languages.

S->aSb | ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abab

One variable on the left of each production, so type 2.

Note that a type 2 grammar with an epsilon rule on a variable other than S is not type 1, because that production shortens. So type 2 is not literally a subset of type 1 as stated; the containment is between the classes of languages, and every context free language does have a type 1 grammar once the epsilon productions are removed by chapter 46's method. That subtlety is worth a sentence in an answer.

Type 3: regular, also called right linear

The rule. Every left hand side is a single variable, and every right hand side is one of:

  • a single terminal;
  • a single terminal followed by a single variable;
  • epsilon.

So A to aB is allowed, A to a is allowed, A to epsilon is allowed, and A to aBc, A to Ba and A to AB are not.

Machine: the finite automaton, chapter 9. Languages: the regular languages.

S->aS | bS | aA
A->bB
B->aB | bB | ε

Accepts: ab, aba, abab, aabab, babb

Rejects: ε, a, b, ba, aa, bb

Every right hand side here is a terminal, or a terminal then one variable, or epsilon. So E3 is type 3. Read the variables as a machine's states: S means no ab has been completed yet, A means an a has just been read and is waiting for its b, and B means ab has been seen and anything may follow. Its language is regular:

munotes.in125

The Chomsky Classification of Grammars and Languages

L(E3) = L(E3R)

(a + b)* a b (a + b)*

Chapter 30 treats the regular grammar properly, including the left linear form and why mixing the two forms escapes the class.

The four types on one page

TypeNameThe restriction on alpha to betaMachineLanguages
0unrestricted, phrase structurealpha holds a variableTuring machinerecursively enumerable
1context sensitivebeta at least as long as alphalinear bounded automatoncontext sensitive
2context freealpha is one variablepushdown automatoncontext free
3regular, right linearalpha is one variable; beta is a terminal, a terminal and a variable, or epsilonfinite automatonregular

Reading the table downward the restriction tightens, the machine weakens, and the class of languages shrinks. Reading it upward the opposite. A student who can reproduce this table has most of what MU asks on this topic.

The containments, and that they are strict

regular inside context free inside context sensitive inside recursively enumerable

Each containment is proper: there is a language in each class that is not in the one below. That is what makes the hierarchy four things rather than one, and this book proves every step.

StepA language on the larger side and not the smallerProved in
regular inside context freea to the n b to the nchapter 21 and chapter 36
context free inside context sensitivea to the n b to the n c to the nchapter 50
context sensitive inside recursively enumerablethe halting languagechapter 74

So each of the three witnesses is a language this book handles in full, and none of the containments is taken on trust.

Deciding which type a grammar is

The examination asks this, and the method is to check the tightest rule first and work outward.

  1. Is every right hand side a terminal, a terminal and a variable, or epsilon, with a single variable on the left? Then type 3.
  2. Otherwise, is every left hand side a single variable? Then type 2.
  3. Otherwise, does every production have a right hand side at least as long as its left? Then type 1.
  4. Otherwise type 0.

A grammar is always classified by the tightest type it satisfies, because the types are nested. A type 3 grammar is also type 2 and also type 0, and calling it type 0 is not wrong but answers nothing.

Worked

GrammarTightest typeWhy
S->aS, S->b3terminal then variable, or terminal
S->aSb, S->ε2left side one variable; aSb is not terminal then variable
S->Sa, S->b2Sa is a VARIABLE then a terminal, which type 3 does not allow
aS->Sa, S->b1left side is two symbols; nothing shortens
aS->b, S->a0aS to b shortens, from two symbols to one
munotes.in126

The Chomsky Classification of Grammars and Languages

The third row is the one that catches people: S->Sa is not type 3. The right hand side must be a terminal first. A grammar all of whose right hand sides are a variable then a terminal is called left linear and generates a regular language too, but a grammar mixing the two forms generally does not, and chapter 30 is where that is settled.

Where the classification came from

Chomsky was not classifying computer languages. He was asking which mathematical models could describe human language, and his "Three Models for the Description of Language" proposes three, in increasing power, and argues that the weakest is not adequate for English. The hierarchy that carries his name came out of that work and was taken up by computer science afterwards, which is why its vocabulary is linguistic: sentence, grammar, derivation, ambiguity.

The fourth type and the machine correspondences were filled in over the following decade by others, and the linear bounded automaton of chapter 60 is the latest of the four to have been matched to its grammar.

Distinctions

Type 1, context sensitiveType 2, context free
Left hand sideany string with a variableexactly one variable
Can shortennoyes, by an epsilon rule
A rule may depend on surroundingsyesno
Machinelinear bounded automatonpushdown automaton
Membership decidableyesyes
Emptiness decidablenoyes
Type 3 right linearLeft linear
Right hand sidea terminal, then at most one variableat most one variable, then a terminal
ExampleA->aBA->Ba
Generatesa regular languagea regular language
Mixing the two in one grammarnot type 3, and generally not regular

What it does NOT mean

Type 0 is not "no rules". The left hand side must still contain a variable. A production of terminals only could never be applied usefully.

Context sensitive does not mean the rules look at context. It means no production shortens the string. The context formulation is equivalent, and the length one is what you check.

Context free does not mean easy. The class contains languages no finite automaton can manage, and equality of two context free languages is undecidable, which chapter 77 proves.

A grammar is not "type 2 or type 3". It is classified by the tightest type it satisfies. Every type 3 grammar is also type 2, and saying so adds nothing.

munotes.in127

The Chomsky Classification of Grammars and Languages

The hierarchy is about grammars, and the containments are about languages. A regular language has a type 3 grammar and also has type 2, 1 and 0 grammars. What makes it regular is that a type 3 grammar for it exists.

Quick revision

  • Type 0, unrestricted or phrase structure: alpha holds a variable, nothing more. Turing machine. Recursively

enumerable languages.

  • Type 1, context sensitive: beta at least as long as alpha, with S to epsilon allowed. Linear bounded

automaton.

  • Type 2, context free: alpha is a single variable. Pushdown automaton.
  • Type 3, regular or right linear: alpha is a single variable, beta is a terminal, a terminal then a variable,

or epsilon. Finite automaton.

  • The containments are proper, with witnesses a to the n b to the n, then a to the n b to the n c to the n,

then the halting language.

  • Classify by the tightest rule satisfied: check type 3, then 2, then 1, else 0.
  • S->Sa is left linear, not type 3, because a right hand side must begin with a terminal.
  • Tightening the restriction loses expressive power and gains decidability, which is the whole point.

Test yourself

1. Give the restriction defining each of the four types. Type 0: the left hand side contains a variable. Type 1: the right hand side is at least as long as the left, with S to epsilon permitted. Type 2: the left hand side is a single variable. Type 3: as type 2, and the right hand side is a terminal, a terminal followed by one variable, or epsilon.

2. Classify S->aSa | bSb | ε. Type 2. Each left hand side is a single variable, so not type 0 or 1 is needed; but aSa is not a terminal followed by a single variable, so it is not type 3. Note also that the epsilon rule on S is permitted in type 1 only because S appears on no right hand side, and here it does, so this grammar is not type 1 as written.

3. Classify aB->Ba, B->b, S->aB. Type 1. The first production has a two symbol left hand side, so it is not type 2; and nothing shortens, since aB to Ba is two symbols to two.

4. Name a language in each of the three gaps between the four classes. a to the n b to the n is context free and not regular. a to the n b to the n c to the n is context sensitive and not context free. The halting language is recursively enumerable and not context sensitive.

5. Why is S->Sa | b not type 3? Because a type 3 right hand side must be a terminal, or a terminal followed by a variable, or epsilon. Sa begins with a variable, so the grammar is left linear rather than right linear, and left linear grammars are not type 3 by the standard definition even though they generate regular languages.

munotes.in128

The Chomsky Classification of Grammars and Languages

6. What is gained by restricting a grammar, and what is lost? Decidability is gained: emptiness is decidable for context free languages and not for context sensitive ones, and membership is decidable for context sensitive languages and not for type 0. Expressive power is lost, since each class is strictly smaller than the one above it.

Contents This chapter on its own page

munotes.in129

Chapter Twenty-Seven

Recursive and Recursively Enumerable Sets

Syllabus topic Module 1, "Formal Languages: Recursive Enumerable Sets"

In one line

A set is recursively enumerable when a machine can list its members, and recursive when a machine can also decide, for anything at all, whether it is in or out.

In the wording a student can write in an examination: a language L is recursively enumerable if there exists a Turing machine that accepts every string of L and does not accept any string outside L, where a machine that does not accept may either reject or run for ever. L is recursive, also called decidable, if there exists a Turing machine that accepts every string of L and rejects every string outside L, always halting. Every recursive set is recursively enumerable, and the converse is false.

Why the two are different, in one sentence

Because a machine that never stops has not said no.

That is the whole of it, and everything else in this chapter is that sentence made precise. A machine handed a string can do three things: accept, reject, or run for ever. A recursive set has a machine that never does the third. A recursively enumerable set has a machine that may do the third on strings outside the set, and that is the gap between the two notions.

The gap is not a technicality. Chapter 74 proves that a particular set is recursively enumerable and not recursive, so the two classes really are different, and the difference is what "no algorithm exists" means.

Where the names come from

Both names are older than computer science and come from the theory of recursive functions, which is why neither sounds like what it means.

Enumerable means listable. A set is recursively enumerable when a machine can be set going and will print its members, one after another, for ever if the set is infinite. Every member appears eventually. Nothing that is not a member ever appears. But if you are waiting to see whether a particular string appears, and it has not appeared yet, you learn nothing, because it may appear tomorrow.

Recursive here means decidable: there is a mechanical method that always answers. The word has nothing to do with a function calling itself, which is an unrelated later use of the same word.

Modern books say decidable for recursive and semi decidable or Turing recognisable for recursively enumerable. MU uses the older names, so this book uses both and says which is which.

The three outcomes, tabulated

The string isA recursive set's machineA recursively enumerable set's machine
in the setaccepts, and haltsaccepts, and halts
not in the setrejects, and haltsrejects and halts, OR runs for ever

The right hand column's second possibility is the entire difference, and every question on this topic is about it.

munotes.in130

Recursive and Recursively Enumerable Sets

The two theorems worth knowing

These are proved properly in chapter 70. They are stated here because they are the form the examination asks for, and because the first one explains the second.

Theorem 1. Every recursive set is recursively enumerable.

Immediate: a machine that always halts with an answer is in particular a machine that accepts exactly the members, so the definition of recursively enumerable is satisfied without changing anything.

Theorem 2. A set is recursive if and only if both it and its complement are recursively enumerable.

The forward half is easy: if the set is recursive, swap the accept and reject outcomes of its machine and the complement is recursive too, hence recursively enumerable by theorem 1.

The backward half is the useful one and the construction is worth remembering. Suppose L has a machine M and the complement of L has a machine N, each accepting its own set and possibly running for ever otherwise. Build a machine that runs M and N side by side, one step of each in turn. Any input is in L or in the complement of L, so one of the two must accept eventually. If M accepts, accept. If N accepts, reject. So the combined machine always halts with the right answer, and L is recursive.

Running two machines in turn rather than one after the other is the whole trick, and it has a name: dovetailing. It matters because running M to completion first is exactly what does not work, since M may never finish.

The consequence that gets examined. If L is recursively enumerable and its complement is not, then L is not recursive. That is how chapter 75 shows a second problem undecidable once the first has been done, and it is why the complement is worth asking about at all.

Which class each type of grammar lands in

This is the connection to chapter 26 and the reason MU puts the label here.

Grammar typeLanguagesRecursive?
type 3, regularregularyes, and decidable very cheaply
type 2, context freecontext freeyes, by the CYK algorithm of chapter 51
type 1, context sensitivecontext sensitiveyes, by chapter 60's machine
type 0, unrestrictedrecursively enumerablenot in general

So the top of the Chomsky hierarchy is exactly the recursively enumerable sets, and the three lower classes are all recursive. The recursive sets sit strictly between the context sensitive and the recursively enumerable ones: every context sensitive language is recursive, not every recursive language is context sensitive, and not every recursively enumerable language is recursive.

That gives a five level picture rather than the four of chapter 26, and it is worth drawing once:

munotes.in131

Recursive and Recursively Enumerable Sets

regular inside context free inside context sensitive inside recursive inside recursively enumerable

with every containment proper.

A worked example of the difference

The example is a problem rather than a machine, because the machines are Module 2's.

A recursive set. The set of strings over {a, b} with equally many a and b. Given a string, count the a, count the b, compare. That procedure always finishes, in one pass, so the set is recursive.

A recursively enumerable set that is not recursive. The set of pairs consisting of a program and an input on which that program eventually stops. A machine can enumerate this set: run every program on every input a few steps at a time, dovetailing as above, and print each pair as soon as its program stops. Every pair where the program does stop is printed eventually. But no machine can decide the set, because deciding it would require saying no for a program that has not stopped yet and never will, and chapter 74 proves that impossible.

So the difference between the two definitions is the difference between a program you can wait on and a question you can answer.

Enumerating in order, which is a third notion

Worth knowing because it explains the word enumerable and because it makes theorem 2 obvious.

A recursively enumerable set can be listed in some order. A set that can be listed in increasing order, by length and then alphabetically, turns out to be exactly a recursive set. The reason is short: to decide whether w is in the set, list the set in order until you reach w or pass it. If the set is listed in order you must eventually do one or the other, so the procedure halts. And conversely a recursive set can be listed in order by testing every string in canonical order and printing the ones that pass.

So the three notions line up:

Can beMeans
accepted, possibly looping otherwiserecursively enumerable
listed in some orderrecursively enumerable
listed in increasing orderrecursive
decided, always haltingrecursive

Distinctions

RecursiveRecursively enumerable
Modern namedecidablesemi decidable, Turing recognisable
On a memberaccepts and haltsaccepts and halts
On a non memberrejects and haltsmay run for ever
Closed under complementyesno
Grammar typeup to type 1, and moreexactly type 0
Every one of the other isrecursively enumerablenot necessarily recursive
Recursive, the setRecursive, the function
Meansdecidable by an always halting machinea function defined in terms of itself
Relatedhistorically, through recursive function theorynot at all, in this chapter
munotes.in132

Recursive and Recursively Enumerable Sets

What it does NOT mean

Recursive here has nothing to do with a function calling itself. It means decidable. The collision of vocabulary is historical.

Recursively enumerable does not mean you can list the non members. That is the complement, and the theorem above says having both lists makes the set recursive. Most recursively enumerable sets do not have both.

A machine that runs for ever has not rejected. It has not answered. Treating a long run as a no is exactly the mistake the definitions exist to prevent.

Not recursive does not mean not describable. The halting set is perfectly describable in one English sentence. It is not decidable, which is a different thing.

Enumerable does not mean finite. Most of these sets are infinite; enumerable is about being listable, which chapter 7 called countable.

Quick revision

  • Recursively enumerable, also semi decidable: a machine accepts every member and never accepts a non member,

but may run for ever on a non member.

  • Recursive, also decidable: a machine accepts every member, rejects every non member, and always halts.
  • Every recursive set is recursively enumerable. The converse fails, and chapter 74 provides the witness.
  • A set is recursive if and only if it and its complement are both recursively enumerable, proved by running

the two machines side by side, which is called dovetailing.

  • So if a set is recursively enumerable and its complement is not, the set is not recursive.
  • Type 0 grammars generate exactly the recursively enumerable languages; types 1, 2 and 3 all generate

recursive ones.

  • The full chain, all containments proper: regular, context free, context sensitive, recursive, recursively

enumerable.

  • Listable in some order means recursively enumerable; listable in increasing order means recursive.

Test yourself

1. Give the two definitions and the one difference between them. Recursively enumerable: a machine accepts exactly the members, and on a non member may reject or run for ever. Recursive: the same, except that it must always halt, so a non member is always rejected. The difference is whether the machine is permitted not to answer.

2. State and prove that a recursive set is recursively enumerable. A recursive set has a machine that always halts and accepts exactly its members. That machine already satisfies the definition of recursively enumerable, since accepting exactly the members is all that is required. No construction is needed.

3. State the complement theorem and sketch the construction. A set is recursive exactly when it and its complement are both recursively enumerable. Given machines M for the set and N for the complement, run one step of each alternately; every input is in one or the other, so one must accept eventually, and the combined machine answers accordingly and always halts.

munotes.in133

Recursive and Recursively Enumerable Sets

4. Why is running M to completion before starting N wrong? Because M may never complete on an input outside the set, in which case N is never started and the combined machine never answers, which is exactly the failure the theorem is meant to avoid.

5. Which class of languages does a type 0 grammar generate, and which classes are recursive? Type 0 generates exactly the recursively enumerable languages. The regular, context free and context sensitive classes are all recursive, and each is properly contained in the next.

6. A student says a set is not recursively enumerable because their machine ran for an hour without answering. Correct them. A long run is not evidence of anything. The definition of recursively enumerable permits the machine to run for ever on a non member, and the definition of recursive requires a proof that the machine always halts, not an observation that it usually does.

Contents This chapter on its own page

munotes.in134

Chapter Twenty-Eight

Operations on Languages

Syllabus topic Module 1, "Formal Languages: Operations on Languages"

In one line

The set operations apply to languages because a language is a set, and two more, concatenation and closure, come from the fact that its members are strings.

In the wording a student can write in an examination: for languages L and M over Sigma, the union, intersection, difference and complement are the corresponding set operations, the complement being taken with respect to Sigma star. The concatenation LM is the set of all xy with x in L and y in M. The Kleene closure L star is the union of L to the i over all i at least 0, where L to the 0 is the language containing only the empty string, and the positive closure L plus is the same union over i at least 1.

The four set operations

Since a language is a subset of Sigma star, all of chapter 2 applies unchanged.

OperationDefinitionNote
unionevery string in L, in M, or in both
intersectionevery string in both
differenceevery string in L and not in M
complementevery string of Sigma star not in Lneeds Sigma to be fixed

The complement needs a universe and the universe is Sigma star. This is the trap chapter 2 warned about, and here it has teeth: the complement of the language of strings over {a, b} containing ab is not "the strings containing ba". It is every string over {a, b} that does not contain ab, which includes epsilon, a, b, ba, bb, and so on.

Two more set operations are used in this book, and both are applied to a single language.

Reversal. The set of reversals of the members, written L to the R.

Prefix, suffix, substring. The set of all prefixes of members of L, and similarly. Chapter 38 proves the regular languages closed under these, and the proof for prefixes is a one line change to a machine's final states.

Concatenation

L M = { x y : x is in L and y is in M }

One string from L, one from M, in that order, joined with nothing in between.

Three properties and three traps.

It is associative. (LM)N equals L(MN), so LMN is unambiguous.

It is not commutative. LM is usually not ML. If L is {a} and M is {b} then LM is {ab} and ML is {ba}.

The identity is the language {epsilon}, not the empty language. L concatenated with {epsilon} is L, because appending the empty string to anything leaves it alone.

The first trap: the empty language is an annihilator. L concatenated with the empty language is the empty language, for every L, because there is no y to choose from M. So the empty language behaves like zero under concatenation and {epsilon} behaves like one. Students who write L concatenated with { } equals L have confused the two.

munotes.in135

Operations on Languages

The second trap: the size of LM can be less than the product. If L is {a, aa} and M is {a, aa}, then LM is {aa, aaa, aaaa}, which has three members, not four, because aa concatenated with a and a concatenated with aa both give aaa.

The third trap: powers are concatenation, not repetition of a single string. L to the 2 is LM with M equal to L, so it is every concatenation of any member with any member, not just each member with itself.

L to the 0 = {ε}

L to the (k plus 1) = ( L to the k ) L

The first line is a definition and it is forced, exactly as w to the 0 being epsilon was forced in chapter 3: it is what makes the law L to the m concatenated with L to the n equals L to the (m plus n) hold when m is zero.

The Kleene closure

L star = the union of L to the i, for every i at least 0

L plus = the union of L to the i, for every i at least 1

So L star is every string that can be split into any number of pieces, each of which is in L, including zero pieces.

The empty string is always in L star, for every L, because zero pieces concatenate to epsilon. That includes the empty language: the closure of the empty language is {epsilon}, which has one member. It is the single most asked trap on this topic.

L plus is L star with epsilon removed, unless epsilon is in L. If epsilon is in L then L plus already contains it, and L plus equals L star. If epsilon is not in L then L plus is L star minus {epsilon}. Stating it as "L plus is L star without epsilon" is wrong in the first case.

L plus equals L concatenated with L star, and also L star concatenated with L. Both are used in chapter 31.

L star star equals L star. Closing something already closed adds nothing, because a string split into pieces each of which splits into pieces of L splits into pieces of L.

Worked computations

Let Sigma be {a, b}, and take

L = {a, ab}

M = {ε, b}

ExpressionValueWorking
L union M{ε, a, b, ab}four distinct members
L intersect M{ }no member is in both
L minus M{a, ab}neither member of L is in M
L M{a, ab, abb}a with ε and with b; ab with ε and with b, giving ab and abb
M L{a, ab, ba, bab}ε with each of L, then b with each
L to the 0{ε}by definition
L to the 2{aa, aab, aba, abab}every ordered pair of members joined
munotes.in136

Operations on Languages

Two things to notice in that table. LM has three members and not four, because a concatenated with b and ab concatenated with epsilon both give ab. And LM is not ML, which the fourth and fifth rows show explicitly.

Now some closures.

ExpressionValue
{ } star{ε}
{ε} star{ε}
{a} star{ε, a, aa, aaa, ...}
{a, b} starevery string over {a, b}, which is Sigma star
{aa} starthe strings of a of even length
{a, b} plusSigma star without epsilon

The first row is the trap. The second is worth seeing beside it: the closure of {epsilon} is also {epsilon}, because concatenating any number of copies of the empty string gives the empty string.

The operations as machines

Everything above is a definition about sets. Chapter 38 proves that if L and M are regular then so is each of these, by building a machine, and chapter 34 uses three of the constructions to turn a regular expression into a machine. Here is one of them in full, so this chapter is not only definitions.

Union, on two machines. Take a machine for L and a machine for M, and build a new one with a fresh start state having an empty move into each of the two old start states. A run then commits to one machine at the outset and never returns, so the new machine accepts exactly what one or the other accepts.

F1Stateεab
starts{p0, q0}--
p0-p1-
finalp1---
q0--q1
finalq1---

Accepts: a, b

Rejects: ε, ab, ba, aa, bb

L(F1) = L(F1R)

a + b

The two sub machines here accept {a} and {b}, so the union accepts {a, b}, which the claim lists and the comparison confirm.

Substitution, homomorphism, and the reverse of each

The operations above build one language out of two. These build one language out of one, by REPLACING each symbol, and MU has asked for them by name.

Substitution

A substitution replaces every symbol by a whole language. Choose, for each symbol a of Sigma, a language f(a) over some alphabet Delta. Then:

f(ε) = {ε}

f(x a) = f(x) f(a), the concatenation of the two languages

f(L) = the union of f(w), over every w in L

munotes.in137

Operations on Languages

Worked. Let Sigma be {a, b}, Delta be {0, 1}, and:

f(a) = {0, 01}

f(b) = {1}

ArgumentValueWorking
f(a){0, 01}given
f(b){1}given
f(ab){01, 011}each member of f(a) joined to each of f(b)
f(ba){10, 101}the other order, and it is a different language
f(aa){00, 001, 010, 0101}four pairs, all distinct
f({ab, b}){01, 011, 1}the union of f(ab) and f(b)

Notice f(ab) and f(ba). Substitution respects order, because concatenation does.

Homomorphism

A homomorphism is a substitution whose every image is a SINGLE string, so h(w) is one string rather than a language. Writing h(a) for that string:

h(ε) = ε

h(x a) = h(x) h(a), the concatenation of the two strings

h(L) = {h(w) : w in L}

Worked, with h(a) = 01 and h(b) = 1:

ArgumentValue
h(a)01
h(b)1
h(ab)011
h(ba)101
h(aab)01011

Every homomorphism is a substitution and not the other way round. The f above is not a homomorphism, because f(a) holds two strings.

The reverse, which examinations call reverse substitution

Going forwards replaces symbols by languages. Going backwards asks which strings would have been sent into a language you already have.

For a homomorphism, where h(w) is one string, the definition is exact and is the one to quote:

h inverse of M = {w : h(w) is in M}

Worked, with the same h and with M being {0101, 011}:

wh(w)In M?
εεno
a01no
b1no
aa0101YES
ab011YES
ba101no
bb11no

So the reverse of M under h is {aa, ab}, and every longer string is excluded because h never shortens: a contributes two symbols and b one, so nothing longer than two symbols can land in a set whose longest member has four.

For a substitution, where f(w) is a whole language, the question is what "lands in M" should mean, and the strict reading is the useful one: w counts when EVERYTHING f can make of it is in M.

f inverse of N = {w : every string in f(w) is in N}

Worked, with the same f and with N being {0, 01, 011, 0101, 01101}:

wf(w)All of it in N?
ε{ε}no, ε is not in N
a{0, 01}YES
b{1}no
ab{01, 011}YES
ba{10, 101}no
aa{00, 001, 010, 0101}no, three of the four are missing

So the reverse of N under f is {a, ab}.

munotes.in138

Operations on Languages

Why these operations matter. The regular languages are closed under all four, and so are the context free languages, which chapters 38 and 51 record. An inverse homomorphism is the standard way to turn a question about one alphabet into a question about another, and it is how several closure proofs are written.

Distinctions

The empty languageThe language {epsilon}
Membersnoneone, the empty string
Size01
Under concatenationannihilates: L times it is emptyidentity: L times it is L
Its closure{ε}{ε}
Machineno reachable final statethe start state final, nothing else
L starL plus
Pieces allowedzero or moreone or more
Contains epsilonalwaysonly if epsilon is in L
RelationshipL plus union {ε}L concatenated with L star
L to the 2{ ww : w in L }
Membersany member of L then any membereach member doubled
For L = {a, b}{aa, ab, ba, bb}{aa, bb}

What it does NOT mean

L concatenated with the empty language is not L. It is the empty language. The identity for concatenation is {epsilon}.

L star does not exclude epsilon. It always contains it, even when L is empty.

L plus is not always L star minus epsilon. Only when epsilon is not in L.

The complement is not "the other obvious language". It is everything in Sigma star that is not in L, and Sigma must be stated for the question to have an answer.

L to the 2 is not the doubled strings. It is all ordered pairs concatenated, and the last table above shows the difference.

Quick revision

  • Union, intersection, difference and complement are the set operations; the complement is with respect to Sigma

star, which must be stated.

  • LM is every x from L followed by every y from M. Associative, not commutative.
  • {epsilon} is the identity for concatenation; the empty language annihilates.
  • L to the 0 is {epsilon}, which is forced by the exponent law.
  • L star is any number of pieces from L, including none, so epsilon is always in it, and the closure of the

empty language is {epsilon}.

  • L plus is one or more pieces, and equals L star minus {epsilon} only when epsilon is not in L.
  • L star star equals L star, and L plus equals L concatenated with L star.
  • The size of LM can be smaller than the product of the sizes, because different pairs can give the same

string.

  • A substitution replaces each symbol by a LANGUAGE and extends to strings by concatenation and to languages

by union. A homomorphism is the case where each image is one string.

  • The reverse of a homomorphism is the set of w whose image lands in the given language. The reverse of a
munotes.in139

Operations on Languages

substitution is the set of w every one of whose images lands in it.

  • Every homomorphism is a substitution; the reverse does not hold.

Test yourself

1. Give L M and M L for L = {a, ab} and M = {b, ε}. LM is {ab, a, abb}, which is {a, ab, abb} with three members, because a with b gives ab and ab with epsilon also gives ab. ML is {ba, bab, a, ab}, four members.

2. What is the closure of the empty language, and its size? {epsilon}, of size 1. Zero pieces concatenate to the empty string, and that is the only string obtainable.

3. When does L plus equal L star? Exactly when epsilon is in L, since L plus then already contains the empty string from a single piece.

4. Give the complement of the language of strings over {a, b} that end in a. Every string over {a, b} that does not end in a, that is, epsilon together with every string ending in b. The answer must include epsilon, and it is meaningless without saying that Sigma is {a, b}.

5. For L = {a, b}, give L to the 2 and the set of doubled strings. L to the 2 is {aa, ab, ba, bb}. The doubled strings are {aa, bb}. The two are different, and only the first is what L to the 2 means.

6. Simplify the closure of the closure of L, and the concatenation of L with the empty language. The closure of the closure is the closure. The concatenation with the empty language is the empty language.

7. Define substitution, and give f(ab) for f(a) = {0, 01} and f(b) = {1}. A substitution assigns a language to each symbol, sends epsilon to {epsilon}, sends x a to the concatenation of f(x) with f(a), and sends a language to the union of the images of its members. Here f(ab) is f(a) concatenated with f(b), which is {01, 011}.

8. Define reverse substitution, and reverse {0101, 011} under h, where h(a) = 01 and h(b) = 1. The reverse of a language M under a homomorphism h is the set of strings w with h(w) in M; for a general substitution it is the set of w all of whose images lie in M. Here h(aa) is 0101 and h(ab) is 011, both in M, and no other string qualifies, so the answer is {aa, ab}.

9. Is every substitution a homomorphism? No. A homomorphism is the special case in which every image is a single string. The f of question 7 is not one, because f(a) holds two strings.

Contents This chapter on its own page

munotes.in140

Chapter Twenty-Nine

Languages and Automata: Which Machine Goes With Which Grammar

Syllabus topic Module 1, "Formal Languages: Languages and Automata"

In one line

Each of the four grammar types is matched by one of the four machines, and the match is exact: the grammar and the machine describe the same collection of languages.

In the wording a student can write in an examination: the four types of the Chomsky hierarchy correspond exactly to four models of computation. Type 3 grammars and finite automata both characterise the regular languages; type 2 grammars and pushdown automata the context free languages; type 1 grammars and linear bounded automata the context sensitive languages; and type 0 grammars and Turing machines the recursively enumerable languages. In each case both directions are constructive: a grammar can be converted into a machine and a machine into a grammar.

Why the correspondences exist at all

There is no obvious reason why a device that reads a string and answers should describe the same languages as a device that writes strings from nothing. They are opposite processes. The fact that they pair off four times over is the central structural result of this subject, and it has a reason worth stating.

A machine's run and a grammar's derivation are the same object read in two directions. A run turns a string into a sequence of states; a derivation turns a start symbol into a string. So a correspondence between them is a matter of recording, in the grammar's variables, whatever the machine keeps in its memory, and of recording, in the machine's memory, whatever the grammar's variables stand for.

And that is exactly why the four rows of the table line up the way they do: each machine's memory matches the shape of its grammar's productions.

The grammar restricts a production toSo a derivation needs to rememberAnd the machine's memory is
a terminal then at most one variableone variable at a timeone state
a single variable on the lefta stack of pending variablesa stack
nothing shortensno more space than the string itselfa tape the length of the input
nothing at allanythingan unbounded tape

Reading that table is the quickest way to remember the correspondence, because it is the reason rather than the fact.

The correspondence

TypeGrammarMachineLanguagesConversions in this book
3regular, right linearfinite automatonregularchapter 30, both ways
2context freepushdown automatoncontext freechapters 58 and 59
1context sensitivelinear bounded automatoncontext sensitivechapter 61
0unrestrictedTuring machinerecursively enumerablechapter 70

Four rows, and each is an if and only if. The smallest case is proved in full in chapter 40, the second in chapters 58 and 59, and the remaining two are stated with their constructions in Module 2.

munotes.in141

Languages and Automata: Which Machine Goes With Which Grammar

Row by row, with the reason

Type 3 and the finite automaton

A type 3 production is A to aB: write one terminal, and hand over to one variable. A derivation therefore has exactly one variable in it at every step, sitting at the right hand end.

So the whole state of a derivation is which variable that is. A finite automaton's whole state is which state it is in. One variable per state and one production per transition, and the correspondence is immediate. Chapter 30 carries it out in both directions.

That is also the reason a finite automaton cannot count: a derivation with one variable can carry only a bounded amount of information forward, and a machine with finitely many states likewise.

Type 2 and the pushdown automaton

A type 2 production is A to gamma, where gamma may hold several variables. A derivation can therefore have many variables outstanding at once, and they must be dealt with in a definite order.

A leftmost derivation deals with the leftmost first, and the ones to its right wait. That is a stack: the variable most recently created and not yet expanded is the next one to be expanded, which is last in, first out. So the machine that matches a context free grammar is a machine with a stack, and the construction of chapter 58 makes the stack hold exactly the outstanding variables.

Type 1 and the linear bounded automaton

A type 1 production never shortens the string. So in a derivation of a string of length n, no sentential form is ever longer than n, because a form longer than n could never come back down.

That is a space bound, and it is the definition of the machine: a linear bounded automaton is a Turing machine whose head may not leave the portion of tape the input occupies. So the derivation and the machine are bounded by the same quantity, and chapter 61 uses that observation in both directions.

Type 0 and the Turing machine

No restriction on the grammar, no restriction on the machine's tape. A derivation may grow without bound and a Turing machine's tape may be used without bound.

Here the correspondence is not quite symmetric, and the asymmetry is the subject of Module 2. A type 0 grammar generates exactly the recursively enumerable languages, and a Turing machine accepts exactly the recursively enumerable languages, where accepting permits the machine to run for ever on a string not in the language. Chapter 27 said why that permission matters.

Four machines, four things they can and cannot do

The other half of the table is the languages that separate the rows, and these are the answers to "give an example".

munotes.in142

Languages and Automata: Which Machine Goes With Which Grammar

MachineA language it accepts that the one below cannotProved in
Turing machinethe halting languagechapter 74
linear bounded automatona to the n b to the n c to the nchapter 50
pushdown automatona to the n b to the nchapters 21 and 36
finite automatonstrings ending in abchapter 13

Every one of those four is worked in this book, so the strictness of the hierarchy is demonstrated rather than asserted.

The correspondence made concrete: a machine and its grammar

The smallest row of the table, shown once here so the correspondence is not abstract. Chapter 30 gives the general construction.

Take a machine for the strings over {a, b} containing ab:

G1Stateab
startSAS
AAB
finalBBB

Accepts: ab, aab, abb, bab, abab

Rejects: ε, a, b, ba, aa, bb, baa

Now write the grammar mechanically: one variable per state, one production per transition, and an epsilon rule on each final state.

S->aA | bS
A->aA | bB
B->aB | bB | ε

Accepts: ab, aab, abb, bab, abab

Rejects: ε, a, b, ba, aa, bb

L(G1) = L(G2)

Read the two side by side. The row for S in the table becomes the two productions of S. The cell S on a holds A, and the production is S to aA: write the a that was read, and hand over to the variable for the state reached. The final state B gets an epsilon rule, because a run may stop there. Nothing else is needed, and the checker confirms the two accept exactly the same language.

Distinctions

A machineA grammar
Given a string, itanswersis asked to produce it
Its memorystates, a stack, or a tapethe variables outstanding in the current form
Running itreading left to rightrewriting
The correspondence recordsthe grammar's variables as its memorythe machine's memory as its variables
RegularContext freeContext sensitiveRecursively enumerable
Memorya fixed number of statesa stacktape as long as the inputunbounded tape
Can match two countsnoyes, one pairyes, severalyes
Membership decidableyesyesyesno
Closed under complementyesnoyesno

What it does NOT mean

The correspondence is not an approximation. Each row is an if and only if with constructions in both directions, not a rough analogy.

A machine is not converted into a grammar by relabelling. The type 3 case looks like relabelling and is not: the epsilon rules on final states are a real step, and the type 2 case in chapter 59 is genuinely involved.

munotes.in143

Languages and Automata: Which Machine Goes With Which Grammar

The four machines are not four programming languages. They are four amounts of memory, and the table above is really a table about memory.

A language having a grammar of one type does not stop it having a grammar of another. Every regular language has a context free grammar. What the table says is which type is the smallest that suffices.

The Turing machine row is not symmetric with the others. A Turing machine accepting a language may loop forever on strings outside it, which is chapter 27's distinction and has no analogue in the three rows below.

Quick revision

  • Four rows: type 3 with the finite automaton, type 2 with the pushdown automaton, type 1 with the linear

bounded automaton, type 0 with the Turing machine. Each is an if and only if.

  • The reason: a machine's memory matches the shape of its grammar's productions. One variable outstanding needs

one state; many outstanding need a stack; a non shortening grammar needs only as much tape as the input; an unrestricted grammar needs an unbounded tape.

  • A type 3 derivation has exactly one variable at every step, at the right hand end.
  • A leftmost derivation of a context free grammar handles outstanding variables last in first out, which is a

stack.

  • A type 1 derivation never produces a form longer than the target string, which is the linear bound.
  • The separating languages: strings ending in ab; a to the n b to the n; a to the n b to the n c to the n; the

halting language.

  • To convert a finite automaton to a grammar: one variable per state, one production per transition, an epsilon

rule on each final state.

Test yourself

1. Give the four pairings. Type 3 with the finite automaton, type 2 with the pushdown automaton, type 1 with the linear bounded automaton, type 0 with the Turing machine, each characterising the regular, context free, context sensitive and recursively enumerable languages respectively.

2. Why does a context free grammar need a stack rather than finitely many states? Because a production may put several variables into the form at once, so a derivation has several outstanding variables and must return to them in a definite order. The most recently created is the next expanded, which is last in, first out.

3. Why is a non shortening grammar matched by a machine with a linear space bound? Because no sentential form in a derivation of a string of length n can be longer than n, since a longer form could never shrink back. So the derivation fits in the space the input occupies, which is what the machine's tape restriction says.

munotes.in144

Languages and Automata: Which Machine Goes With Which Grammar

4. Convert this machine into a grammar: two states p start and q final, p on a to q, p on b to p, q on a to q, q on b to q. P->aQ | bP and Q->aQ | bQ | ε. One variable per state, one production per transition, and the epsilon rule on the variable for the final state.

5. Name the language separating each pair of adjacent classes. Strings ending in ab separates nothing below regular; a to the n b to the n is context free and not regular; a to the n b to the n c to the n is context sensitive and not context free; the halting language is recursively enumerable and not context sensitive.

6. In what way is the Turing machine row unlike the other three? Because a Turing machine accepting a recursively enumerable language may run for ever on a string outside it, so acceptance and decision come apart. In the three rows below, membership is decidable and the machine always answers.

Contents This chapter on its own page

munotes.in145

Chapter Thirty

Regular Grammar: the Right Linear and Left Linear Forms

Syllabus topic Module 1, "Regular Languages: Regular Grammar"

In one line

A regular grammar puts at most one variable in each right hand side, always at the same end, and it generates exactly the languages a finite automaton accepts.

In the wording a student can write in an examination: a grammar is right linear if every production has the form A to wB or A to w, where w is a string of terminals and B is a variable. It is left linear if every production has the form A to Bw or A to w. A grammar is regular if it is right linear or left linear. Both forms generate exactly the regular languages, but a grammar mixing the two forms need not.

Why the form is restricted this way

Because of the reason chapter 29 gave: a derivation in such a grammar has exactly one variable in it at any time, and that variable is always at the same end.

So the whole state of the derivation is which variable it is, and that is a finite amount of information. A finite automaton's state is also a finite amount of information. One variable per state, and the two objects are the same thing written differently.

If a right hand side could hold two variables the derivation could have two outstanding at once, and then the amount of information in a form would grow with the string. That is the context free case of chapter 41, and it needs a stack.

The two forms

Right linear. Every production is A to wB or A to w, with w a string of terminals, possibly empty, and B a single variable. The variable, when there is one, is at the right hand end.

Left linear. Every production is A to Bw or A to w. The variable is at the left hand end.

The strict definition of type 3 in chapter 26 allows only a single terminal, A to aB or A to a. That is a special case of right linear, and any right linear grammar can be put into it by breaking a long w into one production per symbol with a fresh variable each time. This chapter uses the more generous right linear form, because it is what MU's papers write and because the conversion is trivial.

Worked: right linear grammar to a finite automaton

The construction is direct, because the grammar already is a machine.

The procedure.

  1. One state per variable. The start state is the state of the start symbol.
  2. Add one extra state, call it Z, and make it the only final state.
  3. For each production A to aB, add a transition from A on a to B.
  4. For each production A to a, add a transition from A on a to Z.
  5. For each production A to epsilon, mark A as final as well.
  6. For each production A to wB with w longer than one symbol, insert fresh intermediate states, one per extra
munotes.in146

Regular Grammar: the Right Linear and Left Linear Forms

symbol.

The grammar.

S->aA | bS | ε
A->aA | bB
B->aB | bB | ε

Accepts: ε, b, ab, bab, aab, abab

Rejects: a, aa, ba, bba, aaa

Read the variables: S means no ab yet and no a pending, A means an a is pending its b, B means ab has been seen. The epsilon rules on S and B say those two are acceptable places to stop.

The machine, with one state per variable and no extra state needed because every production either names a variable or is epsilon:

H2Stateab
start finalSAS
AAB
finalBBB

Accepts: ε, b, ab, bab, aab, abab

Rejects: a, aa, ba, bba, aaa

L(H2) = L(H1)

The extra state Z of step 2 was not needed here because no production has the form A to a with a bare terminal and no variable. Where one does, Z is required, and the next example shows it.

Worked: the extra final state is sometimes needed

S->aS | bS | aB
B->b

Accepts: ab, aab, bab, abab, bbab

Rejects: ε, a, b, ba, aa, abb

This generates the strings ending in ab. B to b is a production of the form A to a, so B needs somewhere to go after writing the b, and that is Z.

H4Stateab
startS{S, B}S
B-Z
finalZ--

Accepts: ab, aab, bab, abab, bbab

Rejects: ε, a, b, ba, aa, abb

L(H4) = L(H3)

H4 is nondeterministic

The machine is nondeterministic, which is expected: S on a leads to S and to B, because the grammar offered both S to aS and S to aB. That is the general case, and chapter 16 makes it deterministic when a deterministic machine is wanted.

Worked: a finite automaton to a right linear grammar

The reverse of the same construction, and chapter 29 gave it in outline.

The procedure.

  1. One variable per state. The start symbol is the variable of the start state.
  2. For each transition from p on a to q, add the production P to aQ.
  3. For each final state p, add the production P to epsilon.

That is all three steps, and nothing else is needed.

munotes.in147

Regular Grammar: the Right Linear and Left Linear Forms

The machine, accepting strings over {a, b} with an even number of a:

H5Stateab
start finalEOE
OEO

Accepts: ε, b, aa, bb, aab, abab

Rejects: a, ab, ba, aaa, bab

The grammar. Four transitions give four productions, and the one final state gives one epsilon rule:

E->aO | bE | ε
O->aE | bO

Accepts: ε, b, aa, bb, aab, abab

Rejects: a, ab, ba, aaa, bab

L(H6) = L(H5)

Compare the two objects line by line: the row for E becomes the productions of E, the cell E on a holding O becomes E to aO, and E being final becomes E to epsilon. The construction is a change of notation, and that is exactly what the correspondence of chapter 29 predicted for this row.

The left linear form

Everything above has the variable on the right. The left linear form puts it on the left, so the string is built from the right hand end backwards.

S->Sa | Sb | Ba
B->b

Accepts: ba, baa, bab, baab, babb

Rejects: ε, a, b, ab, aa, abb

This generates the strings beginning with ba, which is the mirror image of what H3 did. A derivation grows leftward: S to Sa to Sba wait, no: S to Sa puts an a on the right of the S, so the terminals accumulate on the right and the variable stays at the left, which means the symbols nearest the front of the string are produced last.

Left linear grammars generate exactly the regular languages too. The cleanest proof is by reversal: take a left linear grammar, reverse every right hand side, and the result is right linear and generates the reversal of the original language. Chapter 38 proves the regular languages closed under reversal, so the original language is regular too.

The trap: mixing the two forms

This is the thing to remember from the chapter, and MU's classification question depends on it.

A grammar in which some productions are right linear and others left linear is called linear, and a linear grammar need not generate a regular language.

The witness is small:

S->aSb | ε

Every production has at most one variable in it, so the grammar is linear. But the variable has a terminal on each side, so it is neither right linear nor left linear. And its language is a to the n b to the n, which chapter 21 proved is not regular.

munotes.in148

Regular Grammar: the Right Linear and Left Linear Forms

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb

So the restriction is not "at most one variable". It is "at most one variable, at a fixed end", and the fixed end is what makes the derivation's state finite. In S to aSb the terminals accumulate on both sides at once, and the machine would have to remember how many are still owed on the right, which is unbounded.

FormEvery productionRegular?
right linearA to wB or A to wyes
left linearA to Bw or A to wyes
linearat most one variable, anywherenot necessarily
context freeany right hand sidenot necessarily

Distinctions

Right linearLeft linear
Variable sitsat the right endat the left end
String is builtleft to rightright to left
Converts to a machinedirectly, one variable per stateby reversing first, or by a mirrored construction
Type 3 as strictly definedyes, after splitting long terminal stringsnot by the strict definition, though the language is regular
A regular grammarA linear grammar
Ruleat most one variable, at a fixed endat most one variable, anywhere
Languagealways regularnot always
Witness that they differS to aSb with S to epsilon

What it does NOT mean

At most one variable is not the condition. The variable must also be at a fixed end. S to aSb has one variable and generates a non regular language.

A right linear grammar is not required to use single terminals. A to abcB is right linear. The strict type 3 definition of chapter 26 wants single terminals, and splitting is mechanical.

A left linear grammar is not worse than a right linear one. Both generate the regular languages exactly. Right linear is preferred here only because its conversion to a machine is direct.

The extra state Z is not always needed. It is needed exactly when some production is A to w with no variable and w non empty.

A regular grammar need not be the smallest grammar for its language. Minimality is a question about machines in this book, and chapter 20 answers it there.

Quick revision

  • Right linear: every production is A to wB or A to w. Left linear: A to Bw or A to w. Either one is called a

regular grammar.

  • Both generate exactly the regular languages, and the left linear case is proved by reversal.
  • Grammar to machine, right linear: one state per variable, plus a final state Z; A to aB becomes a transition,

A to a becomes a transition to Z, A to epsilon marks A final.

  • Machine to grammar: one variable per state, a production P to aQ per transition, and P to epsilon for each
munotes.in149

Regular Grammar: the Right Linear and Left Linear Forms

final state.

  • The resulting machine is usually nondeterministic, which chapter 16 fixes when determinism is wanted.
  • A linear grammar allows one variable anywhere and need NOT be regular: S to aSb with S to epsilon generates a

to the n b to the n.

Test yourself

1. Give the two forms of a regular grammar. Right linear, where every production is a string of terminals optionally followed by one variable; and left linear, where every production is one variable optionally preceded by nothing and followed by a string of terminals, that is A to Bw or A to w.

2. Convert this machine to a grammar: states p start and q final, p on 0 to p, p on 1 to q, q on 0 to q, q on 1 to p. P->0P | 1Q and Q->0Q | 1P | ε. One production per transition and an epsilon rule on the variable for q.

3. Why is S->aSb not right linear, and why does it matter? Because the variable has a terminal after it as well as before, so it is not at the right hand end. It matters because the language generated, a to the n b to the n, is not regular, so the restriction to a fixed end is what the theorem needs.

4. When is the extra final state Z needed in the grammar to machine construction? Exactly when some production has the form A to w with w a non empty string of terminals and no variable, since the machine then needs somewhere to be after writing w.

5. Is every grammar with at most one variable per right hand side regular? No. Such a grammar is called linear, and S->aSb | ε is linear and generates a non regular language. Regular requires the variable at a fixed end.

6. Why does the machine produced from a right linear grammar tend to be nondeterministic? Because a variable may have two productions beginning with the same terminal, such as S to aS and S to aB, and those become two transitions from one state on one symbol.

Contents This chapter on its own page

munotes.in150

Chapter Thirty-One

Regular Expressions

Syllabus topic Module 1, "Regular Languages: Regular Expressions"

In one line

A regular expression is an algebraic way of writing a language, built from the symbols with three operations: union, concatenation and closure.

In the wording a student can write in an examination: the regular expressions over an alphabet Sigma are defined recursively. The empty set and the empty string are regular expressions, each symbol of Sigma is a regular expression, and if R and S are regular expressions then so are R plus S denoting the union, RS denoting the concatenation, and R star denoting the Kleene closure. A language is regular if it is denoted by some regular expression.

Why an algebra as well as a machine and a grammar

Because it is the only one of the three that is short enough to write in a sentence.

A machine for "strings ending in ab" is a table with six cells. A grammar for it is four productions. The regular expression is (a + b)* a b, which fits in a line and can be read aloud. That is why every tool that takes a pattern from a user takes a regular expression, and why chapters 34 and 35 are about converting between the expression and the machine in both directions.

It is also an algebra in the strict sense: the identities of chapter 32 let one expression be transformed into another, and that is how the state elimination method of chapter 35 does its work.

The definition, piece by piece

The base cases. Three of them.

ExpressionDenotesSize of the language
the empty set symbolthe empty language0
εthe language {ε}1
a, for a in Sigmathe language {a}1

The first two are the pair chapter 3 warned about, and they are different expressions denoting different languages. In practice the empty set appears rarely and this book writes it out in words when it does.

The three operations.

WrittenCalledDenotes
R + Sunion, or alternationevery string in either language
R Sconcatenationevery string of R's language followed by one of S's
R starthe Kleene closureany number of pieces from R's language, including none

And one abbreviation used everywhere:

WrittenMeans
R plusR R star, that is, one or more pieces

Chapter 28 defined all four operations on languages, and this chapter is the notation for them.

The dialect: what + means here

This is the single most important paragraph of the chapter for a student who also programs.

In this subject, + means union. So a + b denotes the two string language {a, b}.

In a programming language's regular expressions, + means one or more. So a+ there denotes the language of one or more a, which in this subject is written a a star or a plus.

munotes.in151

Regular Expressions

The two dialects collide on exactly that symbol, and they are otherwise similar enough to be confusing. Some textbooks, and some of MU's own papers, use a vertical bar for union instead of a plus, which removes the clash; this book uses + throughout because that is what her papers mostly print, and it uses no other operator that a programming language would read differently.

There is also a great deal in a programming language's dialect that is not in this one: character classes, bounded repeats, anchors, backreferences, lookahead. Some of those are convenience and some genuinely go beyond the regular languages. None is available here, and an answer that uses them has answered a different question.

Precedence, and the brackets

Without a rule, a + b c would be ambiguous. The rule is:

star binds tightest, then concatenation, then union

So:

WrittenMeansNot
a b stara followed by any number of bany number of ab
a + b ca, or b followed by c(a plus b) followed by c
a b + cab, or ca followed by (b plus c)
(a b) starany number of ab
(a + b) ca or b, then c

That is the same precedence as multiplication before addition in ordinary algebra, with star playing the part of an exponent, and the analogy is close enough to be worth using: think of + as plus, concatenation as times, and star as a power.

Brackets override it, and a well written answer uses them freely. (a + b)* a b has two pairs of brackets in this book's printing and needs only the first; the second is habit and costs nothing.

Reading an expression aloud

This is a real skill and it is how you check your own answer. Read from the outside in.

ExpressionRead aloud
(a + b)* a banything at all, then ab
a bany number of a, then any number of b
(a b)*any number of repetitions of the pair ab
a (a + b)* bbegins with a, ends with b, anything between
(a + b) a (a + b)contains at least one a
b (a b)*hmm: any number of b, then any number of blocks each being an a and some b, which is everything
(a a)*an even number of a
a (a a)*an odd number of a

The sixth row is worth pausing on, because it shows that two expressions can look very different and denote the same language, which is what chapter 32's identities are for.

munotes.in152

Regular Expressions

Worked: expression and machine, compared exactly

Every expression in this book is compared against a machine for the same language, and the comparison is exact rather than by examples. Here is the pair for "contains at least one a".

(a + b)* a (a + b)*
J2Stateab
startnonesomenone
finalsomesomesome

Accepts: a, ab, ba, aab, bab, abab

Rejects: ε, b, bb, bbb

L(J2) = L(J1)

Two states for a one line expression. And here is the pair for "an even number of a", where the expression is the one that is hard to write and the machine is easy.

b* (a b* a b*)*
J4Stateab
start finalevenoddeven
oddevenodd

Accepts: ε, b, aa, bb, aab, abab, baab

Rejects: a, ab, ba, aaa, bab

L(J4) = L(J3)

Read J3 aloud: any number of b, then any number of blocks, each block being an a, some b, another a, some b. Each block contributes exactly two a, so the count is even. That reading is an argument, and the comparison is the proof.

The empty set and epsilon, as expressions

Worth one section, because examination questions set them deliberately.

ExpressionLanguageNote
the empty set{ }no strings at all
ε{ε}one string, of length 0
the empty set, starred{ε}zero pieces gives epsilon
ε star{ε}any number of empty strings is empty
a, times the empty set{ }nothing to follow the a with
a, times ε{a}epsilon is the identity
a + the empty set{a}the empty set is the identity for union

The third row is the one that surprises people and it follows straight from chapter 28: the closure of any language contains epsilon, because zero pieces are allowed.

Distinctions

In this subjectIn a programming language
+unionone or more
Union written+ or a vertical bara vertical bar
One or moreR R star, or R plusR+
Character classesnonesquare brackets
Anchors, backreferencesnonepresent, and some go beyond regular
ε as an expressionthe empty set as an expression
Language{ε}{ }
Size10
Identity forconcatenationunion
Its closure{ε}{ε}
a b star(a b) star
Meansone a, then any number of bany number of the pair ab
Contains ayesno, epsilon is there instead
Contains ababnoyes

What it does NOT mean

Plus is not "one or more" here. It is union. This is the error that costs the most marks on this topic.

munotes.in153

Regular Expressions

A regular expression is not a machine. It denotes a language, and chapter 34 builds the machine.

Two different expressions can denote the same language. (a + b) and (a b) denote the same thing. Chapter 32 is about proving such equalities, and the checker behind this book decides them exactly.

Star does not mean "some". It means any number including zero, so every starred expression denotes a language containing epsilon.

There is no subtraction and no complement in the notation. The regular languages are closed under both, which chapter 38 proves, but the closure is a theorem about machines and the notation has no symbol for it.

Quick revision

  • Base cases: the empty set, epsilon, and each symbol of Sigma.
  • Operations: R + S for union, RS for concatenation, R star for the closure. R plus abbreviates R R star.
  • In this dialect + is UNION, not one or more.
  • Precedence: star, then concatenation, then union, like exponent, times, plus. Brackets override.
  • Read an expression aloud from the outside in; it is how you check your own answer.
  • The closure of anything contains epsilon, so the starred empty set denotes {epsilon}.
  • Epsilon is the identity for concatenation, the empty set for union, and the empty set annihilates under

concatenation.

  • Different expressions may denote the same language, which is what chapter 32's identities are for.

Test yourself

1. What does + mean in this subject, and what does it mean in a programming language? Union here, so a + b denotes {a, b}. One or more in a programming language, where a+ denotes one or more a, written a a star here.

2. Give the precedence rule and bracket a + b c star accordingly. Star binds tightest, then concatenation, then union. So it means a, or b followed by any number of c, that is a + (b (c star)).

3. What language does the starred empty set denote, and why? {epsilon}. A closure allows zero pieces, and the concatenation of zero strings is the empty string, so epsilon is in the closure of any language including the empty one.

4. Write a regular expression for the strings over {a, b} of odd length. (a + b) ((a + b)(a + b))*. One symbol, then any number of pairs.

5. Distinguish a b star from (a b) star. The first is one a followed by any number of b, so it contains a, ab, abb. The second is any number of repetitions of ab, so it contains epsilon, ab, abab, and does not contain a.

6. Write an expression for the strings over {a, b} that do not contain bb. (a + b a)* (b + ε). Every b is either followed by an a inside a block or is the single optional b at the very end, so two b can never be adjacent.

Contents This chapter on its own page

munotes.in154

Chapter Thirty-Two

The Identities of Regular Expressions

Syllabus topic Module 1, "Regular Languages: Regular Expressions"

In one line

The identities are the rules that let one regular expression be rewritten as another denoting the same language, and they are what makes the notation an algebra.

In the wording a student can write in an examination: two regular expressions are equivalent if they denote the same language, and an identity is an equivalence holding for all regular expressions substituted into it. The standard identities are the ones governing the empty set and the empty string as identities and annihilators, the commutativity and associativity of union, the associativity of concatenation, the distribution of concatenation over union, the idempotence of union and of closure, and Arden's rule for a recursive equation.

Why identities are worth learning

Three reasons, in increasing order of importance.

To simplify an answer. A conversion from a machine, by the method of chapter 35, produces an expression far longer than it needs to be, and the identities are how it is brought down to something readable.

Because MU asks for them by name. The 2022 paper sets them as a question of their own.

Because chapter 35's method cannot work without them. Arden's theorem, which is the last identity below, is what turns a set of simultaneous equations into a closed form answer, and the state elimination method is identity 11 applied repeatedly. So this chapter is machinery for the next but one, not a list to be memorised in isolation.

The identities

Write R, S and T for any regular expressions, and write the empty set expression out in words as "the empty set" to keep it distinguishable from epsilon. Each identity is numbered, and the numbering is this book's own, for reference in later chapters.

The base cases

Identity 1. The empty set is the identity for union.

R + the empty set = R

Adding a language with no strings in it adds no strings.

Identity 2. The empty set annihilates under concatenation.

R times the empty set = the empty set

R times ε = R

The first is chapter 28's trap: with nothing to choose from the second factor, no string can be formed at all. The second says epsilon is the identity for concatenation.

Identity 3. The closure of the empty set, and of epsilon, are both epsilon.

( the empty set ) star = ε

ε star = ε

Both follow from the closure allowing zero pieces, and the first is the trap of chapter 28.

Union

Identity 4. Union is commutative and associative, and idempotent.

R + S = S + R

(R + S) + T = R + (S + T)

R + R = R

Idempotence is the one to notice: a language does not grow by being united with itself, because a set has no repeats. In practice it is what lets a duplicated alternative be struck out of an answer.

munotes.in155

The Identities of Regular Expressions

Concatenation

Identity 5. Concatenation is associative and is NOT commutative.

(R S) T = R (S T)

There is no identity saying RS equals SR, and there is no such identity to be had, since a b and b a denote different one string languages.

Identity 6. Concatenation distributes over union, on both sides.

R (S + T) = R S + R T

(S + T) R = S R + T R

This is the identity that does the most work in practice, and it works in both directions: read left to right it multiplies out, and read right to left it factors, which is usually what shortens an answer.

Closure

Identity 7. The closure absorbs itself and epsilon.

(R star) star = R star

ε + R star = R star

R star R star = R star

Each says the same thing in a different place: a string split into pieces each of which splits into pieces of R splits into pieces of R.

Identity 8. The closure unfolds, in two directions.

R star = ε + R R star

R star = ε + R star R

This is the identity that lets a closure be peeled: either a string has no pieces, or it has a first piece and then the rest. It is the basis of the induction in most proofs about starred expressions, and it is what identity 11 is a solved form of.

Identity 9. R plus, and its relation to R star.

R plus = R R star = R star R

R star = ε + R plus

R plus = R star, when ε is in the language of R

The third line is chapter 28's trap: the usual statement "R plus is R star without epsilon" fails exactly when R itself denotes a language containing the empty string.

Identity 10. Two closures joined.

(R + S) star = (R star S star) star

(R star S star) star = (R star + S star) star

Both sides of the first denote all strings made of pieces from either language, and the star on the outside is what makes the two equal. It is worth knowing because the right hand form turns up naturally in chapter 35's answers and the left hand form is what a reader wants.

Arden's rule

Identity 11. The solution of a recursive equation.

if X = R X + S and ε is not in the language of R, then X = R star S

munotes.in156

The Identities of Regular Expressions

This is the most useful identity in the chapter and it has a name because chapter 35 is built on it. Read it as an equation about an unknown language X: if X is defined in terms of itself by prefixing something from R, or else being something from S, then X is any number of R pieces followed by an S piece.

Why it is true. Substituting the right hand side into itself repeatedly gives S, then RS, then RRS, and so on, which is exactly R star S. The condition that R does not contain epsilon is what makes the solution unique: without it, X equal to R X plus S has more than one solution, because an R piece could be empty and the recursion could turn over for ever without consuming anything.

The mirror form, which chapter 35 also uses:

if X = X R + S and ε is not in the language of R, then X = S R star

Each identity, proved by machine

Every identity above is an equation between two expressions, and each side denotes a language. The checker builds a machine from each side, forms the product, and searches for any string on which the two disagree. Finding none over every string there is, not over a sample, is a proof.

Here are the identities with a concrete R, S and T, so the proof is visible rather than described. Take R equal to a, S equal to b, T equal to a b.

a (b + a b)
a b + a a b

L(K1) = L(K2)

That is identity 6, distribution, with those three expressions. Now identity 8, unfolding:

(a b)*
ε + a b (a b)*

L(K3) = L(K4)

Now identity 10, two closures joined:

(a + b)*
(a* b*)*

L(K5) = L(K6)

And Arden's rule, identity 11, with R equal to a and S equal to b, so the claim is that the solution of X equals aX plus b is a star b:

a* b
b + a a* b

L(K7) = L(K8)

The second expression is one unfolding of the recursion, and it agrees with the closed form, which is what the rule says.

Now identity 7, and the closure of the empty string:

((a b)*)*

L(K9) = L(K3)

And identity 4, idempotence, which looks too obvious to check and is checked anyway:

a b + a b

L(K10) = L(K11)

a b

Using the identities: a worked simplification

An answer produced by chapter 35's method might come out as:

(a + b)* b (a + b)* + (a + b)* b (a + b)* b (a + b)*
munotes.in157

The Identities of Regular Expressions

Read it: "contains a b", or "contains a b, then later another b". The second alternative is a special case of the first, so the whole thing should collapse to the first.

The reasoning, identity by identity. The second alternative is (a + b) b (a + b) b (a + b). By identity 6 read right to left, both alternatives begin with (a + b) b (a + b), so the expression factors as that followed by (ε + b (a + b)). And ε + b (a + b) is a sub language of (a + b), so the concatenation is contained in (a + b) b (a + b) already. Hence:

(a + b)* b (a + b)*

L(K12) = L(K13)

The checker decides that equality exactly, so the simplification is correct and not merely plausible. Reasoning of this kind is what an examination answer should show, and the machine comparison is what a book owes its reader.

Distinctions

UnionConcatenation
Commutativeyesno
Associativeyesyes
Idempotentyes, R + R = Rno, RR is not R
Identitythe empty setε
Annihilatornonethe empty set
R starR plus
Pieceszero or moreone or more
Contains εalwaysonly if R's language does
Equal to each otherwhen ε is in R's language

What it does NOT mean

RS equals SR is not an identity. Concatenation is not commutative, and the temptation to treat these expressions as ordinary algebra founders here.

R plus R is not R squared. Union is idempotent, so R + R is R; it is concatenation that gives R squared.

Arden's rule needs its condition. Without epsilon being absent from R's language, the equation has more than one solution and the rule gives the wrong one. Chapter 35 states this again, because it is where it bites.

An identity is not a simplification. Identity 8 read left to right makes an expression longer. Which direction shortens an answer depends on the answer.

Equality of expressions is not equality of strings. Two expressions are equal when they denote the same language, and they may look nothing alike: (a + b) and (a b) are equal.

Quick revision

  • The empty set is the identity for union and annihilates under concatenation; epsilon is the identity for

concatenation; the closure of either is epsilon.

  • Union is commutative, associative and idempotent. Concatenation is associative and NOT commutative.
  • Concatenation distributes over union on both sides, and reading that right to left is how answers get shorter.
  • Closure absorbs itself: R star starred is R star, R star R star is R star, and epsilon plus R star is R star.
  • R star unfolds as epsilon plus R R star, and also as epsilon plus R star R.
  • R plus is R R star; R star is epsilon plus R plus; and the two closures coincide exactly when epsilon is in
munotes.in158

The Identities of Regular Expressions

R's language.

  • Arden's rule: X equals RX plus S has the unique solution R star S when epsilon is not in R's language, and

the mirror form X equals XR plus S has S R star.

  • Every identity here is proved by building a machine from each side and deciding equality exactly.

Test yourself

1. State the two identities governing the empty set. The empty set is the identity for union, so R plus the empty set is R; and it annihilates under concatenation, so R times the empty set is the empty set.

2. Is RS equal to SR? Justify. No. Taking R as a and S as b gives ab on one side and ba on the other, which denote different one string languages. Concatenation is associative but not commutative.

3. Simplify ε + a a star. It is a star, by identity 9: a a star is a plus, and epsilon plus R plus is R star.

4. State Arden's rule with its condition, and say why the condition is needed. If X equals RX plus S and the language of R does not contain epsilon, then X equals R star S. The condition makes the solution unique: if R could contribute the empty string, the recursion could turn over without consuming anything and more than one language would satisfy the equation.

5. Show, by reasoning rather than by examples, that (a + b) equals (a b). Every string over {a, b} is a sequence of symbols, each of which is in a or in b, so it is a sequence of pieces from a b and lies in the right hand side. Conversely every piece of a b is a string over {a, b}, so the right hand side is contained in the left. The two containments give equality.

6. Which direction of the distribution identity usually shortens an answer, and why? Right to left, that is, factoring. A conversion from a machine produces many alternatives sharing long common prefixes or suffixes, and pulling the shared part out is what makes the expression readable.

Contents This chapter on its own page

munotes.in159

Chapter Thirty-Three

Writing a Regular Expression for a Language You Are Given

Syllabus topic Module 1, "Regular Languages: Regular Expressions"

In one line

To write a regular expression, say what the string must contain and where, then write each part in turn with "anything at all" standing between the fixed parts.

In the wording a student can write in an examination: a regular expression for a language is constructed by decomposing the description into the fixed substrings the language requires and the unrestricted portions between them, writing the closure of the whole alphabet for each unrestricted portion, and combining the parts by concatenation, with union used where the description offers alternatives.

The four building blocks

Almost every answer is assembled from these, and recognising which is needed is the whole task.

Phrase in the questionWhat to write
anything at all, over Sigma = {a, b}(a + b)*
exactly this substringthe substring itself, as terminals
any number of thesethat expression, starred
either this or thatthe two, joined with +

The first is the one beginners omit. "Contains ab" is not ab; it is (a + b) a b (a + b), because anything may come before and after.

The method

Step 1. Decide what is fixed and what is free. Write the description as a sequence: free part, fixed part, free part, and so on.

Step 2. Write each free part as the closure of the alphabet. For Sigma = {a, b} that is (a + b)*, and this book writes it out every time rather than abbreviating, because an answer that abbreviates it has to define the abbreviation.

Step 3. Write each fixed part literally.

Step 4. Where the description says "or", join with +. Where it says "and", think again, because "and" is usually not expressible directly and needs the structure rearranged. The regular languages are closed under intersection, by chapter 38, but the notation has no symbol for it, so an "and" has to be reasoned into a single pattern.

Step 5. Check the boundary cases. Does the empty string qualify? Does a one symbol string? Those are where a wrong answer shows itself, and they are what the claim lists under each answer below test.

The twelve sets

Throughout, Sigma is {a, b}, except where 0 and 1 are used, and each answer is followed by a machine built independently and an exact comparison.

1. All strings of a and b

(a + b)*
M2Stateab
start finalqqq

Accepts: ε, a, b, ab, bab

L(M1) = L(M2)

One state, and it accepts everything. Trivial, and worth writing once because it is the free part of every other answer.

2. All strings ending in 00

MU's 2022 paper sets exactly this.

munotes.in160

Writing a Regular Expression for a Language You Are Given

(0 + 1)* 0 0
M4State01
startz0z1z0
z1z2z0
finalz2z2z0

Accepts: 00, 000, 100, 1000, 0100

Rejects: ε, 0, 1, 01, 10, 001

L(M4) = L(M3)

The machine's third row is the cell students get wrong: from z2, reading another 0 stays in z2, because a string ending in 000 still ends in 00.

3. All strings beginning with 0 and ending with 1

Also from her 2022 paper.

0 (0 + 1)* 1

The boundary case matters. The shortest string in this language is 01, of length 2, because the first and last symbols are different and so cannot be the same symbol. A student who writes 0 (0 + 1) 1 has this right and a student who writes 0 (0 + 1) 1 + 0 1 has written the same language twice.

M6State01
startsmd
mmf
finalfmf
ddd

Accepts: 01, 001, 011, 0101, 0001

Rejects: ε, 0, 1, 10, 00, 11, 100

L(M6) = L(M5)

State d is the dead state, entered when the first symbol is a 1.

4. All strings of 0 and 1 whose length is odd

Her 2022 paper again.

(0 + 1) ((0 + 1)(0 + 1))*

Accepts: 0, 1, 000, 010, 10101

Rejects: ε, 00, 01, 1010, 0000

L(M7) = L(M8)

M8State01
startevenoddodd
finaloddeveneven

5. All strings containing exactly two a

b* a b* a b*

Accepts: aa, aba, baab, bbabba, ababb

Rejects: ε, a, b, aaa, ab, bab

L(M9) = L(M10)

M10Stateab
startn0n1n0
n1n2n1
finaln2n3n2
n3n3n3

Three b blocks and two a, and the b blocks may each be empty, which is what lets aa qualify.

6. All strings with at least two a

(a + b)* a (a + b)* a (a + b)*

Accepts: aa, aba, aaa, baab, abbaa

Rejects: ε, a, b, ab, ba, bb, bab

L(M11) = L(M12)

M12Stateab
startc0c1c0
c1c2c1
finalc2c2c2

Compare with set 5: exactly two needs b star between the a, at least two needs the whole alphabet starred. That is the commonest single error on this topic.

munotes.in161

Writing a Regular Expression for a Language You Are Given

7. All strings in which every 0 is immediately followed by at least two 1

MU's 2018 paper sets this and then asks for a proof that the expression describes the same set.

(1 + 0 1 1)*

Accepts: ε, 1, 11, 011, 0111, 1011, 011011

Rejects: 0, 01, 10, 0110, 0011, 01101

L(M13) = L(M14)

M14State01
start finalg0g1g0
g1dg2
g2dg0
ddd

Read the expression: a legal string is a sequence of blocks, each block being a single 1 or the three symbols

  1. So every 0 carries its two 1 with it, and a 0 at the end or with one 1 after it cannot be part of any block.

The equality is decided by the checker, which is the proof MU's paper asks for, and the reasoning just given is how it is written out.

8. All strings not containing the substring aa

(b + a b)* (a + ε)

Accepts: ε, a, b, ab, ba, bab, abab

Rejects: aa, baa, aab, aaa, abaa

L(M15) = L(M16)

M16Stateab
start finalh0h1h0
finalh1ddh0
dddddd

Every a is either followed by a b inside a block, or is the single optional a at the very end.

9. All strings whose second symbol from the left is a

(a + b) a (a + b)*

Accepts: aa, ba, aab, bab, aabb

Rejects: ε, a, b, ab, bb, abb

L(M17) = L(M18)

M18Stateab
startp0p1p1
p1yesno
finalyesyesyes
nonono

10. All strings of even length over {0, 1}

((0 + 1)(0 + 1))*

Accepts: ε, 00, 01, 10, 11, 0011

Rejects: 0, 1, 010, 101, 00000

L(M19) = L(M20)

M20State01
start finaleoo
oee

11. All strings over {0, 1} with an even number of 0 and an even number of 1

This is the "and" case of step 4, and it cannot be written by joining two answers with a plus.

(0 0 + 1 1 + (0 1 + 1 0)(0 0 + 1 1)*(0 1 + 1 0))*

Accepts: ε, 00, 11, 0110, 0011, 1001, 0101

Rejects: 0, 1, 01, 10, 001, 111, 0111

L(M21) = L(M22)

M22State01
start finaleeoeeo
eoooee
oeeeoo
ooeooe

The machine is four states and immediate: one state per pair of parities, reading a 0 flips the first and reading a 1 flips the second, and only the both even state is final. The expression is long and hard to be sure of by eye, and it is worth reading once. Each block of the outer closure returns the machine to the both even state: either two of the same symbol, which flips one parity twice, or a pair of different symbols, which flips both, then any number of same symbol pairs, then another pair of different symbols, which flips both back.

munotes.in162

Writing a Regular Expression for a Language You Are Given

That is the general lesson of this set: when a question has two independent conditions, build the machine first, and convert it by chapter 35's method if an expression is what was asked for. Writing the expression directly is possible, as here, and it is where answers go wrong. The comparison above is what makes this one trustworthy.

12. All strings over {a, b} of length at most 3

(ε + a + b)(ε + a + b)(ε + a + b)

Accepts: ε, a, b, ab, aba, bbb

Rejects: abab, aaaa, babab

L(M23) = L(M24)

L(M23) is finite

M24Stateab
start finalk0k1k1
finalk1k2k2
finalk2k3k3
finalk3k4k4
k4k4k4

Three optional symbols. Writing ε + a + b three times is the cleanest answer; writing out all fifteen strings as alternatives is also correct and much longer.

The four traps, collected

Each of these was met above, and each is worth naming.

Forgetting the free part. "Contains ab" is not ab.

Exactly against at least. Exactly two a needs b a b a b*; at least two needs the alphabet starred between them.

The boundary cases. Does the empty string qualify? Does a one symbol string? Set 3's shortest member is of length 2, and set 12's includes epsilon.

An "and" of two conditions. Not expressible by joining two answers. Build the machine and convert, or reason the two conditions into a single pattern as set 11 does.

Distinctions

b a b a b*(a + b) a (a + b) a (a + b)*
Meansexactly two aat least two a
Contains aaanoyes
ab(a + b) a b (a + b)
Meansthe single string abevery string containing ab
Size of the language1infinite
(a + b)(a + b)(ε + a + b)(ε + a + b)
Meansexactly two symbolsat most two symbols
Contains epsilonnoyes

What it does NOT mean

A regular expression cannot say "and". There is no intersection symbol. The class is closed under intersection but the notation is not, so an "and" must be reasoned into one pattern.

munotes.in163

Writing a Regular Expression for a Language You Are Given

A regular expression cannot say "not". Same reason. The complement of a regular language is regular, and the expression for it has to be worked out, usually by building the machine, complementing its final states, and converting back.

Two expressions that look different are not necessarily different. Every answer above is one of many correct ones, and the checker tests the language rather than the form.

The shortest expression is not required. Any correct one earns the marks, and simplifying with chapter 32's identities is a separate exercise.

Quick revision

  • Four blocks: the alphabet starred for a free part; the substring itself for a fixed part; a star for "any

number"; a + for "either".

  • Method: split the description into free and fixed parts, write each, join by concatenation, use + for "or",

then check the boundary cases.

  • "Contains X" needs the alphabet starred on both sides of X.
  • Exactly n of a symbol needs the other symbols starred between them; at least n needs the whole alphabet

starred.

  • Two independent conditions cannot be joined with +. Build the machine and convert it, or reason the two into

one pattern.

  • The notation has no intersection and no complement, though the class is closed under both.
  • Always test the empty string and the one symbol strings against your answer.

Test yourself

1. Write an expression for the strings over {a, b} containing at least one a and at least one b. (a + b) a (a + b) b (a + b) + (a + b) b (a + b) a (a + b). The two orders of the first a and the first b are the two alternatives, and neither alone is enough.

2. Write an expression for the strings over {0, 1} that begin and end with the same symbol, of length at least two. 0 (0 + 1) 0 + 1 (0 + 1) 1. The two cases are the two possible shared symbols, and the middle is free.

3. What is wrong with writing a b for "contains ab"? It denotes the single string ab. "Contains ab" permits anything before and after, so the free parts must be written: (a + b) a b (a + b).

4. Give the difference between b a b a b and (a + b) a (a + b) a (a + b). The first has exactly two a, because only b may appear elsewhere. The second has at least two, because the free parts may contain more a.

5. Why can you not write an expression for "an even number of a and an odd number of b" by joining two expressions with a plus? Because a plus is union, which gives strings satisfying either condition, and what is wanted is strings satisfying both. The notation has no intersection symbol, so the two conditions must be reasoned into one pattern, or a four state machine built and converted.

munotes.in164

Writing a Regular Expression for a Language You Are Given

6. Write an expression for the strings over {a, b} of length at most 2, and say whether epsilon is in it. (ε + a + b)(ε + a + b). Epsilon is in it, obtained by choosing epsilon for both factors.

Contents This chapter on its own page

munotes.in165

Chapter Thirty-Four

From a Regular Expression to a Finite Automaton

Syllabus topic Module 1, "Regular Languages: Finite automata and Regular Expressions"

In one line

Build a small machine for each symbol, then glue the small machines together with empty moves, one rule per operation of the expression.

In the wording a student can write in an examination: for every regular expression R there is a nondeterministic finite automaton with epsilon moves accepting the language denoted by R. It is constructed by structural induction on R: base machines are given for the empty set, for epsilon and for each symbol, and machines for R plus S, RS and R star are built from machines for R and S by adding a new start state, a new final state, and epsilon transitions.

Why the construction is worth having

Because a regular expression is what a person writes and a machine is what a program runs. Every tool that takes a pattern from a user does this conversion, and chapter 16 then makes the result deterministic. The two chapters together are what happens inside a search tool when a pattern is compiled.

And because it proves half of Kleene's theorem. Chapter 40 assembles the four equivalences of Module 1, and this chapter supplies the step from an expression to a machine.

The invariant that makes it work

Every machine the construction builds has three properties, and keeping them is what lets the rules be applied blindly.

Exactly one start state, with no arrow entering it.

Exactly one final state, with no arrow leaving it.

The two are different states.

That is the whole trick. Because nothing enters the start state and nothing leaves the final state, two such machines can be joined at those points without any risk of a path sneaking through in an unintended direction. A machine with several final states, or with an arrow back into its start state, cannot be glued safely, which is why the construction insists on the invariant at every step.

The cost is extra states: each operation adds two, so the machine has about twice as many states as the expression has symbols and operators. Chapter 20 removes them afterwards if a small machine is wanted.

The six rules

Write the two sub machines as A, for R, and B, for S, each with its own start and final state.

Rule 1. A single symbol a. Two states, one arrow labelled a from the first to the second.

N1Statea
starts1f1
finalf1-

Accepts: a

Rejects: ε, aa

L(N1) = L(N1R)

a

Rule 2. Epsilon. Two states, one empty arrow between them.

N2Stateεa
starts2f2-
finalf2--

Accepts: ε

Rejects: a, aa

Rule 3. The empty set. Two states and no arrow at all, so nothing is ever accepted.

munotes.in166

From a Regular Expression to a Finite Automaton

N3Statea
starts3-
finalf3-

Rejects: ε, a, aa

L(N3) is empty

Rule 4. Union, R plus S. A new start state with an empty arrow into each sub machine's start state, and a new final state with an empty arrow from each sub machine's final state.

Rule 5. Concatenation, RS. An empty arrow from A's final state to B's start state. A's start state becomes the start, B's final state becomes the final.

Rule 6. Closure, R star. A new start state and a new final state. An empty arrow from the new start to A's start, from A's final to the new final, from the new start straight to the new final, which allows zero pieces, and from A's final back to A's start, which allows more than one piece.

The fourth arrow of rule 6 is the one to write down carefully: it goes from A's final state back to A's start state, not to the new start state, and getting that wrong is the commonest slip in the construction.

Worked: union, in full

Build a machine for a + b, with the sub machines of rule 1.

N4Stateεab
starts{p0, q0}--
p0-p1-
p1f--
q0--q1
q1f--
finalf---

Accepts: a, b

Rejects: ε, ab, ba, aa

L(N4) = L(N4R)

a + b

Six states for a two symbol expression, which is the cost the invariant buys: two for each of the two sub machines and two new ones.

Worked: concatenation, in full

Build a machine for a b.

N5Stateεab
startu0-u1-
u1v0--
v0--v1
finalv1---

Accepts: ab

Rejects: ε, a, b, ba, abb

L(N5) = L(N5R)

a b

Four states, and the single empty arrow from u1 to v0 is the join. No new start or final state is needed for concatenation, which is why it is the cheapest of the three operations.

Worked: closure, in full

Build a machine for a star.

N6Stateεa
startw0{x0, w1}-
x0-x1
x1{x0, w1}-
finalw1--

Accepts: ε, a, aa, aaa, aaaa

There is nothing over this machine's alphabet for it to reject, since it accepts every string of a, and a rejection claim naming a b would be a claim about a symbol the machine has never heard of. What is worth proving instead is that it is exactly the closure and that its language is infinite.

munotes.in167

From a Regular Expression to a Finite Automaton

L(N6) = L(N6R)

L(N6) is infinite

a*

All four arrows of rule 6 are visible in the table. From w0 to x0 enters the sub machine; from w0 to w1 skips it, which is what admits epsilon; from x1 to w1 leaves it; and from x1 back to x0 is the loop that admits more than one piece.

Worked: a whole expression

Build a machine for (a + b)* a b b, which is the language of strings ending in abb. The construction proceeds from the inside out: rule 1 four times for the symbols, rule 4 for the union, rule 6 for the closure, then rule 5 three times to concatenate the four parts.

Carrying that out gives fourteen states. Rather than print all fourteen and their empty moves, here is the result after chapter 16's subset construction and chapter 20's minimisation, together with the original expression, and the checker decides that the two agree over every string there is.

N7Stateab
startt0t1t0
t1t1t2
t2t1t3
finalt3t1t0

Accepts: abb, aabb, babb, abbabb, bbabb

Rejects: ε, a, b, ab, abba, aab

L(N7) = L(N7R)

N7 is deterministic

(a + b)* a b b

Four states after minimising, against fourteen from the construction. That gap is normal and it is the reason the three chapters run in the order they do: build blindly, determinise, then minimise.

The whole pipeline

This is worth having on one page, because MU's questions can start at any point in it.

FromToByChapter
a regular expressionan NFA with empty movesThompson's constructionthis chapter
an NFA with empty movesan NFAthe closure method17
an NFAa DFAthe subset construction16
a DFAthe minimal DFApartition refinement20
a DFAa regular expressionArden or state elimination35
a DFAa regular grammarone variable per state30
a regular grammaran NFAone state per variable30

Every arrow in that table is a construction this book carries out, and the whole table closing on itself is Kleene's theorem, which chapter 40 states.

Distinctions

Rule 4, unionRule 5, concatenationRule 6, closure
New states added202
Empty arrows added414
The loopnonenonefrom A's final back to A's start
Admits epsilononly if a sub machine doesonly if both doalways
Thompson's constructionDesigning by hand
Inputan expressiona description
Resultan NFA with empty moves, about twice the sizeusually the minimal DFA
Can be done blindlyyes, that is the pointno
Needs the invariantyesno
munotes.in168

From a Regular Expression to a Finite Automaton

What it does NOT mean

The machine produced is not deterministic. It has empty moves and usually several arrows on one symbol. Chapters 17 and 16 make it deterministic.

The machine produced is not minimal. It is about twice the size of the expression, and chapter 20 shrinks it.

The loop in rule 6 does not go to the new start state. It goes from the sub machine's final state back to the sub machine's start state. An arrow to the new start state would also work here but breaks the invariant when several closures are nested.

The invariant is not decoration. Without exactly one final state and no arrow into the start state, rules 4, 5 and 6 can create paths the expression does not denote.

The construction does not need the expression simplified first. It works on any expression, and a longer expression simply gives a bigger machine.

Quick revision

  • Every machine built keeps the invariant: one start state with nothing entering it, one different final state

with nothing leaving it.

  • Rule 1: a symbol is two states and one labelled arrow. Rule 2: epsilon is two states and one empty arrow. Rule

3: the empty set is two states and no arrow.

  • Rule 4, union: a new start with empty arrows into both sub machines, a new final with empty arrows from both.
  • Rule 5, concatenation: one empty arrow from the first sub machine's final to the second's start. No new states.
  • Rule 6, closure: a new start and a new final, with four empty arrows, including the loop from the sub machine's

final back to its start, and the arrow from new start to new final that admits epsilon.

  • The result is an NFA with empty moves, roughly twice the size of the expression. Chapters 17, 16 and 20 turn it

into the minimal DFA.

  • The whole pipeline closes on itself, and that is Kleene's theorem.

Test yourself

1. State the invariant the construction maintains, and why it matters. One start state with no arrow entering it, one final state with no arrow leaving it, and the two distinct. It matters because it lets two machines be glued at those points without any path running in an unintended direction, which is what makes the rules applicable without looking inside a sub machine.

2. How many states does rule 6 add, and where does each of its four arrows go? Two. From the new start to the sub machine's start; from the sub machine's final to the new final; from the new start straight to the new final, which admits epsilon; and from the sub machine's final back to its start, which admits more than one piece.

munotes.in169

From a Regular Expression to a Finite Automaton

3. Which rule adds no new states, and why can it not? Concatenation. The first machine's final state and the second's start state are joined by one empty arrow, and the invariant is preserved with the first machine's start and the second machine's final serving as the new ones.

4. Build the machine for a b and say how many states it has. Four: two from each symbol's base machine, joined by one empty arrow from the first's final state to the second's start state.

5. Is the machine the construction produces deterministic or minimal? Neither. It has empty moves and is about twice the size of the expression. Chapter 17 removes the empty moves, chapter 16 determinises, and chapter 20 minimises.

6. Why does the construction need empty moves at all? Because gluing two machines together requires an arrow that consumes no input: the symbol the first machine finished on has already been read, so there is nothing left for the joining arrow to consume.

Contents This chapter on its own page

munotes.in170

Chapter Thirty-Five

From a Finite Automaton to a Regular Expression: Arden's Theorem and State Elimination

Syllabus topic Module 1, "Regular Languages: Finite automata and Regular Expressions"

In one line

Write one equation per state saying how that state is reached, solve them with Arden's rule, and the equation for the final state is the answer.

In the wording a student can write in an examination: for every finite automaton there is a regular expression denoting the language it accepts. One method writes a simultaneous equation for each state, whose unknown is the set of strings that take the machine from the initial state to that state, and solves the system using Arden's rule, that X equals RX plus S has the unique solution R star S when the language of R excludes the empty string. A second method deletes states one at a time, replacing the paths through each deleted state by labelled transitions carrying regular expressions.

Arden's theorem

Statement. Let R and S be regular expressions over an alphabet, and suppose the language of R does not contain the empty string. Then the equation

X = R X + S

has the unique solution

X = R star S

Proof that R star S is a solution. Substitute it into the right hand side:

R (R star S) + S = (R R star) S + S = (R plus) S + S = (R plus + ε) S = R star S

The first step is associativity of concatenation, the second is identity 9 of chapter 32, the third is distribution read right to left, and the fourth is identity 9 again. So R star S satisfies the equation.

Proof that it is the only solution. Suppose X is any solution. Substituting the equation into itself n times gives

X = S + R S + R squared S + ... + R to the n S + R to the (n plus 1) X

Now take any string w in X, of length k. Because the language of R excludes the empty string, every string in R to the m has length at least m. So for n greater than k the last term contributes nothing of length k, and w must come from one of the earlier terms, all of which are inside R star S. So X is contained in R star S, and the first half gave the other containment. Hence X equals R star S.

Where the condition bites. If the language of R does contain epsilon, the argument above fails at "every string in R to the m has length at least m", and the equation genuinely has more than one solution. Take R equal to epsilon and S equal to a. Then X equals X plus a is satisfied by {a} and by every larger language, so there is no unique answer.

munotes.in171

From a Finite Automaton to a Regular Expression: Arden's Theorem and State Elimination

The mirror form, used when the equations are written the other way round:

X = X R + S has the unique solution X = S R star

Method 1: the equation method

The procedure.

  1. For each state q write an unknown X for the set of strings taking the machine from the start state to q.
  2. Write one equation per state: X for q equals the sum, over every transition into q, of the expression for the

source state concatenated with the symbol on that transition. For the start state, add epsilon, because the empty string reaches it.

  1. Solve the system by substitution, using Arden's rule whenever an unknown appears on both sides.
  2. The answer is the sum of the expressions for the final states.

Step 2 is where the direction has to be watched. The equation for q is about arrows into q, not out of it. A student who writes the equations from the outgoing arrows has written a different and wrong system.

Worked, in full

Take the machine for the strings over {a, b} ending in ab.

P1Stateab
starts0s1s0
s1s1s2
finals2s1s0

Accepts: ab, aab, bab, abab, bbab

Rejects: ε, a, b, ba, abb, aa

Step 1 and 2. The equations. Write X0, X1, X2 for the three states, and read off the arrows into each.

Into s0: from s0 on b, and from s2 on b. And s0 is the start state, so epsilon is added.

Into s1: from s0 on a, from s1 on a, and from s2 on a.

Into s2: from s1 on b, and nothing else.

X0 = ε + X0 b + X2 b

X1 = X0 a + X1 a + X2 a

X2 = X1 b

Step 3. Solve. Take them in the easiest order, which is the one with the fewest unknowns on the right.

The third equation gives X2 directly in terms of X1, so substitute it into the second:

X1 = X0 a + X1 a + X1 b a

= X0 a + X1 (a + b a)

The second line factors X1 out by identity 6 read right to left. Now X1 appears on both sides in the mirror form, so Arden's rule gives

X1 = X0 a (a + b a) star

Substitute that into the first equation, remembering that X2 is X1 b:

X0 = ε + X0 b + X1 b b

= ε + X0 b + X0 a (a + b a) star b b

= ε + X0 ( b + a (a + b a) star b b )

munotes.in172

From a Finite Automaton to a Regular Expression: Arden's Theorem and State Elimination

Arden's rule again, in the mirror form with S equal to epsilon:

X0 = ( b + a (a + b a) star b b ) star

and then working back:

X1 = ( b + a (a + b a) star b b ) star a (a + b a) star

X2 = ( b + a (a + b a) star b b ) star a (a + b a) star b

Step 4. The answer is X2, because s2 is the only final state.

(b + a (a + b a)* b b)* a (a + b a)* b

L(P2) = L(P1)

The checker decides that equality exactly, so the solution is right and not merely plausible. And the two intermediate expressions are right too, which is worth checking separately because an error in X0 or X1 would usually still give a plausible looking X2.

(b + a (a + b a)* b b)*
(b + a (a + b a)* b b)* a (a + b a)*
P5Stateab
start finals0s1s0
s1s1s2
s2s1s0

Accepts: ε, b, abb, bb, babb

Rejects: a, ab, aab, ba, aa

P6Stateab
starts0s1s0
finals1s1s2
s2s1s0

Accepts: a, aa, ba, aba, abba

Rejects: ε, b, ab, bb, abb

L(P3) = L(P5)

L(P4) = L(P6)

P5 and P6 are the same machine with a different final state, so their languages are exactly the sets X0 and X1 were defined to be. Both expressions check out, so the whole solution is verified and not just its last line.

And the answer is not the shortest one. The machine plainly accepts (a + b)* a b, and so:

(a + b)* a b

L(P7) = L(P1)

Both expressions are correct. Arden's method is mechanical and gives something correct; chapter 32's identities are what get it down to something short, and an examiner will accept either.

Method 2: state elimination

Often quicker by hand, and it is the method to use when the machine has more than three or four states.

The idea. Allow the arrows of the machine to carry whole regular expressions rather than single symbols. Then a state can be deleted by replacing every path through it with a single arrow carrying the expression for those paths.

The procedure.

  1. Add a new start state with an empty arrow to the old start state, and a new single final state with empty

arrows from every old final state. Neither new state is ever deleted. This is chapter 34's invariant again, and it is needed for the same reason.

munotes.in173

From a Finite Automaton to a Regular Expression: Arden's Theorem and State Elimination

  1. Pick any state other than the two new ones and delete it. For every pair of a state p with an arrow in and a

state q with an arrow out, replace those by a single arrow from p to q carrying

(the arrow from p in) (the loop on the deleted state) star (the arrow out to q)

where the loop's star is epsilon if the deleted state has no loop. Where p already had an arrow to q, the two expressions are joined with a plus.

  1. Repeat until only the two new states remain.
  2. The single arrow between them carries the answer.

Worked, on the same machine

Take P1 again, and add the new start N and the new final F.

Arrows at the start. N to s0 on epsilon. s0 to s0 on b, s0 to s1 on a. s1 to s1 on a, s1 to s2 on b. s2 to s0 on b, s2 to s1 on a. s2 to F on epsilon.

Delete s1. Arrows in: from s0 on a, from s2 on a. Arrows out: to s2 on b. The loop on s1 is a, so its star is a star.

  • s0 to s2 gains a a star b.
  • s2 to s2 gains a a star b.

Now the arrows are: N to s0 on epsilon; s0 to s0 on b; s0 to s2 on a a star b; s2 to s0 on b; s2 to s2 on a a star b; s2 to F on epsilon.

Delete s0. Arrows in: from N on epsilon, from s2 on b. Arrows out: to s2 on a a star b. Loop on s0 is b, so its star is b star.

  • N to s2 gains b star a a star b.
  • s2 to s2 gains b b star a a star b, joined with the a a star b it already had.

Now: N to s2 on b star a a star b; s2 to s2 on a a star b + b b star a a star b; s2 to F on epsilon.

Delete s2. Arrows in: from N. Arrows out: to F. The loop on s2 is the union just written, so:

  • N to F gains b star a a star b ( a a star b + b b star a a star b ) star.

The answer:

b* a a* b (a a* b + b b* a a* b)*

L(P8) = L(P1)

A third correct expression for the same language, arrived at by a different route, and checked. Which of P2, P7 and P8 a student produces depends on the method and the order of elimination, and all three earn the marks.

munotes.in174

From a Finite Automaton to a Regular Expression: Arden's Theorem and State Elimination

The order matters for length, not for correctness. Deleting the most connected state last usually gives the shortest answer, and deleting in a different order here gives a different expression for the same language.

Distinctions

The equation methodState elimination
What you writeone equation per statea diagram with expressions on the arrows
The toolArden's rulethe path replacement formula
Directionequations are about arrows INarrows are followed forward
Best forthree or four states, and when Arden is asked for by namelarger machines
The answer's lengthusually longdepends on the deletion order
X = RX + SX = XR + S
SolutionR star SS R star
Which ariseswhen equations are written with the source on the rightwhen written with the source on the left, as here

What it does NOT mean

Arden's rule does not need the machine to be deterministic. The equations can be written for any finite automaton, and an NFA gives a union of terms on the right hand side.

The equation for a state is not about its outgoing arrows. It is about the arrows into it. This is the error that produces a system with a plausible solution and the wrong language.

Epsilon is added only to the start state's equation. Adding it elsewhere claims the empty string reaches that state, which it does not.

The answer is not unique. Three correct expressions for one machine appear above. Only the language is determined.

Arden's condition is not a formality. Without epsilon being absent from R, the equation has many solutions and the rule picks the smallest, which may not be the one the machine denotes.

Quick revision

  • Arden's theorem: X equals RX plus S has the unique solution R star S when the language of R excludes epsilon.

The mirror form X equals XR plus S has S R star.

  • The uniqueness proof: substitute repeatedly, and use the fact that every string of R to the m has length at

least m, which is exactly what the condition guarantees.

  • Equation method: one unknown per state, one equation per state written from the arrows IN, epsilon added to the

start state's equation, solved by substitution and Arden. The answer is the sum over the final states.

  • State elimination: add a new start and a new single final state, then delete states one at a time, replacing

paths through a deleted state by an arrow carrying in, loop starred, out; join parallel arrows with a plus.

munotes.in175

From a Finite Automaton to a Regular Expression: Arden's Theorem and State Elimination

  • Both methods give correct expressions, usually different ones, and neither gives the shortest.
  • Chapter 32's identities are what shorten the answer.

Test yourself

1. State Arden's theorem, both forms, with the condition. If the language of R does not contain epsilon, then X equals RX plus S has the unique solution R star S, and X equals XR plus S has the unique solution S R star.

2. Prove that R star S satisfies X equals RX plus S. Substitute: R R star S plus S is R plus S plus S, which is (R plus plus epsilon) S by distribution, which is R star S by identity 9.

3. Why does the uniqueness proof need epsilon to be absent from R? Because it relies on every string of R to the m having length at least m, so that for large enough n the term R to the n plus one X cannot contribute a string of a given length. If R contained epsilon that bound fails and the equation has many solutions.

4. In the equation method, which arrows does a state's equation use, and what extra term does the start state get? The arrows INTO that state, each contributing the source state's unknown concatenated with the transition's symbol. The start state's equation gains an epsilon term, because the empty string reaches it.

5. In state elimination, what expression replaces the paths through a deleted state? The expression on the incoming arrow, then the loop on the deleted state starred, then the expression on the outgoing arrow. Parallel arrows between the same pair are joined with a plus.

6. Two students produce different expressions for the same machine. Are both wrong? Neither need be. The language is determined but the expression is not, and the chapter gives three correct expressions for one machine, from Arden, from inspection and from state elimination.

Contents This chapter on its own page

munotes.in176

Chapter Thirty-Six

The Pumping Lemma for Regular Languages

Syllabus topic Module 1, "Regular Languages: Pumping Lemma and its Applications"

In one line

Every long enough string of a regular language has a short piece near the front that can be repeated any number of times without leaving the language.

In the wording a student can write in an examination: let L be a regular language. Then there exists a constant n, called the pumping length, such that every string z in L with length at least n can be written as z equal to uvw with the length of uv at most n, the length of v at least 1, and u v to the i w in L for every i at least 0.

Why the lemma exists

Because up to this chapter there is no way to prove a language is not regular except by Myhill Nerode.

Chapter 21 gave one method and it is the stronger one, being an if and only if. But it works on an infinite family of strings and an extension argument, and examinations set the pumping lemma instead, because it works on one string and a case analysis. Both are worth having, and a student who knows only one is missing the better half.

The statement, clause by clause

Each of the three conditions is doing something and an answer that drops one has not stated the lemma.

There exists a constant n. It depends on L and you do not get to choose it. In a proof you must handle whatever n the adversary gives you, which is why the string chosen in chapter 37 is always written in terms of n.

Every z in L with length at least n. Short strings are exempt. The lemma says nothing about them, which is why the string you choose must be long enough.

z equals uvw. A splitting into three pieces, in order, with u possibly empty and w possibly empty.

The length of uv is at most n. So the piece that gets repeated lies within the first n symbols of z. This is the condition students drop, and it is the one that does the work: it limits where v can be, which is what lets an argument rule out every possible v.

The length of v is at least 1. The repeated piece is not empty. Without this the lemma would be trivially true and useless.

u v to the i w is in L for every i at least 0. Every number of repetitions, including zero, which deletes v. The i equal to 0 case is often the easiest one to use and is frequently forgotten.

The proof

The whole proof is the pigeonhole principle of chapter 6 applied to states, and it is four sentences long.

munotes.in177

The Pumping Lemma for Regular Languages

Proof. L is regular, so some DFA M accepts it. Let n be the number of states of M, and let z be any string of L with length at least n.

Run M on z. The run visits the state it starts in and one state after each symbol, so on a string of length at least n it visits at least n plus 1 states. M has only n states. By the pigeonhole principle some state is visited twice, and we may take two such visits among the first n plus 1, so both occur within the first n symbols.

Say the state q is visited after reading u and again after reading uv, where v is non empty because the two visits are different points in the run, and the length of uv is at most n because both visits are within the first n symbols. Write w for the rest of z, so z is uvw.

Now v takes M from q back to q. So it may be traversed any number of times, and the machine is in q afterwards whatever that number is. Reading w from q then ends where reading w from q ended when z was accepted, namely in a final state. So u v to the i w is accepted for every i at least 0, including i equal to 0, which skips the loop entirely.

That is the lemma.

The proof in a picture, and why it is called pumping

The run of M on z looks like this: from the start state, u takes it to q; then v takes it from q round a loop back to q; then w takes it from q to a final state.

start, then u, reaches q

q, then v, returns to q

q, then w, reaches a final state

Once a loop is on the path, the machine cannot tell how many times it has been round. So the loop can be pumped: traversed twice, three times, or skipped. Each choice gives a different string, and all of them are accepted.

The lemma made concrete

Take the machine for the strings over {a, b} containing an even number of a, which has 2 states, so n is 2.

Q1Stateab
start finaleoe
oeo

Accepts: ε, b, aa, bb, aab, abab

Rejects: a, ab, ba, aaa, bab

Take z equal to bb, which is in the language and has length 2, so the lemma applies. The run is e, e, e, so the state e is visited three times and certainly twice. Take u empty, v equal to b, w equal to b. Then the length of uv is 1, which is at most 2; the length of v is 1; and

munotes.in178

The Pumping Lemma for Regular Languages

iu v to the i wIn the language?
0byes
1bbyes
2bbbyes
3bbbbyes

All in, as the lemma promises. The loop being pumped is the loop on e labelled b.

Now take z equal to aa. The run is e, o, e. No state repeats within the first n plus 1 states in an obvious place except that e occurs at positions 0 and 2, so u is empty, v is aa, w is empty. But the length of uv would be 2, which is at most n equal to 2, so that splitting is legal. Pumping gives aa, aaaa, aaaaaa, all with an even number of a, all in the language. Again as promised.

Notice what has not been shown. Nothing in either example proves the language is regular. The lemma is a property regular languages have, and observing it hold tells you nothing, because non regular languages can have it too.

The one thing the lemma is for

This is the paragraph that decides marks on this topic.

The lemma says if L is regular then L pumps. Chapter 6 gave the contrapositive, and it is the only useful form:

if L does not pump, then L is not regular

So the lemma is used in exactly one way: assume L is regular, obtain the n it promises, produce a string that cannot be pumped, and conclude a contradiction.

Three consequences follow, and each is a trap.

It can never prove a language regular. Showing that some strings pump proves nothing at all. To prove a language regular you build a machine, an expression or a regular grammar, or you use Myhill Nerode.

Its failure to prove irregularity proves nothing either. There are non regular languages that satisfy the pumping condition. So the lemma is a sufficient test for irregularity and not a necessary one, and when it fails to give a contradiction the right move is Myhill Nerode.

The direction of the quantifiers is not negotiable. The lemma says: there exists n, such that for all long z, there exists a splitting, such that for all i, the pumped string is in L. To contradict it you must show: for every n, there is a long z, such that for every legal splitting, there is an i with the pumped string outside L. Chapter 37 writes that as a game, which is the easiest way to keep the order straight.

munotes.in179

The Pumping Lemma for Regular Languages

A worked contradiction, in short

The full treatment is chapter 37. One example here, so the shape of an argument is visible before the chapter of them.

Claim. The language of strings a to the k followed by b to the k, for k at least 0, is not regular.

Proof. Suppose it is. Let n be the pumping length. Take z equal to a to the n followed by b to the n, which is in the language and has length 2n, at least n.

By the lemma z splits as uvw with the length of uv at most n and v non empty. The first n symbols of z are all a, so uv lies entirely inside the a block, and therefore v is a to the m for some m at least 1.

Pump with i equal to 2. The string u v squared w is a to the (n plus m) followed by b to the n, which has more a than b and so is not in the language. That contradicts the lemma.

Hence the language is not regular.

Where the condition earned its keep: the length of uv being at most n is what forced v to be all a. Without it v could straddle the boundary and the argument would fail.

Distinctions

The pumping lemmaMyhill Nerode
Formif regular then pumpsregular if and only if finite index
Proves irregularitysometimesalways, in principle
Proves regularityneveryes
What you produceone string, in terms of n, and a case analysisan infinite family and a distinguishing extension
Gives the state countnoyes, it is the index
The condition uv at most nThe condition v at least 1
Does whatconfines v to the first n symbolsstops v being empty
If droppedmost proofs fail, because v could sit anywherethe lemma is trivially true and useless

What it does NOT mean

It does not say every splitting pumps. It says some splitting does. So in a proof you must rule out every legal splitting, not just the obvious one.

It does not apply to short strings. Only to strings of length at least n.

It does not include i equal to 1 only. Every i at least 0, and i equal to 0, which deletes v, is often the easiest case to use.

Satisfying the lemma does not make a language regular. The implication runs one way.

You do not choose n. It is given by the supposed machine, and the string you pick must be written in terms of it.

The lemma is not about a machine you have. It is about a machine assumed to exist, which is why the proof is by contradiction.

munotes.in180

The Pumping Lemma for Regular Languages

Quick revision

  • Statement: if L is regular there is an n such that every z in L of length at least n splits as uvw with uv at

most n long, v at least 1 long, and u v to the i w in L for every i at least 0.

  • Proof: take n to be the number of states; a run on a string of length at least n visits at least n plus 1

states; by the pigeonhole principle a state repeats within the first n symbols; the piece between the two visits is a loop and can be traversed any number of times.

  • The only use is the contrapositive: if the condition fails, L is not regular.
  • It can never prove a language regular, and failing to give a contradiction proves nothing.
  • uv at most n is the condition that does the work, because it confines v to the first n symbols.
  • i equal to 0 is allowed and deletes v.
  • Some splitting pumps, not every splitting, so a proof must rule out all legal splittings.

Test yourself

1. State the pumping lemma in full, with all three conditions. If L is regular there is a constant n such that every string z in L with length at least n can be written uvw where the length of uv is at most n, the length of v is at least 1, and u v to the i w is in L for every i at least 0.

2. Prove it. Let M be a DFA for L with n states and let z in L have length at least n. The run on z visits at least n plus 1 states among n, so some state q repeats within the first n symbols. Write u for the part before the first visit, v for the part between the visits, w for the rest. Then v takes q to q, so it can be traversed any number of times, and w still leads to a final state. Hence every u v to the i w is accepted.

3. Which condition forces v to lie in the first block of a string like a to the n b to the n, and why does it matter? The condition that the length of uv is at most n. Since the first n symbols are all a, v must be a block of a, and pumping then changes the a count without changing the b count. Without the condition v could straddle the boundary and the argument would not close.

4. A student shows that every string of some language pumps and concludes it is regular. What is wrong? The implication runs from regular to pumping, not back. Non regular languages can satisfy the pumping condition, so satisfying it proves nothing. Regularity is proved by exhibiting a machine, expression or regular grammar, or by Myhill Nerode.

munotes.in181

The Pumping Lemma for Regular Languages

5. Why is i equal to 0 worth remembering? Because it deletes v altogether, and for many languages the shortened string is the easiest one to show lies outside the language.

6. In a proof by the lemma, who chooses n and who chooses z? n is given by the supposed machine and cannot be chosen; z is chosen by you, and must therefore be written in terms of n so that the argument works whatever n turns out to be.

Contents This chapter on its own page

munotes.in182

Chapter Thirty-Seven

Applications of the Pumping Lemma

Syllabus topic Module 1, "Regular Languages: Pumping Lemma and its Applications"

In one line

To prove a language is not regular, assume it is, take the pumping length it promises, choose a string in terms of that length, and show every legal splitting can be pumped out of the language.

In the wording a student can write in an examination: suppose L is regular and let n be the constant of the pumping lemma. Exhibit a string z in L with length at least n. For every decomposition z equal to uvw with the length of uv at most n and the length of v at least 1, exhibit an i for which u v to the i w is not in L. This contradicts the lemma, so L is not regular.

The argument as a game

The lemma has four quantifiers in a fixed order and getting them backwards is what produces a wrong answer that looks right. Reading the proof as a game between you and an adversary keeps them straight, because each quantifier becomes a move.

MoveWhoseWhat happens
1the adversarychooses n, the pumping length. You know nothing about it.
2youchoose z, in L, of length at least n. It must be written in terms of n.
3the adversarychooses the splitting uvw, subject to uv at most n and v at least 1.
4youchoose i, and must land outside L.

You win if you have a winning reply at every move 4, for every splitting the adversary could have picked. The language is then not regular.

Two consequences of the order.

Move 2 is where the work is. A badly chosen z lets the adversary pick a splitting you cannot beat. A well chosen z leaves the adversary no room, usually because the first n symbols are all the same and so v must be made of them.

Move 3 is the adversary's, so you must handle every case. "Take v to be a single a" is not a proof. Either argue about an arbitrary v, or enumerate the cases, and the first is shorter when it can be done.

The standard trick for choosing z

Choose z so that its first n symbols are all the same symbol. Then the condition that uv has length at most n forces v to be a block of that symbol, and there is only one case to handle rather than several.

That is why every proof below chooses a string beginning with n copies of one symbol, and it is the single most useful habit in this topic.

Language 1: a to the k followed by b to the k

Claim. The language of a to the k followed by b to the k, for k at least 0, is not regular.

munotes.in183

Applications of the Pumping Lemma

Proof. Suppose it is regular, and let n be the pumping length.

Choose z equal to a to the n followed by b to the n. It is in the language and its length is 2n, which is at least n.

Any splitting. z equals uvw with the length of uv at most n and the length of v at least 1. The first n symbols of z are all a, so uv consists of a alone, and therefore v is a to the m for some m with 1 at most m at most n.

Choose i equal to 2. Then

u v squared w = a to the (n plus m) then b to the n

which has n plus m occurrences of a and n of b. Since m is at least 1 these counts differ, so the string is not in the language.

Contradiction, so the language is not regular.

The alternative reply, worth knowing because some questions make it the easier one: choose i equal to 0 instead, giving a to the (n minus m) followed by b to the n, which again has unequal counts.

Language 2: the palindromes over {a, b}

Claim. The language of strings equal to their own reversal is not regular.

Proof. Suppose it is, with pumping length n.

Choose z equal to a to the n, then b, then a to the n. It is a palindrome, and its length is 2n plus 1.

Any splitting. The first n symbols are all a, so v is a to the m with m at least 1.

Choose i equal to 0. Then u w is a to the (n minus m), then b, then a to the n. Reading it backwards gives a to the n, then b, then a to the (n minus m), which is a different string because m is at least 1. So u w is not a palindrome and is not in the language.

Contradiction. Chapter 56 builds the pushdown automaton that does accept this language.

Language 3: the strings with equally many a and b

Claim. The language of strings over {a, b} with as many a as b is not regular.

Proof. Suppose it is, with pumping length n.

Choose z equal to a to the n followed by b to the n, which is in this language too.

Any splitting. As in language 1, v is a to the m with m at least 1.

Choose i equal to 2. The string has n plus m occurrences of a and n of b, so the counts differ and it is not in the language.

munotes.in184

Applications of the Pumping Lemma

Contradiction.

Note how little changed from language 1. The same z and the same v work, because both languages contain that z and both are broken by the same pumping. That is common, and it is worth trying a proof you already have before inventing a new one.

Language 4: a to the k squared

Claim. The language of a to the k squared, that is the strings of a whose length is a perfect square, is not regular.

Proof. Suppose it is, with pumping length n.

Choose z equal to a to the n squared. It is in the language and its length is n squared, which is at least n.

Any splitting. Every symbol is an a, so v is a to the m with 1 at most m at most n.

Choose i equal to 2. The pumped string has length n squared plus m. Now the point: the next perfect square after n squared is n squared plus 2n plus 1, and

n squared < n squared plus m and n squared plus m at most n squared plus n < n squared plus 2n plus 1

because m is at most n. So the length lies strictly between two consecutive perfect squares and is therefore not a perfect square. So the pumped string is not in the language.

Contradiction.

This is the proof worth studying, because it is the one where the arithmetic carries the argument, and because it shows why the bound on the length of uv matters twice over: it forces v to be all a, and it bounds m by n, which is what puts the pumped length below the next square.

Language 5: a to the p for p prime

Claim. The language of strings of a whose length is a prime number is not regular.

Proof. Suppose it is, with pumping length n.

Choose z equal to a to the p, where p is any prime at least n plus 2. There is one, because there are infinitely many primes.

Any splitting. v is a to the m with 1 at most m at most n. Write the length of u w as p minus m, so the pumped string u v to the i w has length p minus m plus i m, that is, p plus (i minus 1) m.

Choose i equal to p minus m. Then the length is

p plus (p minus m minus 1) m = p (1 plus m) minus m (m plus 1) = (p minus m)(1 plus m)

which is a product of two factors. The first is p minus m, which is at least 2 because p is at least n plus 2 and m is at most n. The second is 1 plus m, which is at least 2 because m is at least 1. A number that is a product of two factors each at least 2 is not prime.

munotes.in185

Applications of the Pumping Lemma

Contradiction.

Why the choice of p mattered. Taking p merely at least n would allow p minus m to be 1 or 0, and the factoring argument would break. The habit is to give yourself room, and here two extra symbols is enough.

The four ways these proofs go wrong

Each of these loses marks, and each can be avoided by naming it.

Choosing z without n in it. "Take z equal to aaabbb" is not a proof, because the adversary chose n first and may have chosen n larger than 6, in which case the lemma says nothing about your string.

Handling one splitting. The adversary chooses the splitting. A proof must cover every legal one, which is why the standard trick of putting n identical symbols at the front is worth using: it reduces every legal splitting to one case.

Forgetting that uv is at most n. Without invoking it there is no reason v should sit in the first block, and almost every proof above depends on that.

Pumping in the wrong direction. Both i equal to 0 and i equal to 2 are available, and for some languages only one of them lands outside. Try both.

Distinctions

Languagez to choosei to chooseWhy it works
a to the k b to the ka to the n b to the n2, or 0v is all a, so the counts separate
palindromesa to the n b a to the n0the two a blocks stop matching
equal countsa to the n b to the n2as the first
perfect squaresa to the n squared2the new length sits between consecutive squares
primesa to the p, p at least n plus 2p minus mthe new length factors as two numbers at least 2

What it does NOT mean

A proof does not choose the splitting. The adversary does, and every legal splitting must be beaten.

A proof does not need a machine. The machine is assumed to exist and is never exhibited. That is what makes it a proof by contradiction.

The lemma failing to give a contradiction does not make a language regular. It means the lemma is not the right tool and Myhill Nerode should be tried.

The same z does not work for every language. It works for languages broken by the same pumping, which is a large family, and when it does not the choice has to be rethought.

munotes.in186

Applications of the Pumping Lemma

Quick revision

  • The four moves: the adversary picks n, you pick z in terms of n, the adversary picks uvw with uv at most n and v

at least 1, you pick i and must land outside L.

  • Choose z with its first n symbols all the same, so the condition on uv forces v to be a block of that symbol and

there is one case, not several.

  • a to the k b to the k, equal counts, and palindromes: v is a block of a, and pumping up or deleting breaks the

match.

  • Perfect squares: pump to i equal to 2 and the length lands strictly between n squared and the next square,

because m is at most n.

  • Primes: take p at least n plus 2 and pump to i equal to p minus m, so the length factors as (p minus m)(1 plus

m), both factors at least 2.

  • Four errors: a z with no n in it; handling one splitting; not using uv at most n; pumping in the wrong direction.

Test yourself

1. Whose move is the choice of n, and whose is the choice of the splitting? n is the adversary's, being given by the supposed machine. The splitting is also the adversary's. You choose z and you choose i.

2. Prove that the language of a to the k followed by b to the k is not regular. Suppose regular with pumping length n, and take z equal to a to the n followed by b to the n. Since uv is at most n long and the first n symbols are a, v is a to the m with m at least 1. Pumping to i equal to 2 gives n plus m of a and n of b, which is outside the language. Contradiction.

3. Why choose z to begin with n identical symbols? Because the condition that uv has length at most n then forces v to lie inside that block, so every legal splitting has the same shape and one argument covers all of them.

4. For the perfect squares, why does pumping to i equal to 2 land outside the language? The new length is n squared plus m with m at most n, so it is greater than n squared and less than n squared plus 2n plus 1, which is the next perfect square. A number strictly between consecutive squares is not a square.

5. For the primes, why is p chosen at least n plus 2 rather than at least n? So that p minus m is at least 2. With p only at least n, p minus m could be 1 or 0 and the factorisation would not show the number composite.

munotes.in187

Applications of the Pumping Lemma

6. A student writes "let v be a single a" and completes the argument. What is missing? Every legal splitting must be handled, and the adversary may choose any v of length between 1 and n inside the first block. The fix is to write v as a to the m with m at least 1 and at most n, and argue for arbitrary m.

Contents This chapter on its own page

munotes.in188

Chapter Thirty-Eight

Closure Properties of the Regular Languages

Syllabus topic Module 1, "Regular Languages: Closure Properties"

In one line

The regular languages are closed under every operation in this book, and each closure is proved by building a machine.

In the wording a student can write in an examination: a class of languages is closed under an operation if applying the operation to members of the class always yields a member of the class. The regular languages are closed under union, concatenation, Kleene closure, complement, intersection, difference, reversal, and the prefix, suffix and substring operations, and in each case the proof is a construction on finite automata or on regular expressions.

Why closure properties are worth proving

Two reasons, and the second is the one that earns marks.

They make the class coherent. A class of languages that lost membership under an ordinary operation would be an accident of definition rather than a natural object. That all ten operations stay inside the class is what makes "regular" worth a name.

They are a proof technique. Suppose you want to show a language L is not regular, and you already know some language K is not regular. If you can obtain K from L by operations the class is closed under, then L cannot be regular, because otherwise K would be. Chapter 37's method needs a fresh string and a case analysis every time; this method needs one line. Both are examined and the second is often quicker.

The three that come free from regular expressions

Union, concatenation and closure are immediate, because the notation has a symbol for each.

Union. If L is denoted by R and M by S, then L union M is denoted by R plus S.

Concatenation. L M is denoted by R S.

Kleene closure. L star is denoted by R star.

That is a complete proof in three lines, and it is the shortest route to those three closures. The machine constructions of chapter 34's rules 4, 5 and 6 are the same three results proved the other way, and either is acceptable in an answer.

Complement, by swapping the final states

Construction. Let M be a deterministic and complete machine for L. Build M' with the same states, the same start state and the same transitions, and with the final states swapped: every state that was final becomes non final and every state that was not becomes final.

Why it works. M is deterministic, so every string has exactly one run, and that run ends in exactly one state. The string is in L exactly when that state was final, so it is accepted by M' exactly when it was not accepted by M. Hence M' accepts the complement.

The two conditions are not decoration.

Deterministic. On a nondeterministic machine a string may have several runs, some ending finally and some not, so swapping the final states does not swap acceptance. The machine must be determinised first, by chapter 16.

munotes.in189

Closure Properties of the Regular Languages

Complete. A missing transition means a string has no run at all and is rejected by both machines. The dead state of chapter 5 must be added first, and it then becomes a final state of M', which is exactly right: a string that fell off the old machine belongs in the complement.

Worked

A machine for the strings over {a, b} containing ab:

R1Stateab
startc0c1c0
c1c1c2
finalc2c2c2

Accepts: ab, aab, bab, abab

Rejects: ε, a, b, ba, aa, bb

Swap the final markings, leaving everything else alone:

R2Stateab
start finalc0c1c0
finalc1c1c2
c2c2c2

Accepts: ε, a, b, ba, aa, bb, bba

Rejects: ab, aab, bab, abab

L(R2) = L(R2R)

b* a*

The complement of "contains ab" is "does not contain ab", and a string over {a, b} with no ab is a block of b followed by a block of a, which is what the expression says. The claim lists are exactly swapped, which is the construction working.

Intersection, two ways

Way one, by De Morgan. The class is closed under union and complement, and

L intersect M = (L' union M')'

so it is closed under intersection with no new construction at all. This is the answer to give when the question asks for a proof and nothing more: it is three symbols long and it is complete.

Way two, the product construction. Sometimes the machine is wanted rather than the existence.

Let A and B be deterministic machines for L and M. Build a machine whose states are pairs, one component from each, the Cartesian product of chapter 2.

PartDefinition
statesevery pair (p, q) with p from A and q from B
start statethe pair of the two start states
transition on afrom (p, q) go to (delta A of p and a, delta B of q and a)
final states, for intersectionpairs with BOTH components final
final states, for unionpairs with EITHER component final
final states, for differencepairs with the first final and the second not

One machine, three results, and only the final states change. That is why the product construction is worth knowing even though De Morgan is shorter: it gives union, intersection and difference at once, and it is the construction the checker behind this book uses to decide whether two machines are equal.

munotes.in190

Closure Properties of the Regular Languages

Worked

Intersect "an even number of a" with "ends in b".

R3Stateab
start finaleoe
oeo

Accepts: ε, b, aa, bb, aab

Rejects: a, ab, ba, aaa

R4Stateab
startnny
finalyny

Accepts: b, ab, bb, aab, abb

Rejects: ε, a, aa, ba, aba

The product has four states, one per pair, and only the pair (e, y) is final:

R5Stateab
startenoney
finaleyoney
onenoy
oyenoy

Accepts: b, bb, aab, aabb, abab

Rejects: ε, a, ab, aa, ba, abb

L(R5) = L(R5R)

(b + a b* a)* b b*

Reading the answer: an even number of a arranged as blocks, and then at least one b at the end. The claim lists show the intersection at work: aab is in both R3 and R4 and is in R5; abb is in R4 and not in R3 and is not in R5.

Reversal

Construction. Reverse every arrow of the machine, make the old final states into start states, and the old start state into the only final state. A machine with several start states is handled by adding one new start state with empty moves into all of them, as chapter 34 did.

Why it works. A path from the start to a final state spelling w becomes, with every arrow reversed, a path from that final state to the start spelling w reversed.

The result is nondeterministic even when the original was deterministic, which is expected and harmless: chapter 16 determinises it.

Prefix, suffix and substring

These three are one line each and are worth knowing because they are the cheapest closure proofs in the chapter.

Prefix. Take a machine for L and make final every state from which some final state is reachable. A string is then accepted exactly when it can be continued to a member of L, which is exactly what being a prefix means.

Suffix. Make the start state into a new start state with empty moves to every state from which the start state is reachable... more simply, add a new start state with empty moves to every state reachable from the old start, and keep the final states. A string is accepted exactly when it finishes a member of L.

Substring. Both at once: start anywhere reachable, and accept anywhere from which a final state is reachable.

Each of these changes only the markings, never the transitions, which is why they are so cheap.

munotes.in191

Closure Properties of the Regular Languages

The whole table

OperationClosed?Construction
unionyesR plus S, or the product with either component final
concatenationyesR S, or rule 5 of chapter 34
Kleene closureyesR star, or rule 6 of chapter 34
complementyesswap the final states of a complete DFA
intersectionyesDe Morgan, or the product with both components final
differenceyesthe product with the first final and the second not
reversalyesreverse the arrows, swap start and final
prefixyesmake final every state that can reach a final state
suffixyesstart from every reachable state
substringyesboth of the above

Ten operations, ten closures, and no exceptions. Chapter 51 gives the same table for the context free languages, where three of the rows say no, and the contrast is the point of that chapter.

Using closure to prove a language is not regular

This is the technique worth practising, and it is often three lines where chapter 37 takes half a page.

The pattern. To show L is not regular: find an operation the class is closed under, and a language K known not to be regular, such that applying the operation to L and to some regular languages yields K. If L were regular, K would be regular. It is not, so L is not.

Example 1

Claim. The language of strings over {a, b} with equally many a and b is not regular.

Proof. Call it L, and intersect it with the regular language denoted by a b. The result is the set of strings of the form a to the k followed by b to the k, which chapter 37 proved is not regular. The regular languages are closed under intersection. So if L were regular the intersection would be regular, and it is not. Hence L is not regular.

Three lines, no pumping, no case analysis. And it reuses a result already proved instead of proving a new one.

Example 2

Claim. The language of strings over {a, b} that are NOT palindromes is not regular.

Proof. The regular languages are closed under complement. The complement of this language is the palindromes, which chapter 37 proved is not regular. So if this language were regular its complement would be, and it is not. Hence it is not regular.

One line, and a pumping lemma proof of this language directly is genuinely awkward, because a non palindrome usually stays a non palindrome when pumped.

Example 3

Claim. The language of a to the i b to the j with i not equal to j is not regular.

Proof. Suppose it is. Then so is its complement, by closure. Intersect that complement with a b, which is regular, and the result is the strings a to the k followed by b to the k, which is not regular. Two closure steps from a supposed regular language to a known non regular one, so the supposition fails.

munotes.in192

Closure Properties of the Regular Languages

Note why this one needs the complement. Pumping this language directly is a well known trap: the obvious choices of z and i often land back inside it, because making the counts more unequal keeps the string in the language. Closure avoids the trap entirely.

Distinctions

De Morgan for intersectionThe product construction
Length of the proofone linea table
Gives you a machinenoyes
Also gives union and differencenoyes, by changing the final states only
Number of statesnot applicablethe two counts multiplied
Closure as a factClosure as a technique
Used toshow the class is well behavedprove a language is not regular
Needsa constructiona known non regular language and an operation
Compared with pumpingusually much shorter, and reuses earlier work

What it does NOT mean

The complement construction does not work on an NFA. Determinise first, or the swap does not swap acceptance.

The complement construction needs the machine complete. Add the dead state first, and note that it becomes a final state of the complement, which is correct.

Closure does not mean the operation is easy to carry out. The product of two machines with twenty states each has four hundred, and chapter 20 is usually wanted afterwards.

Closure under an operation does not mean the notation has a symbol for it. There is no complement or intersection symbol in a regular expression, and chapter 33 said why that matters when writing one.

A closure argument still needs the direction checked. "L is regular because K is regular and K comes from L" is not an argument. The implication runs the other way: if L were regular then K would be, and K is not.

Quick revision

  • Union, concatenation and closure come free from the regular expression notation, and also from chapter 34's rules

4, 5 and 6.

  • Complement: swap the final states of a COMPLETE DETERMINISTIC machine. Both conditions are needed.
  • Intersection: De Morgan in one line, or the product construction, whose states are pairs and which gives union,

intersection and difference by changing only the final states.

  • Reversal: reverse every arrow, make the old final states the start states and the old start state the only final

one.

  • Prefix, suffix and substring: change only the markings, never the transitions.
  • All ten operations stay inside the class. Chapter 51 shows the context free languages do not.
  • To prove a language not regular by closure: derive a known non regular language from it using closed
munotes.in193

Closure Properties of the Regular Languages

operations, and conclude by contradiction. It is usually much shorter than pumping, and it reuses earlier work.

Test yourself

1. Give the construction for the complement and both conditions on the machine. Swap the final and non final states. The machine must be deterministic, so that each string has exactly one run, and complete, so that no string falls off; the dead state added to complete it becomes final in the complement.

2. Prove closure under intersection in one line. L intersect M equals the complement of the union of the complements, and the class is closed under union and complement, so it is closed under intersection.

3. Describe the product construction and say what changes between union, intersection and difference. States are pairs, one component from each machine, the start state is the pair of start states, and a transition moves both components. Only the final states change: both components final for intersection, either for union, and the first final with the second not for difference.

4. Prove by closure that the strings over {a, b} with equally many a and b are not regular. Intersect with the regular language a b. The result is a to the k followed by b to the k, which is not regular. Since the class is closed under intersection, the original cannot be regular.

5. Why does the complement construction fail on a nondeterministic machine? Because a string may have several runs, some ending in a final state and some not. Swapping the markings then changes which runs succeed rather than negating acceptance, so the new machine does not accept the complement.

6. Which of the three prefix, suffix and substring constructions changes a transition? None. All three change only which states are marked start or final, which is why they are the cheapest closure proofs in the chapter.

Contents This chapter on its own page

munotes.in194

Chapter Thirty-Nine

The Decision Problems of the Regular Languages

Syllabus topic Module 1, "Regular Languages: Closure Properties"

In one line

For a regular language given as a machine, every natural question about it can be answered by an algorithm that always terminates.

In the wording a student can write in an examination: a decision problem for a class of languages is a question about a member of the class admitting a yes or no answer. For the regular languages the membership, emptiness, finiteness, equivalence and containment problems are all decidable, and each has an algorithm whose running time is bounded in the size of the machine given.

Why these five and why now

Because each of them is a question a person actually asks, and because the contrast with the later chapters is the whole point of putting them here.

Every question in this chapter is decidable. In chapter 51 the same list for the context free languages has three entries that say no. In chapter 77 the list for type 0 languages says no to almost everything. So this chapter is the baseline against which the negative results of Module 2 are measured, and a student who has seen the easy case first understands what is being lost.

The five problems

Membership. Given a machine M and a string w, is w in L(M)?

Emptiness. Given M, is L(M) empty?

Finiteness. Given M, is L(M) finite?

Equivalence. Given two machines, do they accept the same language?

Containment. Given two machines, is the language of the first contained in that of the second?

Membership

Algorithm. Run the machine on the string. Accept if the last state is final.

Why it terminates. A DFA makes exactly one move per symbol, so the run is as long as the string and no longer.

Cost. One step per symbol, and each step is one table lookup. For an NFA, chapter 14's set of states method keeps a set of at most as many states as the machine has, so the cost is the length of the string multiplied by the number of states.

This is the problem that is decidable most obviously, and it is the one that becomes undecidable last as the classes grow: it is still decidable for context sensitive languages in chapter 61 and fails only for type 0.

Emptiness

Algorithm. Search the machine's graph from the start state. If any final state is reached, the language is not empty; otherwise it is.

Why it terminates. A reachability search visits each state at most once.

Why it is right. A string is accepted exactly when there is a path from the start state to a final state spelling it, so the language is non empty exactly when such a path exists, and any path will do.

The short answer to give: the language is empty exactly when no final state is reachable from the start state.

munotes.in195

The Decision Problems of the Regular Languages

Careful. Emptiness is not "the machine has no final states". A machine may have final states that are unreachable, and its language is then empty anyway. Chapter 20's first step is the same reachability search, which is why minimising a machine for the empty language leaves exactly one state.

Finiteness

Algorithm. Delete the states that are not reachable from the start, and the states from which no final state is reachable. In what remains, look for a cycle. The language is infinite exactly when there is one.

Why it is right. A cycle on a path from the start to a final state can be gone round any number of times, each giving a different accepted string, so the language is infinite. Conversely if there is no such cycle then every accepting path has length less than the number of states, so there are finitely many accepted strings.

Why both deletions are needed. A cycle among unreachable states contributes nothing, and neither does a cycle from which no final state can be reached. Testing for a cycle without pruning first gives the wrong answer, and this is the step that is usually forgotten.

The bound worth quoting. If the machine has n states and the language is non empty, it contains a string of length less than n. If the language is infinite, it contains a string of length at least n and less than 2n. Both follow from the pumping lemma of chapter 36, and together they turn finiteness into a finite check: test every string of length between n and 2n minus 1.

Worked

S1Stateab
startf0f1f1
finalf1dddd
dddddd

Accepts: a, b

Rejects: ε, aa, ab, ba, bb

L(S1) is finite

Prune: every state is reachable, but from dd no final state is reachable, so dd goes. What remains is f0 and f1 with arrows from f0 to f1 and nothing else. No cycle, so the language is finite, and indeed it has two members.

S2Stateab
startg0g1g0
finalg1g1g0

Accepts: a, aa, ba, aba, bba

Rejects: ε, b, ab, bb

L(S2) is infinite

Here both states survive the pruning and there is a loop on g1 labelled a, on a path from the start to a final state, so the language is infinite.

Equivalence

Algorithm one, by minimising. Minimise both machines by chapter 20. Chapter 21 proved the minimal DFA is unique up to the names of its states, so the two languages are equal exactly when the two minimal machines are the same machine with its states renamed. Checking that is a graph isomorphism question, but on a minimal DFA it is settled by walking both machines from their start states together, which takes one pass.

munotes.in196

The Decision Problems of the Regular Languages

Algorithm two, by the product. Build the product machine of chapter 38 with the final states set to the symmetric difference: pairs where exactly one component is final. That machine accepts exactly the strings on which the two disagree, so the two languages are equal exactly when the product's language is empty, which is the emptiness problem already solved.

Algorithm two is the one to know, for three reasons. It reduces a new problem to one already solved, which is the technique the whole of Module 2 rests on. It gives a witness when the answer is no: the shortest string reaching a disagreeing pair is the shortest string telling the two machines apart. And it is what the checker behind this book actually runs: every claim in this book of the form "these two accept the same language" is decided by this algorithm, breadth first, so the witness it reports when a claim is wrong is the shortest one there is.

Containment

Algorithm. The language of A is contained in that of B exactly when

L(A) intersect (complement of L(B)) is empty

Both operations are closures from chapter 38, so the left hand side is a regular language with a machine, and emptiness is already solved. Build the product with the final states set to pairs where the first component is final and the second is not, then test for emptiness.

Equivalence from containment. Two languages are equal exactly when each contains the other, so equivalence can be answered by two containment tests. The product method above answers it in one, which is why it is preferred.

The table

ProblemDecidable?Reduced toCost
membershipyesrun the machinethe length of the string
emptinessyesreachability from the startone pass over the states
finitenessyespruning, then a cycle searchone pass, after pruning
equivalenceyesemptiness of the symmetric difference productthe two state counts multiplied
containmentyesemptiness of the difference productthe two state counts multiplied

Every row says yes, and every row but the first is an emptiness test in disguise. That is the pattern to take away: emptiness is the problem the others reduce to.

Reduction, introduced here on purpose

The word reduction is used above without ceremony, and it is worth naming, because it is the single most important technique in Module 2.

To reduce problem X to problem Y is to give a method that converts any instance of X into an instance of Y whose answer is the same. Then an algorithm for Y gives an algorithm for X: convert, and run.

munotes.in197

The Decision Problems of the Regular Languages

That is what happened four times in this chapter. Equivalence was converted into emptiness by building the symmetric difference product. Containment was converted into emptiness by building the difference product.

And the technique has a mirror image which chapter 75 uses throughout: if X is known to have no algorithm, and X reduces to Y, then Y has no algorithm either, for otherwise X would. Reductions carry decidability forwards and undecidability backwards, and keeping the direction straight is what that chapter is about.

Distinctions

EmptinessFiniteness
Questionis any string accepted?are only finitely many accepted?
Methodreachability from the startprune, then look for a cycle
Pruning needednoyes, both directions
Answer for a machine with no reachable final stateemptyfinite, trivially
Minimise and compareThe symmetric difference product
Gives a witness when unequalnoyes, and the shortest one
Reduces to an earlier problemnoyes, to emptiness
Used by this book's checkernoyes

What it does NOT mean

Emptiness is not about having no final states. It is about having no reachable final state.

Finiteness is not about having no cycles. It is about having no cycle on a path from the start to a final state, which is why both prunings come first.

Decidable does not mean cheap. The product of two machines with a hundred states each has ten thousand, and that is still an algorithm.

These answers do not carry over to bigger classes. Chapter 51 shows equivalence undecidable for context free languages, and chapter 77 proves it. The whole point of this chapter is that the regular case is the easy one.

A reduction is not a proof that two problems are the same. It says an algorithm for one gives an algorithm for the other, in that direction only, and the direction is what chapter 75 warns about.

Quick revision

  • Five problems, all decidable for the regular languages: membership, emptiness, finiteness, equivalence,

containment.

  • Membership: run the machine; the run is as long as the string.
  • Emptiness: is any final state REACHABLE from the start? Not "are there final states".
  • Finiteness: prune the unreachable states and the states that cannot reach a final state, then look for a cycle. A

cycle in what remains means infinite.

  • Equivalence: build the product with the symmetric difference as final states, and test emptiness. It gives the

shortest witness when the answer is no, and it is what this book's checker runs.

  • Containment: L(A) inside L(B) exactly when L(A) intersect the complement of L(B) is empty; build the difference
munotes.in198

The Decision Problems of the Regular Languages

product and test emptiness.

  • Emptiness is the problem the others reduce to, and reduction is the technique Module 2 is built on.
  • If the machine has n states: a non empty language holds a string shorter than n, and an infinite language holds

one of length between n and 2n minus 1.

Test yourself

1. Give the algorithm for emptiness and state it in one sentence. Search the graph from the start state; the language is empty exactly when no final state is reachable.

2. Why must a machine be pruned before testing for a cycle to decide finiteness? Because a cycle among states that cannot be reached from the start, or from which no final state can be reached, contributes no accepted strings. Testing for a cycle without pruning can report an infinite language when the language is finite or even empty.

3. Describe the symmetric difference product and say what it decides. The product machine of the two, with final states the pairs in which exactly one component is final. It accepts exactly the strings on which the two machines disagree, so its emptiness decides equivalence, and the shortest string it accepts is the shortest witness of inequality.

4. Reduce containment to emptiness. L(A) is contained in L(B) exactly when L(A) intersected with the complement of L(B) is empty. Both operations are closures, so build the product with the first component final and the second not, and test emptiness.

5. A machine has 4 states and its language is infinite. What can you say about the lengths of its strings? It contains a string of length less than 4, and a string of length at least 4 and less than 8. Both follow from the pumping lemma, and together they make finiteness a finite check.

6. Define a reduction, and say which way it carries decidability and which way undecidability. A reduction from X to Y converts every instance of X into an instance of Y with the same answer. It carries decidability from Y to X, since an algorithm for Y then solves X; and it carries undecidability from X to Y, since an algorithm for Y would solve X.

Contents This chapter on its own page

munotes.in199

Chapter Forty

Regular Sets and Regular Grammar: Kleene's Theorem Assembled

Syllabus topic Module 1, "Regular Languages: Regular Sets and Regular Grammar"

In one line

Four different ways of describing a language turn out to describe the same collection of languages, and that collection is the regular sets.

In the wording a student can write in an examination: the following four statements about a language L are equivalent. L is accepted by some deterministic finite automaton; L is accepted by some nondeterministic finite automaton; L is denoted by some regular expression; L is generated by some regular grammar. A language satisfying any one of them, and therefore all of them, is called a regular set or a regular language. The equivalence of the second and third statements is Kleene's theorem.

Why this is worth a chapter of its own

Because a theorem with four equivalent parts is proved by a cycle of constructions, and the point is easy to miss while the constructions are being learned one at a time.

Each chapter of this block proved one arrow. This chapter shows that the arrows join up into a circle, so that any description can be turned into any other, and therefore the four notions collapse into one.

It is also the chapter where the word regular set is finally defined. MU uses "regular set" and "regular language" interchangeably, as does this book, and the reason there are two names is historical: Kleene called them regular events, the literature turned that into regular sets, and the modern habit is regular languages.

The four statements and the arrows between them

FromToByChapter
a regular expressionan NFA with empty movesThompson's construction34
an NFA with empty movesan NFAthe closure method17
an NFAa DFAthe subset construction16
a DFAa regular expressionArden, or state elimination35
a DFAa regular grammarone variable per state30
a regular grammaran NFAone state per variable30
a DFAan NFAtrivially, singleton cells14

Seven arrows, and they are more than enough: any of the four descriptions can be turned into any other by following them.

The cycle, in the order the chapters built it

Read this as one proof in four steps, because that is what it is.

Step 1. An expression gives a machine. Chapter 34's six rules turn any regular expression into an NFA with empty moves. So every language denoted by an expression is accepted by an NFA.

Step 2. Any machine gives a deterministic machine. Chapter 17 removes the empty moves and chapter 16 removes the nondeterminism, and neither changes the language. So every language accepted by any finite automaton is accepted by a DFA.

Step 3. A deterministic machine gives an expression. Chapter 35's equation method or state elimination turns any machine into a regular expression denoting the same language. So every language accepted by a DFA is denoted by an expression.

munotes.in200

Regular Sets and Regular Grammar: Kleene's Theorem Assembled

Step 4. And the grammar joins in. Chapter 30 converts a DFA into a right linear grammar and a right linear grammar into an NFA, both directly. So the fourth description sits on the same cycle.

Steps 1, 2 and 3 form a closed loop through three of the descriptions, and step 4 attaches the fourth to it. Since every arrow preserves the language, following the loop from any starting point and back proves all four descriptions equivalent.

The whole thing on one machine

Here is one language in all four descriptions, so the theorem is concrete rather than a diagram. The language is the strings over {a, b} ending in ab.

As a deterministic finite automaton:

T1Stateab
starts0s1s0
s1s1s2
finals2s1s0

Accepts: ab, aab, bab, abab, bbab

Rejects: ε, a, b, ba, abb, aa

T1 is deterministic

As a nondeterministic finite automaton, which needs only two states because it may guess where the final ab begins:

T2Stateab
startu0{u0, u1}u0
u1-u2
finalu2--

Accepts: ab, aab, bab, abab, bbab

Rejects: ε, a, b, ba, abb, aa

T2 is nondeterministic

L(T2) = L(T1)

As a regular expression:

(a + b)* a b

L(T3) = L(T1)

As a regular grammar, right linear, one variable per state of T1 and an epsilon rule on the final state:

S->aA | bS
A->aA | bB
B->aA | bS | ε

Accepts: ab, aab, bab, abab, bbab

Rejects: ε, a, b, ba, abb, aa

L(T4) = L(T1)

Four descriptions of one language, and three exact comparisons against the first. That is Kleene's theorem shown rather than stated, and every comparison is decided over every string there is.

What the theorem is worth

Three things follow, and each is used later.

A definition that does not depend on a description. "Regular" can be defined by any one of the four and means the same thing. So a proof may pick whichever description is convenient, and the proofs in this book do exactly that: chapter 38 proves closure under union from the expression and closure under complement from the machine, because each is easy in one description and awkward in the other.

A pipeline. A pattern written by a person becomes a machine a computer runs, and chapters 34, 17, 16 and 20 are the four stages of it. That is not an illustration; it is what happens when a regular expression is compiled.

munotes.in201

Regular Sets and Regular Grammar: Kleene's Theorem Assembled

A boundary. Because all four descriptions give the same class, a language shown to be outside the class by any method is outside all four. So when chapter 37 proves that a to the k followed by b to the k is not regular, it has proved at once that there is no machine for it, no expression for it, and no regular grammar for it. One proof, four consequences.

Where the name and the result came from

Kleene's RAND memorandum RM-704 is titled "Representation of Events in Nerve Nets and Finite Automata", and its summary states the question it sets out to answer: to what kinds of events can a nerve net respond by firing a certain neuron, and more generally, to what kinds of events can any finite automaton respond by assuming one of certain states. The summary dates the investigations it reports to August 1951.

Two things about that are worth noticing. The question is posed about nerve nets, following McCulloch and Pitts, so the finite automaton arrived in this literature as a model of a neuron and not of a computer. And the word Kleene used for what we call a language was event, which is why the class is called the regular sets rather than the regular languages in the older books, MU's included.

Chapter 26's Chomsky classification arrived from a third direction, linguistics, a few years later, and the identification of Kleene's regular events with Chomsky's type 3 grammars is the fourth arrow of this chapter. Three fields, three vocabularies, one class of languages.

Distinctions

DescriptionBest forWorst for
DFAmembership, complement, minimisingwriting down by hand
NFAdesigning, and the constructionsrunning
regular expressionwriting, reading, union and concatenationcomplement and intersection
regular grammarconnecting to the rest of the hierarchyeverything else
Kleene's theoremThe whole four part equivalence
Statesexpressions and finite automata describe the same languagesall four descriptions do
Proved bychapters 34 and 35those, plus 16, 17 and 30

What it does NOT mean

The four descriptions are not equally convenient. They describe the same class and are wildly different to work with, which is the table above.

The conversions do not preserve size. An expression becomes a machine about twice as large, and an NFA can become a DFA exponentially larger. Only the language is preserved.

A regular set is not a set of regular things. It is a language, and the older word for a language in this literature was event, which is where the mismatch of vocabulary comes from.

The theorem does not say the conversions are cheap. Chapter 16's bound is exponential and is achieved.

munotes.in202

Regular Sets and Regular Grammar: Kleene's Theorem Assembled

It does not extend upwards. Chapter 56 shows that the corresponding statement for pushdown automata fails, because a deterministic pushdown automaton is strictly weaker than a nondeterministic one.

Quick revision

  • Four equivalent statements: accepted by a DFA; accepted by an NFA; denoted by a regular expression; generated by a

regular grammar. A language satisfying any is a regular set, also called a regular language.

  • The equivalence of expressions and machines is Kleene's theorem.
  • The proof is a cycle: expression to NFA by chapter 34; NFA to DFA by chapters 17 and 16; DFA to expression by

chapter 35; and the grammar joins the cycle by chapter 30, both ways.

  • Every arrow preserves the language, so following the cycle proves all four equivalent.
  • What it buys: a definition independent of description, a compilation pipeline, and one irregularity proof serving

all four descriptions.

  • Kleene's memorandum RM-704 asked the question about nerve nets, and called a language an event, which is why the

class is called the regular sets in the older books.

  • The conversions preserve the language and not the size; chapter 16's blow up is exponential and achieved.

Test yourself

1. State the four equivalent conditions. L is accepted by a DFA; L is accepted by an NFA; L is denoted by a regular expression; L is generated by a regular grammar. Any one implies all the others.

2. Which pair of them is Kleene's theorem? That a language is denoted by a regular expression exactly when it is accepted by a finite automaton, proved by chapter 34 in one direction and chapter 35 in the other.

3. Name the four constructions that close the cycle. Thompson's construction from an expression to an NFA with empty moves; the closure method removing the empty moves; the subset construction giving a DFA; and Arden's method or state elimination giving an expression back.

4. Why does one proof that a language is not regular settle four questions at once? Because the four descriptions define the same class, so a language outside the class has no machine of either kind, no expression and no regular grammar.

5. Convert this machine to a right linear grammar: p start, q final, p on a to q, p on b to p, q on a to q, q on b to p. P->aQ | bP and Q->aQ | bP | ε. One production per transition, and the epsilon rule on the variable for the final state.

6. Does the four part equivalence hold for pushdown automata? No. A deterministic pushdown automaton is strictly weaker than a nondeterministic one, so the analogue of the first two statements fails, and chapter 57 gives the witness language.

Contents This chapter on its own page

munotes.in203

Chapter Forty-One

Context Free Grammars and Context Free Languages

Syllabus topic Module 1, "Context Free Languages: Context-free Languages"

In one line

A context free grammar allows any right hand side at all, provided the left hand side is a single variable, and the languages it can generate include ones no finite automaton can accept.

In the wording a student can write in an examination: a context free grammar is a quadruple G = (V, T, P, S) in which every production has the form A to alpha with A a single variable in V and alpha any string over V union T. A language is context free if it is generated by some context free grammar. Every regular language is context free and the converse is false.

What the relaxed restriction buys

Chapter 30's regular grammar allowed at most one variable, at a fixed end. Drop both restrictions and two things become possible at once, and they are the two things a finite automaton cannot do.

A variable may have terminals on both sides. The rule S to aSb adds an a on the left and a b on the right in one step, which ties the two counts together. Chapter 30 proved that this single change escapes the regular languages.

A right hand side may hold several variables. The rule S to SS or S to AB puts two independent sub problems into the form at once, and they can be expanded in either order. That is what makes a stack rather than a state the right memory for this class, as chapter 29 explained.

So the class grows, and it grows exactly as far as the pushdown automaton of chapter 52 reaches. That equality is chapters 58 and 59.

The first language a finite automaton cannot do

S->aSb | ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

Two productions, and the language is a to the n followed by b to the n, which chapters 21 and 37 each proved is not regular, by different methods. So this grammar generates something outside the whole of Module 1 up to chapter 40, and it does so in two lines.

That is the single most useful fact to hold on to about this class: it can match one pair of counts. It cannot match two, which chapter 50 proves, and that is exactly the boundary between this class and the next.

Why programming languages live here

Because the shapes that a finite automaton cannot check are exactly the shapes a programming language is full of.

Balanced brackets. Every opening bracket has a matching closing one, and the matching is nested rather than merely counted.

S->(S) | SS | ε

Accepts: ε, (), (()), ()(), (()())

munotes.in204

Context Free Grammars and Context Free Languages

Rejects: (, ), )(, ((), ())

A finite automaton cannot do this, by the argument of chapter 37 applied to the depth. A context free grammar does it in three productions, and every compiler's parser is doing this at bottom.

Arithmetic expressions. The grammar below is the one every textbook on compilers begins with, and it is worth seeing here because chapter 44 will use it to explain ambiguity and chapter 42 to explain the derivation tree.

E->E+T | T
T->T*F | F
F->(E) | i

Accepts: i, i+i, ii, i+ii, (i+i)*i

Rejects: ε, +, i+, +i, i**i, (i

Read the variables: E is an expression, T is a term, F is a factor, and i stands for any identifier or number. The three levels are what make multiplication bind tighter than addition, and chapter 44 shows what happens to a grammar that tries to do without them.

Every regular language is context free

Proof. A regular grammar is a context free grammar: every production of a right linear grammar has a single variable on the left, which is all the context free restriction requires. So a language with a regular grammar has a context free grammar.

That gives the containment. Chapter 26 said it is proper, and the witness is the language of U1 above.

The class is genuinely larger, and here is a second witness

The palindromes over {a, b}, which chapter 37 proved not regular:

S->aSa | bSb | a | b | ε

Accepts: ε, a, b, aa, aba, abba, ababa

Rejects: ab, ba, aab, abb, aabb

Three terminating rules, and chapter 25 explained why all three are needed. The machine that accepts this language is built in chapter 56, and it is the machine that also shows nondeterminism matters for pushdown automata.

What a context free language can and cannot do

This table is worth having before the rest of the block, because every chapter from here to 51 fills in a row of it.

TaskContext free?Where
match one pair of countsyesthis chapter, U1
nest brackets to any depthyesthis chapter, U2
enforce operator precedenceyesthis chapter, U3
recognise palindromesyesthis chapter, U4
match two pairs of counts at oncenochapter 50
require a repeated block, w then wnochapter 50
intersect two context free languagesnot in generalchapter 51
complement a context free languagenot in generalchapter 51

The first four are what a stack does. The last four are what a stack cannot do, and they are why Module 2 needs a stronger machine.

munotes.in205

Context Free Grammars and Context Free Languages

The notation, and the things students get wrong

Only the left hand side is restricted. A right hand side may be as long as you like, may mix terminals and variables freely, and may be empty.

A context free grammar may be ambiguous. Nothing in the definition forbids it, and chapter 44 is about it. U2 above is ambiguous, and so is U3 if the level structure is removed.

A language may have many grammars. U2's language, the balanced brackets, also has the unambiguous grammar

S->(S)S | ε

Accepts: ε, (), (()), ()(), (()())

Rejects: (, ), )(, ((), ())

L(U5) = L(U2)

U2 is ambiguous

U5 is unambiguous

Two grammars for one language, one ambiguous and one not, and the checker decides all three claims: that the languages agree, that the first is ambiguous by finding a string with two derivations, and that the second is not, over every string up to its bound. That pair is the whole subject of chapter 44 in miniature.

Context free is a property of the language, not of a description. A regular language is context free, and asking whether a grammar is context free is a question about its productions, which is chapter 26's classification.

Distinctions

Regular grammarContext free grammar
Left hand sideone variableone variable
Right hand sidea terminal, then at most one variable, at a fixed endanything
Variables outstanding in a formexactly oneany number
Memory the machine needsa statea stack
Can match a pair of countsnoyes
Context free languageRegular language
Every one iscontext free
Closed under intersectionnoyes
Closed under complementnoyes
Membership decidableyes, by CYKyes, by running the machine
Equivalence decidablenoyes

What it does NOT mean

Context free does not mean unambiguous. Ambiguity is a property of a grammar and most context free grammars are ambiguous. Chapter 44 is about it.

Context free does not mean easy to parse. It means a pushdown automaton exists. Parsing efficiently is a separate subject, and CYK in chapter 51 is the general method.

The restriction is on the left hand side only. A long or complicated right hand side is perfectly context free.

Not every context free language has an unambiguous grammar. Some are inherently ambiguous, which chapter 44 defines.

A context free grammar cannot match two pairs of counts. One pair only. That is the boundary, and chapter 50 proves it.

Quick revision

  • A context free grammar restricts only the left hand side: one variable. The right hand side is anything.
  • A language is context free if some context free grammar generates it.
  • Every regular language is context free, because a regular grammar already satisfies the restriction, and the
munotes.in206

Context Free Grammars and Context Free Languages

containment is proper with witness a to the n b to the n.

  • The relaxed restriction buys two things: terminals on both sides of a variable, which ties a pair of counts, and

several variables in a form, which needs a stack.

  • What the class does: one pair of matched counts, nesting to any depth, operator precedence, palindromes.
  • What it does not do: two pairs of counts at once, a repeated block, intersection, complement.
  • A language may have several grammars, some ambiguous and some not.

Test yourself

1. Give the definition of a context free grammar, and say what is NOT restricted. A quadruple (V, T, P, S) in which every production has a single variable as its left hand side. The right hand side is unrestricted: any string of variables and terminals, including the empty string.

2. Prove that every regular language is context free. A right linear grammar has a single variable on each left hand side, which is all the context free restriction demands. So a regular grammar is already a context free grammar, and the language it generates is context free.

3. Write a context free grammar for the balanced bracket strings, and give a second one that is unambiguous. S->(S) | SS | ε is the obvious one and is ambiguous. S->(S)S | ε generates the same language and is unambiguous, because the position of the first matching close bracket is forced.

4. What two things does the relaxed restriction make possible, and which machine do they need? Terminals on both sides of a variable, which ties two counts together; and several variables in one sentential form, which must be expanded in a definite order. The second is why the machine needs a stack rather than finitely many states.

5. Name two things a context free language cannot do. Match two pairs of counts at once, as in a to the n b to the n c to the n; and require a repeated block, as in w followed by w. Both are proved in chapter 50.

6. Is a regular language context free? Is a context free grammar regular? Every regular language is context free. A context free grammar is regular only if it happens to satisfy the tighter right linear or left linear restriction, which most do not.

Contents This chapter on its own page

munotes.in207

Chapter Forty-Two

The Derivation Tree

Syllabus topic Module 1, "Context Free Languages: Derivation Tree"

In one line

A derivation tree records which production was applied to which symbol, and forgets the order in which the applications were made.

In the wording a student can write in an examination: a derivation tree, also called a parse tree, for a context free grammar G is an ordered tree in which the root is labelled with the start symbol, every interior vertex is labelled with a variable, every leaf is labelled with a terminal or with epsilon, and if an interior vertex labelled A has children labelled X1 to Xk from left to right then A to X1 ... Xk is a production of G. The yield of the tree is the string of leaf labels read left to right.

Why a tree as well as a derivation

Because a derivation records too much.

Chapter 23 showed that one string can have many derivations differing only in the order the variables were taken. For the grammar S to AB with A to a and B to b, the string ab has two derivations: expand A first, or expand B first. They are different sequences and they say exactly the same thing about the structure of ab.

The tree throws away the order and keeps the structure. So two derivations that differ only in order give the same tree, and two derivations that differ in which production was applied to which symbol give different trees. That is precisely the distinction ambiguity is about, and it is why chapter 44 is defined on trees and not on derivations.

The tree is also what a compiler actually builds. A parser's output is not a sequence of rewriting steps; it is a tree, and the next stage of the compiler walks it.

The four conditions

Each is a condition on the labels, and an answer that drops one has not defined the tree.

The root is the start symbol.

Every interior vertex is labelled with a variable. A terminal can never have children, because no production rewrites a terminal.

Every leaf is labelled with a terminal or with epsilon. A leaf labelled epsilon is the child of a variable that was rewritten by a null production, and it is the only child of that vertex.

The children of a vertex spell a production's right hand side, in order. If A has children X1 to Xk left to right, then A to X1 ... Xk must be a production. The order matters, which is why the tree is called ordered.

A tree satisfying all four is a derivation tree. A tree whose root is some other variable A rather than S is called a subtree or an A-tree, and it is what the induction in chapter 50's proof works with.

munotes.in208

The Derivation Tree

The yield

The yield of a tree is the string obtained by reading its leaves from left to right, with epsilon leaves contributing nothing.

That definition is what connects the tree to the language: the strings of L(G) are exactly the yields of the derivation trees of G. A tree whose yield is w is a proof that w is in the language, and finding such a tree is what parsing means.

Worked: one tree, read three ways

Take the grammar for a to the n followed by b to the n.

S->aSb | ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abab

The derivation of aabb, from chapter 23:

S
aSb
aaSbb
aabb

The tree, written with indentation, one vertex per line, children indented under their parent:

S
  a
  S
    a
    S
      ε
    b
  b

Read it against the four conditions. The root is S. The interior vertices are the three S. The leaves are a, a, epsilon, b, b, all terminals or epsilon. And each S's children spell a right hand side: the outer two spell a, S, b, which is the first production, and the innermost spells epsilon, which is the second.

The yield is read off the leaves left to right: a, a, epsilon, b, b, and the epsilon contributes nothing, so the yield is aabb. Which is the string the derivation produced, as it must be.

Worked: building the tree from a derivation

The procedure, and it is mechanical.

  1. Write the start symbol as the root.
  2. Take the derivation's first step. It rewrote some occurrence of some variable; find the vertex for that

occurrence and give it one child per symbol of the right hand side, in order.

  1. Repeat for each step, always attaching to the vertex for the occurrence that step rewrote.
  2. When the derivation ends, every leaf is a terminal or epsilon.

Take a grammar with two variables so the occurrences are worth tracking.

S->AB
A->aA | a
B->bB | b

Accepts: ab, aab, abb, aabb, aaabbb

Rejects: ε, a, b, ba, bab

A derivation of aabb:

S
AB
aAB
aaB
aabB
aabb

The tree it builds:

S
  A
    a
    A
      a
  B
    b
    B
      b

Step by step: the first step gives S the children A and B. The second rewrites A as aA, so A gets children a and A. The third rewrites the inner A as a, so it gets one child a. The fourth rewrites B as bB. The fifth rewrites the inner B as b.

The yield is a, a, b, b, which is aabb.

munotes.in209

The Derivation Tree

Notice what the tree shows and the string does not: that the two a belong to A and the two b belong to B. That structure is the whole reason a compiler wants a tree.

Worked: reading a tree back into a derivation

The reverse, which MU's papers also set.

Take this tree, for the grammar V2 above:

S
  A
    a
  B
    b
    B
      b

The yield is a, b, b, which is abb.

A derivation, obtained by choosing an order in which to expand the interior vertices. Expanding leftmost first:

S
AB
aB
abB
abb

And there are other derivations of the same tree, differing only in the order. Chapter 43 counts them.

The tree and the expression grammar

The tree is what makes operator precedence visible, so here is the arithmetic grammar of chapter 41 with a tree.

E->E+T | T
T->T*F | F
F->(E) | i

Accepts: i, i+i, ii, i+ii, (i+i)*i

Rejects: ε, +, i+, +i, i**i, (i

The tree for i+i*i:

yield i+i*i
E
  E
    T
      F
        i
  +
  T
    T
      F
        i
    *
    F
      i

Read the shape rather than the labels. The top of the tree is a plus, with i on its left and the whole of i*i on its right. So the tree says: add i to the product of i and i. It does not say: multiply the sum of i and i by i.

That is precedence, and it is a fact about the shape of the tree, not about the string. The three levels E, T, F are what force it: a plus can only appear at an E vertex, a times only at a T vertex, and a T is below an E, so a plus is always higher in the tree than a times. Chapter 44 shows what happens to a grammar that tries to do without the levels.

Distinctions

A derivationA derivation tree
Recordsthe order of the stepswhich production was applied where
Two of them differ whenthe order differs, or the productions dothe productions differ
Number for one stringmanyone, if the grammar is unambiguous
What a parser producesnoyes
Ambiguity is defined onnoyes
An interior vertexA leaf
Labela variablea terminal, or epsilon
Has childrenyes, spelling a production's right hand sideno
Contributes to the yieldnoyes, unless it is epsilon

What it does NOT mean

A tree is not a derivation. It is what many derivations have in common, and the order is gone.

Two derivations of one string do not make a grammar ambiguous. Two trees do. V2 above has several derivations of abb and exactly one tree, and it is unambiguous.

munotes.in210

The Derivation Tree

The yield is not the label of the root. It is the leaves, read left to right.

An epsilon leaf is not nothing. It is a leaf, it is the only child of its parent, and it records that a null production was used. Omitting it makes the tree violate the fourth condition, because its parent would then have no children at all.

A terminal never has children. Interior vertices are variables, always, because only a variable can be rewritten.

Quick revision

  • A derivation tree has the start symbol at its root, variables at every interior vertex, terminals or epsilon at

every leaf, and each vertex's children spelling a production's right hand side in order.

  • The yield is the leaves read left to right, epsilon leaves contributing nothing. The strings of L(G) are exactly

the yields of the derivation trees.

  • The tree records which production went where and forgets the order, which is why ambiguity is defined on trees.
  • Building the tree from a derivation: attach children to the vertex for the occurrence each step rewrote.
  • Reading a derivation from the tree: choose an order for expanding the interior vertices.
  • Precedence is a fact about the SHAPE of the tree, forced by the level structure of the grammar.
  • A subtree rooted at a variable A is an A-tree, and it is what the induction in the CFG pumping lemma works with.

Test yourself

1. Give the four conditions defining a derivation tree. The root is labelled with the start symbol; every interior vertex is labelled with a variable; every leaf is labelled with a terminal or epsilon; and if a vertex labelled A has children X1 to Xk left to right then A to X1 ... Xk is a production.

2. What is the yield, and what does it have to do with the language? The string of leaf labels read left to right, epsilon leaves contributing nothing. L(G) is exactly the set of yields of derivation trees of G, so a tree with yield w proves w is in the language.

3. For S->aSb | ε, draw the tree for ab. S at the root with three children: a, then S, then b. The inner S has one child, epsilon. The yield is a, epsilon, b, which is ab.

4. Why is ambiguity defined on trees rather than on derivations? Because one tree corresponds to many derivations, which differ only in the order the variables were taken. Counting derivations would make almost every grammar ambiguous; counting trees counts genuinely different structures.

5. Can an interior vertex be labelled with a terminal? Can a leaf be labelled with a variable? No to both. Only a variable can be rewritten, so only a variable can have children; and a finished tree has no variable left to rewrite, so no leaf carries one.

munotes.in211

The Derivation Tree

6. In the arithmetic grammar, what makes a plus sit higher in the tree than a times? The level structure. A plus can only appear at an E vertex and a times only at a T vertex, and every T vertex is below some E vertex, so any plus is nearer the root than any times in the same expression. That is what encodes precedence.

Contents This chapter on its own page

munotes.in212

Chapter Forty-Three

Leftmost and Rightmost Derivations

Syllabus topic Module 1, "Context Free Languages: Derivation Tree"

In one line

A leftmost derivation always rewrites the leftmost variable and a rightmost derivation always the rightmost, and each tree has exactly one of each.

In the wording a student can write in an examination: a derivation is leftmost if at every step the variable rewritten is the leftmost variable in the current sentential form, and rightmost if it is the rightmost. Every derivation tree of a context free grammar has exactly one leftmost and exactly one rightmost derivation, and conversely every leftmost derivation determines a unique tree.

Why impose a discipline at all

Three reasons, and the third is the one that matters for the next chapter.

Because a derivation has too much freedom. Chapter 42 showed that one tree gives many derivations, differing only in the order the variables were taken. Fixing the order picks one of them, which makes the derivation a canonical object rather than an arbitrary one.

Because a parser has to choose. A program building a tree cannot expand variables in an arbitrary order; it proceeds in a definite one. Expanding leftmost first is what a top down parser does, and it corresponds to reading the input left to right and predicting what comes next. Recognising rightmost, in reverse, is what a bottom up parser does.

Because ambiguity has to be counted. If ambiguity were defined by counting derivations, every grammar with two variables in a right hand side would be ambiguous, which would make the notion useless. Defining it by counting trees is right, and the theorem below is what lets trees be counted by counting leftmost derivations, which is something a machine can do.

The two definitions

A derivation is leftmost when, at every step, the variable rewritten is the leftmost variable of the current sentential form.

A derivation is rightmost when, at every step, it is the rightmost variable.

Note what is fixed and what is not: the position is fixed, the production is not. At each step there may be several productions for the variable in question, and a leftmost derivation may use any of them. Different choices give different leftmost derivations, and that is exactly what ambiguity is.

Worked: one tree, both derivations

Take the grammar of chapter 42 with two variables.

S->AB
A->aA | a
B->bB | b

Accepts: ab, aab, abb, aabb, aaabbb

Rejects: ε, a, b, ba, bab

The tree for aabb:

S
  A
    a
    A
      a
  B
    b
    B
      b

The leftmost derivation. At each step, take the leftmost variable.

S
AB
aAB
aaB
aabB
aabb

Line 3 is the step to watch: the form is aAB, its variables are A and B, and A is the leftmost, so A is rewritten.

munotes.in213

Leftmost and Rightmost Derivations

The rightmost derivation. At each step, take the rightmost variable.

S
AB
AbB
Abb
aAbb
aabb

Line 3 here: the form is AB, its rightmost variable is B, so B is rewritten, and A is left alone until the b are finished.

Five steps each, and they are the same five steps in a different order. Both use the productions S to AB, A to aA, A to a, B to bB, B to b, exactly once each. That is not a coincidence: both derivations are readings of the same tree, and the tree records which productions were used.

The checker proves the discipline of each, so a derivation labelled leftmost that rewrites the wrong occurrence fails the build.

The theorem

Statement. For a context free grammar G and a string w in L(G), the derivation trees of w, the leftmost derivations of w, and the rightmost derivations of w are in one to one correspondence. In particular each tree has exactly one leftmost derivation and exactly one rightmost derivation.

Proof that each tree has at least one leftmost derivation. Expand the interior vertices of the tree in the following order: always the leftmost vertex that still has an unexpanded variable label. That produces a derivation, and by construction it is leftmost.

Proof that each tree has at most one. Suppose two leftmost derivations came from the same tree. At the first step where they differ, both must rewrite the same occurrence, because both are leftmost and the forms up to that point agree. So they must differ in the production used. But the tree fixes which production was applied at each vertex, and the occurrence in question corresponds to a definite vertex. So they cannot differ. Hence there is exactly one.

And conversely a leftmost derivation determines a tree, by chapter 42's construction: attach children to the vertex for the occurrence each step rewrote.

The rightmost case is the mirror image, word for word.

The consequence used in chapter 44: two different leftmost derivations of the same string mean two different trees, which is the definition of ambiguity. So ambiguity can be detected by counting leftmost derivations, and counting is something a program can do.

Worked: two leftmost derivations, and hence two trees

The grammar that shows it, and the one MU's 2022 paper sets.

S->SbS | a

Accepts: a, aba, ababa, abababa

Rejects: ε, b, ab, ba, abab, aa

The string ababa has two leftmost derivations. Here is the first, which builds the plus on the left:

S
SbS
abS
abSbS
ababS
ababa

And the second, which builds it on the right:

S
SbS
SbSbS
abSbS
ababS
ababa
munotes.in214

Leftmost and Rightmost Derivations

Both are leftmost, and the checker proves it for both. They differ at the second step: the first rewrites the leftmost S by the production S to a, and the second rewrites it by S to SbS.

Their trees, which are therefore different:

S
  S
    a
  b
  S
    S
      a
    b
    S
      a
S
  S
    S
      a
    b
    S
      a
  b
  S
    a

The first groups the last two a together; the second groups the first two. Both yield ababa, both satisfy all four conditions of chapter 42, and they are different trees.

W2 is ambiguous

So the grammar is ambiguous, and the checker confirms it by finding two derivations for itself.

Which parser uses which

Worth a paragraph because it is why the two names exist outside this paper.

Leftmost, forwards, is top down parsing. Begin with the start symbol and predict: given the next input symbols, which production should the leftmost variable use? Recursive descent and the LL family of parsers work this way, and the L in LL stands for the leftmost derivation being produced.

Rightmost, backwards, is bottom up parsing. Begin with the input and reduce: find a right hand side in what you have and replace it by its left hand side, until the start symbol is reached. Reading such a reduction sequence backwards gives a rightmost derivation, which is why the LR family produces one, and the R stands for it.

Neither is in MU's syllabus beyond the definitions, and both are the reason the definitions are set.

Distinctions

LeftmostRightmost
Rewritesthe leftmost variablethe rightmost variable
Number per treeexactly oneexactly one
Parser that produces ittop down, LLbottom up, LR, in reverse
Fixesthe position, not the productionthe position, not the production
Two derivations of one stringTwo LEFTMOST derivations of one string
Meanspossibly nothing; the order may differtwo trees, hence ambiguity
Grammar W1 abovehas severalhas exactly one
Grammar W2 abovehas severalhas two

What it does NOT mean

A leftmost derivation is not unique to a string. It is unique to a tree. A string with two trees has two leftmost derivations, which is the whole of chapter 44.

The discipline fixes the position, not the production. At each step the leftmost variable is rewritten, and which of its productions is used is still a free choice.

Leftmost is not shorter or better. Both derivations of a given tree have exactly the same number of steps, one per interior vertex, and use exactly the same multiset of productions.

Rightmost is not a derivation read backwards. It is a derivation, from the start symbol to the string, in which the rightmost variable is always the one rewritten. It is a bottom up parser's REDUCTION that is a rightmost derivation read backwards.

munotes.in215

Leftmost and Rightmost Derivations

Quick revision

  • Leftmost: always rewrite the leftmost variable. Rightmost: always the rightmost. The position is fixed, the

production is not.

  • Every derivation tree has exactly one leftmost and exactly one rightmost derivation, and each determines the tree.
  • Both derivations of a tree have the same number of steps and use the same productions, in a different order.
  • Two different LEFTMOST derivations of one string mean two different trees, which is ambiguity, and that is what

makes ambiguity countable by machine.

  • Top down parsers produce a leftmost derivation, bottom up parsers a rightmost one in reverse, and that is where LL

and LR get their second letters.

Test yourself

1. Define both disciplines, and say what they do NOT fix. Leftmost rewrites the leftmost variable of the current form at every step; rightmost the rightmost. Neither fixes which production is used, so several derivations of each kind may exist for one string.

2. How many leftmost derivations does a tree have, and why? Exactly one. At least one, by expanding the leftmost unexpanded vertex repeatedly; at most one, because two would have to differ in a production at some vertex, and the tree fixes the production at every vertex.

3. For S->AB, A->a, B->b, give both derivations of ab and say whether the grammar is ambiguous. Leftmost: S, AB, aB, ab. Rightmost: S, AB, Ab, ab. Two derivations, one tree, so the grammar is unambiguous.

4. Why would defining ambiguity by counting derivations be useless? Because almost every grammar with two variables in a right hand side has several derivations of the same string, differing only in the order the variables were taken. Every such grammar would be called ambiguous and the notion would distinguish nothing.

5. Give the two leftmost derivations of ababa for S->SbS | a, and say where they differ. Both begin S, SbS. The first then rewrites the leftmost S by S to a, giving abS, and continues. The second rewrites it by S to SbS, giving SbSbS, and continues. They differ at the second step, in the production chosen.

6. Which parser produces a rightmost derivation, and in what order? A bottom up parser, of the LR family. It reduces from the input towards the start symbol, and reading its sequence of reductions backwards gives a rightmost derivation.

Contents This chapter on its own page

munotes.in216

Chapter Forty-Four

Ambiguity of a Grammar

Syllabus topic Module 1, "Context Free Languages: Ambiguity of Grammar"

In one line

A grammar is ambiguous when some string it generates has two different derivation trees.

In the wording a student can write in an examination: a context free grammar G is ambiguous if there exists a string in L(G) having two or more distinct derivation trees, equivalently two or more distinct leftmost derivations. A context free language is inherently ambiguous if every grammar generating it is ambiguous.

Why ambiguity matters

Because a tree is a meaning, and two trees are two meanings.

Chapter 42 showed that the tree for i+i*i says "add i to the product" and not "multiply the sum by i". If a grammar gives that string two trees, then it does not say which of the two is meant, and a compiler using it would be free to compute either. So ambiguity in a programming language's grammar is a defect, and removing it is part of designing the language.

It matters in this paper for a second reason: chapter 77 proves that deciding whether an arbitrary context free grammar is ambiguous is undecidable. So there is no general test, and showing a particular grammar ambiguous is done by exhibiting two trees, which is always possible when it is true and is the only method there is.

How to show a grammar ambiguous

The method is short and there is only one.

Find a string with two derivation trees, and show both.

That is the whole answer. Three things follow, and each loses marks if ignored.

Two derivations are not enough. They must be two trees, or equivalently two leftmost derivations. Chapter 43 explained why: one tree has many derivations differing only in order, and counting those would make almost every grammar ambiguous.

An argument is not enough. "The grammar is ambiguous because S appears twice on the right" is not a proof. Some grammars with a variable twice on the right are unambiguous. The two trees must be exhibited.

The string must be in the language. Two trees for a string the grammar does not generate proves nothing, and it cannot happen, since a tree's yield is generated by construction.

MU's own grammar, worked

S->SbS | a

Accepts: a, aba, ababa, abababa

Rejects: ε, b, ab, ba, abab, aa

The string. The shortest string with two trees is ababa, of length 5. The strings a and aba have one tree each, so the shortest is not the shortest string of the language.

The first tree, which puts the first b at the root:

S
  S
    a
  b
  S
    S
      a
    b
    S
      a

The second tree, which puts the second b at the root:

munotes.in217

Ambiguity of a Grammar

S
  S
    S
      a
    b
    S
      a
  b
  S
    a
Two derivation trees for ababa under the grammar S to SbS or a

Figure 44.1 The two trees for ababa. They have the same yield and different shapes, which is what ambiguity means.

The two leftmost derivations, one per tree, which is the other acceptable way to present the answer:

S
SbS
abS
abSbS
ababS
ababa
S
SbS
SbSbS
abSbS
ababS
ababa

Therefore ababa has two distinct derivation trees, and G is ambiguous.

X1 is ambiguous with ababa

Both trees are checked against the grammar, both derivations are checked for the leftmost discipline, and the ambiguity claim is checked by the tool finding two derivations independently. So this answer is verified three ways over.

What the ambiguity means here

Read the grammar as arithmetic with b as an operator. The first tree computes a, then b, then (a b a). The second computes (a b a), then b, then a. If b were subtraction those two are different numbers, and the grammar does not say which is meant. That is the practical content of ambiguity, and it is why the next section matters.

Removing ambiguity

Sometimes possible, and the standard technique is forcing a level structure.

The ambiguity in X1 comes from b being able to appear at any depth. Force it to associate one way and the ambiguity goes:

S->a b S | a

Accepts: a, aba, ababa, abababa

Rejects: ε, b, ab, ba, abab, aa

L(X2) = L(X1)

X2 is unambiguous

The same language, and now every string has exactly one tree, because the grammar can only build the string from the left: read an a, then a b, then the rest. The checker confirms both claims.

That is the general recipe. Decide the shape you want and write the grammar so that no other shape is possible. The arithmetic grammar of chapter 41 is the same idea carried further:

E->E+T | T
T->T*F | F
F->(E) | i

Accepts: i, i+i, ii, i+ii, (i+i)*i, i+i+i

Rejects: ε, +, i+, +i, i**i, (i

X3 is unambiguous

Three levels, and each operator can appear at exactly one of them, so the shape of every tree is forced. Compare with the grammar that tries to do without them:

E->E+E | E*E | (E) | i

Accepts: i, i+i, ii, i+ii, (i+i)*i, i+i+i

Rejects: ε, +, i+, +i, i**i, (i

L(X4) = L(X3)

X4 is ambiguous with i+i*i

Same language, and ambiguous, which the checker proves by finding two derivations. The string i+i*i has two trees in X4: one where the plus is at the root, giving addition of i to a product, and one where the times is at the root, giving multiplication of a sum by i. And i+i+i has two as well, one grouping to the left and one to the right. Both kinds of ambiguity are removed by the level structure of X3.

munotes.in218

Ambiguity of a Grammar

Three things to notice about that pair. The two grammars generate the same language, which the checker decides, so removing ambiguity is a change of grammar and not of language. X4 is shorter and easier to write, which is why it is the one a beginner produces. And a real compiler uses something like X3 for exactly this reason.

Inherent ambiguity

Some context free languages have no unambiguous grammar at all. Such a language is called inherently ambiguous, and the standard example is worth knowing by name if not by proof.

the strings a to the i b to the j c to the k, where i equals j OR j equals k

Any grammar for it must be ambiguous, because a string where i, j and k are all equal satisfies the condition for two different reasons, and any grammar has to be able to build it for each reason, giving two trees. The full proof is beyond this syllabus; the fact and the example are not.

So the picture has three levels, and MU's questions distinguish them:

An ambiguous grammarAn inherently ambiguous language
What is ambiguousthe grammarevery grammar for the language
Can be repairedsometimes, by forcing a structurenever
Shown bytwo trees for one stringa proof about all grammars
ExampleS->SbSaa to the i b to the j c to the k with i equal to j or j equal to k

Is a grammar ambiguous? There is no general test

This is worth stating plainly because students look for an algorithm.

Deciding whether an arbitrary context free grammar is ambiguous is undecidable. Chapter 77 proves it, by reducing the Post correspondence problem to it. So no method can answer the question for every grammar, and none ever will.

What can be done, and what this book's checker does, is a bounded search: look at every string up to some length and count its derivations. Finding a string with two proves ambiguity outright, because a witness is a proof. Finding none up to that length proves nothing in general, and the checker says so: a claim of unambiguity in this book is a claim about strings up to its bound, and every such claim in this chapter is true for a reason a reader can also see from the grammar's shape.

munotes.in219

Ambiguity of a Grammar

That asymmetry is the shape of the undecidable: the positive answer has a finite witness and the negative answer does not.

Distinctions

Two derivationsTwo leftmost derivationsTwo trees
Proves ambiguitynoyesyes
Common in ordinary grammarsverynono
Related byone to one with trees
AmbiguousUnambiguous
Some string hastwo or more treesexactly one tree
Repairforce a level structure, if possiblenothing to do
Shown byexhibiting two treesa bounded check, or an argument about the grammar's shape

What it does NOT mean

Ambiguity is a property of the grammar, not of the language. X3 and X4 above generate the same language, and one is ambiguous. A language is only called ambiguous when EVERY grammar for it is, which is the inherent case.

Two derivations do not prove ambiguity. Two trees, or two leftmost derivations, do.

A variable appearing twice on a right hand side does not make a grammar ambiguous. It often does, and X2 above shows a grammar that avoids it, and there are grammars with a doubled variable that are unambiguous. The trees must be exhibited.

Removing ambiguity does not change the language. It changes the grammar. The checker's equality claim on X2 against X1 is exactly that point.

Failing to find two trees does not prove a grammar unambiguous. The question is undecidable in general, so a search that finds nothing is evidence and not proof, and this book's unambiguity claims are bounded and say so.

Quick revision

  • A grammar is ambiguous when some string of its language has two or more derivation trees, equivalently two or more

leftmost derivations.

  • To show it: exhibit the string and both trees, or both leftmost derivations. An argument is not enough and two

ordinary derivations are not enough.

  • MU's grammar S->SbS | a is ambiguous, and the shortest witness is ababa, not the shortest string of the language.
  • Repair by forcing a structure: S->abS | a for the same language, or the three level E, T, F grammar for arithmetic.
  • Removing ambiguity changes the grammar and not the language.
  • An inherently ambiguous language has no unambiguous grammar; the standard example is a to the i b to the j c to the

k with i equal to j or j equal to k.

  • Deciding ambiguity for an arbitrary grammar is undecidable, so a witness proves ambiguity and a failed search proves

nothing.

Test yourself

1. Define an ambiguous grammar, and say what is NOT sufficient to show one. A grammar is ambiguous when some string of its language has two or more distinct derivation trees. Two ordinary derivations are not sufficient, because one tree yields many of them; two trees or two leftmost derivations are.

munotes.in220

Ambiguity of a Grammar

2. Show that S->SbS | a is ambiguous. Take ababa. One tree has the first b at the root, with a on the left and a tree for aba on the right. The other has the second b at the root, with a tree for aba on the left and a on the right. Both yield ababa, so the grammar is ambiguous.

3. Give an unambiguous grammar for the same language. S->abS | a. Every string is built from the left in one way, so each has exactly one tree, and the language is unchanged.

4. Which string shows the one line grammar E->E+E | EE | (E) | i ambiguous? i+ii. One tree has the plus at the root, reading it as i added to the product of i and i; the other has the times at the root, reading it as the sum of i and i multiplied by i. The string i+i+i also has two trees, grouping left or right.

5. What is an inherently ambiguous language, and give the standard example. A context free language for which every grammar is ambiguous. The standard example is the set of a to the i b to the j c to the k with i equal to j or j equal to k, where a string with all three equal qualifies for two reasons.

6. Is there an algorithm to decide whether a grammar is ambiguous? No. The question is undecidable, by a reduction from the Post correspondence problem, so a witness of ambiguity is a proof and a failed search is only evidence.

Contents This chapter on its own page

munotes.in221

Chapter Forty-Five

Simplifying a Grammar, One: Useless Symbols

Syllabus topic Module 1, "Context Free Languages: CFG simplification"

In one line

A symbol is useless if it derives no terminal string, or if the start symbol cannot reach it, and useless symbols and their productions can be deleted.

In the wording a student can write in an examination: a variable A of a context free grammar is generating if A derives some string of terminals, and reachable if S derives some sentential form containing A. A symbol is useful if it is both generating and reachable, and useless otherwise. Deleting the useless symbols and every production mentioning one leaves a grammar generating the same language.

Why simplify at all

Two reasons, and the second is the one that makes these three chapters compulsory rather than tidy.

Because a grammar produced by a construction is full of rubbish. Chapter 59 converts a pushdown automaton into a grammar and the result has variables for state triples that can never be used. Chapter 48's normal form conversion creates fresh variables, some of which turn out unreachable.

Because the later constructions require it. Chomsky normal form in chapter 48 and Greibach normal form in chapter 49 both assume the grammar has no useless symbols, no null productions and no unit productions. If it has, the conversions either fail or produce something that is not in the normal form. So simplification is a prerequisite and not an optional polish.

The two kinds of uselessness

They are genuinely different and the two tests have nothing in common.

Not generating. The variable derives no terminal string at all. Every derivation from it gets stuck with a variable that cannot be removed. Such a variable can never appear in a completed derivation, so it contributes nothing.

Not reachable. No sentential form derivable from S contains the variable. It may be perfectly capable of deriving terminal strings; nothing can ever ask it to.

A variable may fail either test, or both. It is useful only if it passes both.

The order matters, and this is the step people get wrong

Remove the non generating symbols first, then the unreachable ones.

The reason is that removing a non generating variable can make another variable unreachable, since the production that reached it may have been deleted. Doing it the other way round leaves such variables behind.

Worked, in miniature. Take

S->aA | b
A->aA

Accepts: b

Rejects: ε, a, ab, aa, ba

A is not generating: every production of A contains A, so no derivation from A ever produces a terminal string. Remove A and every production mentioning it, which removes S to aA. What is left is S to b, and nothing is unreachable.

Now do it in the wrong order. First look for unreachable symbols: S reaches A through S to aA, so A is reachable, and nothing is removed. Then look for non generating symbols: A is one, so A goes and with it S to aA. The answer happens to be the same here because nothing depended on A, but in a grammar where a third variable was reached only through a production mentioning A, that third variable would survive the reachability pass and then be stranded.

munotes.in222

Simplifying a Grammar, One: Useless Symbols

The safe rule is one line: generating, then reachable, and never the other way.

Finding the generating variables

An upward closure. Start from what is obviously generating and grow.

  1. Mark every variable having a production whose right hand side is all terminals, or is empty.
  2. Repeatedly, mark any variable having a production whose right hand side consists only of terminals and variables

already marked.

  1. Stop when nothing new is marked. The unmarked variables are not generating.

It terminates because there are finitely many variables and the marked set only grows.

Finding the reachable symbols

A downward closure, and it is the reachability search of chapter 39 in another costume.

  1. Mark S.
  2. Repeatedly, for every marked variable, mark every symbol appearing on the right hand side of any of its

productions.

  1. Stop when nothing new is marked. The unmarked symbols are unreachable.

Note that this pass marks terminals as well as variables, and an unreachable terminal can be removed from T. That is usually of no consequence and is worth mentioning because an examiner may ask for the full quadruple after simplification.

Worked example

S->AB | a
A->aA | a
B->bB
C->cC | c
D->BC

Accepts: a

Rejects: ε, ab, b, c, aa, abc

Step one: which variables are generating?

Pass 1. S has the production S to a, all terminals, so S is marked. A has A to a, so A is marked. C has C to c, so C is marked. B's only production is B to bB, which contains B, unmarked, so B is not marked yet. D's only production is D to BC, which contains B, unmarked, so D is not marked yet.

Marked so far: S, A, C.

Pass 2. B still has only B to bB, and B is not marked, so B is still unmarked. D has D to BC, and B is not marked, so D stays unmarked.

Pass 3 marks nothing new, so the process stops. B and D are not generating.

Delete them and every production mentioning either. That removes B to bB, D to BC, and S to AB, because that production mentions B. What is left:

S->a
A->aA | a
C->cC | c

Accepts: a

Rejects: ε, aa, c, cc, ac

munotes.in223

Simplifying a Grammar, One: Useless Symbols

Step two: which symbols are reachable?

Mark S. S's only production is S to a, so mark a. Nothing else is reached.

A and C are unreachable, and so is the terminal c, and so is b, which has already gone with B.

Delete them:

S->a

Accepts: a

Rejects: ε, aa, aaa

Note that the terminal set has shrunk with the variables. The original grammar's b and c are gone from T as well, because the only productions that used them were the ones deleted, and a full answer gives the whole quadruple: V is {S}, T is {a}, P is the single production, and the start symbol is S.

L(Y4) = L(Y2)

The answer. The grammar reduces to the single production S to a, and its language is the one string language {a}, which the checker confirms is exactly what the original five variable grammar generated. Four variables and eight productions were doing nothing at all.

That is an extreme case chosen to make the two passes visible. A real grammar loses less, and a grammar produced by a construction often loses a great deal.

A second example, where the order shows

S->AB | CD
A->a
B->b
C->cC
D->d

Accepts: ab

Rejects: ε, a, b, cd, abd, cccd

Generating. A, B and D are generating at once, from A to a, B to b, D to d. S is generating because S to AB has A and B both marked. C is not: its only production is C to cC. So C goes, and with it C to cC and S to CD.

S->AB
A->a
B->b
D->d

Accepts: ab

Rejects: ε, a, b, d, ad, abd

Reachable. Mark S, then A and B from S to AB, then a and b. D is not reached, because the only production that mentioned D was S to CD, which the first pass deleted.

S->AB
A->a
B->b

Accepts: ab

Rejects: ε, a, b, ba, aab, abb

The terminals c and d have gone from T with the variables that used them, so the reduced grammar is over {a, b} alone, and its language is the single string ab.

L(Y7) = L(Y5)

This is the example that shows the order. D was reachable in the original grammar, through S to CD. It became unreachable only because the generating pass deleted that production. A student who ran the reachability pass first would have kept D and its production, and the answer would carry a variable that can never be used.

Distinctions

Not generatingNot reachable
Meansderives no terminal stringS derives no form containing it
Found bymarking upward from terminal productionsmarking downward from S
Test directionfrom the leaves towards Sfrom S towards the leaves
Which pass firstthis onethis one second
Effect of getting the order wrongnonea stranded variable survives
munotes.in224

Simplifying a Grammar, One: Useless Symbols

A useless symbolA redundant production
Deleted hereyesonly when it mentions a useless symbol
Changes the languagenono
Chapters 46 and 47 removenull and unit productions

What it does NOT mean

Useless does not mean unused in one derivation. It means it can never appear in any completed derivation from S.

Not generating is not the same as not reachable. A variable can be reachable and useless, as B was in the first example, and it can be generating and useless, as A and C were.

The order is not a matter of taste. Generating first. The second example shows why.

Simplifying does not change the language. Every claim in this chapter is checked for that, and if a simplification changed the language it would be a mistake and not a simplification.

A grammar with no useless symbols is not yet reduced. Chapters 46 and 47 remove null and unit productions, and all three are needed before the normal forms of chapters 48 and 49.

Quick revision

  • A variable is generating when it derives some terminal string, and reachable when S derives some form containing it.

Useful means both; useless means either fails.

  • Generating: mark upward from productions whose right hand sides are all terminals or empty, and repeat.
  • Reachable: mark S, then everything on the right hand sides of marked variables, and repeat.
  • Do generating FIRST, then reachable. Deleting a non generating variable can strand another variable, and the other

order leaves it behind.

  • Delete every production mentioning a deleted symbol.
  • The language is unchanged, and simplification is a prerequisite for chapters 48 and 49, not a polish.

Test yourself

1. Define generating and reachable. A variable is generating if it derives at least one string of terminals, and reachable if S derives at least one sentential form containing it. A symbol is useful only if it is both.

2. Give the algorithm for the generating variables. Mark every variable with a production whose right hand side is all terminals or empty; then repeatedly mark any variable with a production whose right hand side contains only terminals and already marked variables; stop when nothing new is marked.

3. Why must the generating pass run before the reachability pass? Because deleting a non generating variable deletes the productions that mention it, and those may have been the only route to some other variable. That variable then becomes unreachable, and a reachability pass run earlier would have kept it.

munotes.in225

Simplifying a Grammar, One: Useless Symbols

4. Simplify S->AB | a, A->a, B->bB. B is not generating, so B and the productions B to bB and S to AB go. That leaves S to a and A to a. Then A is unreachable, so it goes, leaving S to a.

5. Is a reachable variable always useful? No. It must also be generating. A variable reached from S but deriving no terminal string is useless, and the first worked example has one.

6. Does simplification ever change the language? No, and if it does the work is wrong. Every simplification in this chapter is checked for exact equality of the languages before and after.

Contents This chapter on its own page

munotes.in226

Chapter Forty-Six

Simplifying a Grammar, Two: Null Productions

Syllabus topic Module 1, "Context Free Languages: CFG simplification"

In one line

A null production is one whose right hand side is empty, and it can be removed by adding, for every production, the variants with the nullable symbols left out.

In the wording a student can write in an examination: a production of the form A to epsilon is called a null or epsilon production. A variable A is nullable if A derives epsilon in zero or more steps. Given a context free grammar G, a grammar without null productions generating L(G) minus the empty string is obtained by replacing every production with all the versions of it in which any subset of its nullable symbols is omitted, discarding those whose right hand side becomes empty, and deleting the null productions.

Why remove them

Because the normal forms of chapters 48 and 49 have no room for them. Chomsky normal form allows a right hand side of exactly two variables or exactly one terminal, and epsilon is neither. So the conversion cannot begin until the null productions are gone.

And because they are what makes the CYK algorithm of chapter 51 and the pushdown construction of chapter 58 work: both rely on every step of a derivation making the string longer or consuming a symbol, and a null production makes a step that does neither.

The one thing that is lost

A grammar with no null productions cannot generate the empty string, since every derivation from S produces at least one terminal.

So if epsilon is in L(G), it cannot be in the language of the simplified grammar. The standard remedy is to add a fresh start symbol:

S0->S | ε

with S0 the new start symbol and S the old one, and with S0 appearing on no right hand side. That single null production is then the only one in the grammar, it is on the start symbol alone, and chapter 26 noted that type 1 grammars allow exactly that exception for exactly this reason.

The convention in this chapter, and in most textbooks, is to remove the null productions first and add the fresh start symbol afterwards if the empty string is wanted. A question asking for the removal expects the first part; a question asking for an equivalent grammar expects both.

Finding the nullable variables

An upward closure, the same shape as chapter 45's generating test.

  1. Mark every variable having the production A to epsilon.
  2. Repeatedly, mark any variable having a production whose right hand side consists only of marked variables.
  3. Stop when nothing new is marked.

Step 2 is the step that is forgotten. A variable with no epsilon production of its own can still be nullable: if B is nullable and C is nullable then A to BC makes A nullable, because both can vanish.

munotes.in227

Simplifying a Grammar, Two: Null Productions

The replacement

For each production, look at which symbols on its right hand side are nullable. Then write out every version of the production obtained by omitting some subset of those occurrences, including the empty subset, which is the production itself.

Two rules govern it.

Discard a version whose right hand side becomes empty. That would be a new null production, which defeats the purpose. The only exception is the fresh start symbol above.

Each occurrence is decided separately. If A to BB and B is nullable, the versions are A to BB, A to B taking the first out, A to B taking the second out, and A to nothing, which is discarded. The two middle versions are the same production, so it is written once: a grammar is a set of productions and duplicates collapse.

A production with k nullable occurrences gives 2 to the k versions before discarding, which is chapter 2's count of subsets and is why the grammar can grow.

Worked example 1

S->ABaC
A->BC
B->b | ε
C->D | ε
D->d

Accepts: a, ba, bba, ad, bad, da, bdbad

Rejects: ε, b, d, ab, dab, badd

Step 1: the nullable variables.

Pass 1. B has B to epsilon, so B is nullable. C has C to epsilon, so C is nullable.

Pass 2. A has A to BC, and both B and C are now marked, so A is nullable. D has only D to d, so not. S has S to ABaC, which contains the terminal a, so S is not nullable however nullable the variables are.

Pass 3 marks nothing new. Nullable: A, B, C.

Step 2: the replacement, production by production.

S to ABaC. Three nullable occurrences: A, B and C. So 2 to the 3, that is 8 versions:

OmitVersion
nothingS to ABaC
AS to BaC
BS to AaC
CS to ABa
A and BS to aC
A and CS to Aa
B and CS to Aa... no: omitting B and C from ABaC leaves S to Aa
A, B and CS to a

Eight versions and one duplicate pair in the listing above, so seven distinct productions for S. None is empty, because the terminal a is always there.

A to BC. Both occurrences nullable, so four versions: A to BC, A to C, A to B, and A to nothing. The last is discarded.

B to b. No nullable occurrence, so it stands. B to epsilon is deleted.

C to D. D is not nullable, so it stands. C to epsilon is deleted.

munotes.in228

Simplifying a Grammar, Two: Null Productions

D to d. Stands.

The result:

S->ABaC | BaC | AaC | ABa | aC | Aa | a
A->BC | B | C
B->b
C->D
D->d

Accepts: a, ba, bba, ad, bad, da, bdbad

Rejects: ε, b, d, ab, dab, badd

L(Z2) = L(Z1)

The original grammar does not generate epsilon, because S to ABaC always leaves the terminal a, so nothing is lost and the two languages are exactly equal, which the checker decides.

Seven productions for S where there was one. That growth is normal and is the price of the normal forms.

Worked example 2, where the empty string is in the language

S->aSb | ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab

Nullable: S, from S to epsilon. Nothing else, there being nothing else.

The replacement. S to aSb has one nullable occurrence, S, so two versions: S to aSb and S to ab. Neither is empty. And S to epsilon is deleted.

S->aSb | ab

Accepts: ab, aabb, aaabbb

Rejects: ε, a, b, ba, aab

L(Z4) = L(Z4R)

S->aSb | ab

The language is now a to the n followed by b to the n for n at least 1, and the empty string is gone. That is correct and expected.

To get it back, add the fresh start symbol:

S0->S | ε
S->aSb | ab

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab

L(Z5) = L(Z3)

Now the language is exactly the original one, and the grammar has exactly one null production, on a start symbol appearing on no right hand side. The checker decides that equality, so the remedy is verified rather than asserted.

The three errors

Missing a nullable variable. A variable with no epsilon production of its own can be nullable through its productions. Step 2 of the closure is not optional.

Keeping a version whose right hand side is empty. That reintroduces a null production and the grammar is not simplified.

Forgetting the empty string. If epsilon was in the language it is not any more, and a question asking for an equivalent grammar wants the fresh start symbol.

Distinctions

A null productionA nullable variable
Isa production A to epsilona variable deriving epsilon in any number of steps
Found byreading the grammaran upward closure
Every one ison a nullable variablenot necessarily the subject of a null production
Removing null productionsAdding the fresh start symbol
Effect on the languageremoves the empty string, if it was thereputs it back
Number of null productions afterzeroexactly one, on the start symbol
Needed for chapters 48 and 49yesthe single exception is tolerated
munotes.in229

Simplifying a Grammar, Two: Null Productions

What it does NOT mean

Nullable does not mean having a null production. A derives epsilon through A to BC when B and C are both nullable, without A having a null production of its own.

The replacement does not delete the original production. It keeps it and adds the shortened versions, unless the original was itself null.

Removing null productions does not preserve the language exactly when epsilon is in it. It removes epsilon, and the fresh start symbol is how it is restored.

Two identical versions are one production. A grammar is a set of productions, so duplicates from the replacement collapse.

The grammar gets bigger, and that is not a fault. A production with k nullable occurrences gives up to 2 to the k versions.

Quick revision

  • A null production is A to epsilon. A variable is nullable when it derives epsilon in any number of steps.
  • Find the nullable variables by an upward closure: mark those with a null production, then any whose right hand side

is entirely marked variables, and repeat.

  • For each production, add every version omitting some subset of its nullable occurrences; discard a version whose

right hand side is empty; then delete the null productions.

  • A production with k nullable occurrences gives up to 2 to the k versions, and duplicates collapse.
  • The empty string is lost. Restore it with a fresh start symbol S0 to S or epsilon, with S0 on no right hand side.
  • This is a prerequisite for chapters 48 and 49, not a polish.

Test yourself

1. Define nullable, and give a variable that is nullable without having a null production. A variable is nullable if it derives epsilon in zero or more steps. In S->AB, A->ε, B->ε, the variable S is nullable through S to AB, and has no null production of its own.

2. Remove the null productions from S->ABC, A->a | ε, B->b | ε, C->c. Nullable: A and B. S to ABC has two nullable occurrences, giving S to ABC, S to BC, S to AC and S to C. Then A to a, B to b and C to c stand, and the two null productions are deleted.

3. How many versions does a production with 3 nullable occurrences give, and how many survive? Eight, being the subsets of three occurrences. All survive unless one has an empty right hand side, which happens only when every symbol of the production is nullable.

4. What happens to the empty string, and how is it restored? It leaves the language, because no derivation can now produce nothing. Add a fresh start symbol S0 with the productions S0 to S and S0 to epsilon, and make sure S0 appears on no right hand side.

munotes.in230

Simplifying a Grammar, Two: Null Productions

5. Why must a version with an empty right hand side be discarded? Because it is itself a null production, which is what the step exists to remove. The only tolerated null production is the one on the fresh start symbol.

6. Why is this step needed before Chomsky normal form? Because that form permits a right hand side of exactly two variables or exactly one terminal, and epsilon is neither, so a grammar with null productions cannot be put into it.

Contents This chapter on its own page

munotes.in231

Chapter Forty-Seven

Simplifying a Grammar, Three: Unit Productions, and the Reduced Grammar

Syllabus topic Module 1, "Context Free Languages: CFG simplification"

In one line

A unit production replaces one variable by one variable and says nothing, and it can be removed by giving each variable the productions of every variable it can reach through unit productions.

In the wording a student can write in an examination: a production of the form A to B, with A and B both variables, is called a unit production. The pair (A, B) is a unit pair if A derives B using unit productions only. To remove unit productions, compute all unit pairs, and for each pair (A, B) add to A every non unit production of B; then delete all unit productions.

Why remove them

Because they carry no information and get in the way of everything else.

A unit production adds a level to a derivation tree and a step to a derivation while producing nothing. Chomsky normal form in chapter 48 forbids them, because its right hand sides are two variables or one terminal and a single variable is neither. Greibach normal form in chapter 49 forbids them too. And the CYK algorithm of chapter 51 needs a grammar in Chomsky normal form, so this step is upstream of it as well.

They also create cycles. If A to B and B to A are both productions, a derivation can go round for ever without producing anything, which is the pathology chapter 44's checker had to guard against.

Unit pairs

The definition is a closure, and it must be a closure rather than a single step, because unit productions chain.

(A, A) is a unit pair for every variable A, by zero steps. That is the reflexive part and it is what students forget, and it is the part that keeps A's own non unit productions.

If (A, B) is a unit pair and B to C is a unit production, then (A, C) is a unit pair.

The algorithm

  1. Put (A, A) in the set for every variable A.
  2. Repeatedly: if (A, B) is in the set and B to C is a production with C a single variable, add (A, C).
  3. Stop when nothing new is added.

The replacement

For each unit pair (A, B), add to A every production of B whose right hand side is not a single variable. Then delete every unit production.

The pair (A, A) contributes A's own non unit productions, which is why they survive.

Worked example

S->AB
A->a
B->C | b
C->D
D->E | d
E->a

Accepts: aa, ab, ad

Rejects: ε, a, b, ba, abd

Step 1: the unit pairs.

Reflexive: (S, S), (A, A), (B, B), (C, C), (D, D), (E, E).

Now the unit productions are B to C, C to D, D to E. Chaining:

munotes.in232

Simplifying a Grammar, Three: Unit Productions, and the Reduced Grammar

PairFrom
(B, C)B to C
(C, D)C to D
(D, E)D to E
(B, D)(B, C) and C to D
(C, E)(C, D) and D to E
(B, E)(B, D) and D to E

Nothing more, so twelve pairs in all, six reflexive and six chained. Note (B, E): B reaches E through three unit productions, and a single pass would have missed it.

Step 2: the replacement. For each pair, take the non unit productions of the second variable.

PairNon unit productions of the secondAdded to
(S, S)S to ABS
(A, A)A to aA
(B, B)B to bB
(C, C)noneC
(D, D)D to dD
(E, E)E to aE
(B, C)noneB
(B, D)D to dB
(B, E)E to aB
(C, D)D to dC
(C, E)E to aC
(D, E)E to aD

Deleting the unit productions B to C, C to D and D to E, the result is:

S->AB
A->a
B->b | d | a
C->d | a
D->d | a
E->a

Accepts: aa, ab, ad

Rejects: ε, a, b, ba, abd

L(AA2) = L(AA1)

The language is unchanged, which the checker decides. And notice that C, D and E are now unreachable: nothing mentions them any more, because the productions that did were the unit productions just deleted.

That is the observation which fixes the order of the three steps.

The order of the three steps

Null productions, then unit productions, then useless symbols.

Two reasons, one for each adjacency.

Null before unit. Removing null productions can create unit productions. If A to BC and C is nullable, the replacement adds A to B, which is a unit production that was not there before. So doing unit first would leave the new ones behind.

Unit before useless. Removing unit productions can make variables useless, as C, D and E just became. So the useless pass must come last to sweep them up.

A grammar with the three steps done, in that order, is called reduced. Some books use reduced for the useless symbol step alone, so an answer is safer saying which steps were done.

The whole reduction, end to end

This is the examination question, so here it is on one grammar with all three steps in order.

S->AB | a
A->aA | B | ε
B->b | C
C->cC | ε

Accepts: ε, a, b, c, ab, ac, aab, cc, abc

munotes.in233

Simplifying a Grammar, Three: Unit Productions, and the Reduced Grammar

Rejects: ba, ca, bca

Step one: null productions.

Nullable: A, from A to epsilon. C, from C to epsilon. B, because B to C and C is nullable. And S, because S to AB with both A and B nullable.

So all four variables are nullable, and the empty string is in the language.

The replacements:

  • S to AB has two nullable occurrences, giving S to AB, S to A, S to B, and S to nothing, which is discarded. And S to

a stands.

  • A to aA gives A to aA and A to a. A to B gives A to B and A to nothing, discarded. A to epsilon is deleted.
  • B to b stands. B to C gives B to C and B to nothing, discarded.
  • C to cC gives C to cC and C to c. C to epsilon is deleted.
S->AB | A | B | a
A->aA | a | B
B->b | C
C->cC | c

Accepts: a, b, c, ab, ac, aab, cc, abc

Rejects: ε, ba, ca

The empty string has gone, as chapter 46 said it would.

Step two: unit productions.

The unit productions are S to A, S to B, A to B, B to C. Unit pairs, beyond the six reflexive ones:

PairFrom
(S, A)S to A
(S, B)S to B
(A, B)A to B
(B, C)B to C
(S, C)(S, B) and B to C, or (S, A) then (A, B) then B to C
(A, C)(A, B) and B to C

The replacement, taking non unit productions:

  • S keeps S to AB and S to a; gains A's non unit ones, aA and a; gains B's, which is b; gains C's, cC and c.
  • A keeps aA and a; gains B's b and C's cC and c.
  • B keeps b; gains C's cC and c.
  • C keeps cC and c.
S->AB | a | aA | b | cC | c
A->aA | a | b | cC | c
B->b | cC | c
C->cC | c

Accepts: a, b, c, ab, ac, aab, cc, abc

Rejects: ε, ba, ca

L(AA5) = L(AA4)

Step three: useless symbols.

Generating: all four, each having a production of terminals alone. Reachable: S, then A and B from S to AB, then C from cC. So nothing is useless here, and step three removes nothing.

That is a perfectly good outcome and worth seeing: a grammar can need one step and not another. The point of running all three in order is that you do not know in advance which will bite.

munotes.in234

Simplifying a Grammar, Three: Unit Productions, and the Reduced Grammar

The reduced grammar is AA5, and if the empty string is wanted back, chapter 46's remedy applies:

S0->S | ε
S->AB | a | aA | b | cC | c
A->aA | a | b | cC | c
B->b | cC | c
C->cC | c

Accepts: ε, a, b, c, ab, ac, aab, cc, abc

Rejects: ba, ca, bca

L(AA6) = L(AA3)

That final equality is the whole answer: the reduced grammar, with the empty string restored, generates exactly what the original did. The checker decides it over every string up to its bound, so the three steps have been verified and not merely performed.

Distinctions

A unit productionA unit pair
Isa production A to B with B a single variablea relation, A derives B by unit productions only
Reflexivenoyes, (A, A) always
Found byreading the grammara closure
Used fornothing; it is what is removeddeciding what to copy where
StepRemovesCan createSo it comes
null productionsA to epsilonunit productionsfirst
unit productionsA to Buseless symbolssecond
useless symbolsunreachable and non generating symbolsnothingthird

What it does NOT mean

A unit production is not a production with one symbol. A to a is a production with one symbol and is not a unit production, because a is a terminal. Only variable to variable counts.

The reflexive pair is not optional. Without (A, A) the variable loses its own productions and the grammar generates nothing.

A single pass is not a closure. Unit productions chain, and the worked example has a pair three links long.

The order is not a matter of taste. Null, unit, useless. Doing unit first leaves the unit productions that null removal creates; doing useless first leaves the symbols that unit removal orphans.

Reduced does not mean minimal. A reduced grammar can still be larger than necessary, and nothing in this book minimises a grammar the way chapter 20 minimises a machine.

Quick revision

  • A unit production is A to B with both variables. A unit pair is (A, B) with A deriving B by unit productions alone,

and (A, A) is always a pair.

  • Compute the pairs by closure, then for each pair (A, B) add to A every NON unit production of B, then delete the

unit productions.

  • The pair (A, A) is what keeps A's own productions.
  • The three steps run in the order null, unit, useless: null removal creates unit productions, and unit removal

creates useless symbols.

  • A grammar with all three done is reduced. The empty string, if it was in the language, is restored with a fresh
munotes.in235

Simplifying a Grammar, Three: Unit Productions, and the Reduced Grammar

start symbol.

  • Reduced does not mean minimal.

Test yourself

1. Define a unit production and a unit pair, and say why the pairs must be computed as a closure. A unit production is A to B with A and B both variables. A unit pair (A, B) holds when A derives B using unit productions only, including zero of them. The pairs must be closed because unit productions chain: A to B and B to C make (A, C) a pair, and a single pass would miss it.

2. Why is (A, A) always a unit pair, and what does it contribute? Because A derives A in zero steps. It contributes A's own non unit productions to the new grammar, and without it every variable would lose them.

3. Remove the unit productions from S->A | a, A->B | b, B->c. Pairs beyond the reflexive: (S, A), (A, B), (S, B). Then S gains a from itself, b from A and c from B; A gains b from itself and c from B; B keeps c. So S->a | b | c, A->b | c, B->c, and A and B are now unreachable.

4. Give the order of the three steps and the reason for each adjacency. Null, then unit, then useless. Null removal can create a unit production, as A to BC with C nullable gives A to B, so unit must come after. Unit removal can orphan a variable, as the worked example shows, so useless must come last.

5. Is A->a a unit production? No. Its right hand side is a terminal, not a variable. Only a production whose right hand side is a single variable counts.

6. After the three steps, what is the grammar called, and does it generate the empty string? Reduced. It does not generate the empty string, since null productions are gone; if the original language contained it, a fresh start symbol S0 with S0 to S and S0 to epsilon restores it.

Contents This chapter on its own page

munotes.in236

Chapter Forty-Eight

Chomsky Normal Form

Syllabus topic Module 1, "Context Free Languages: Normal Forms"

In one line

A grammar is in Chomsky normal form when every production is either two variables or one terminal, and every context free language has one.

In the wording a student can write in an examination: a context free grammar is in Chomsky normal form if every production has the form A to BC with B and C variables, or A to a with a a terminal, together with the single production S to epsilon when the empty string is in the language and S appears on no right hand side. Every context free language is generated by a grammar in Chomsky normal form.

Why the form is worth having

Because it makes every derivation tree the same shape, and two of the most important results in this block depend on that shape.

Every interior vertex has exactly two children, or one leaf child. So a tree for a string of length n is a binary tree with n leaves, and a binary tree with n leaves has exactly n minus 1 interior vertices with two children. The number of steps in a derivation is therefore fixed by the length of the string, which is what makes CYK in chapter 51 terminate in a predictable number of steps.

A tall tree must have a repeated variable on some path. That is the pigeonhole principle again, and it is the whole of the pumping lemma for context free languages in chapter 50. The proof needs the two children per vertex property to bound the height of the tree by the length of the string, and without the normal form there is no such bound.

So this is not tidying. It is the preparation for the two theorems that follow.

The form, precisely

Every production must be one of:

FormMeaning
A to BCexactly two symbols, both variables
A to aexactly one symbol, a terminal
S to epsilononly for the start symbol, and only if S appears on no right hand side

Three things are therefore forbidden: a right hand side of three or more symbols; a mixture of variables and terminals; and a single variable, which is a unit production.

Note that A to BC permits B and C to be the same variable, so A to AA is in the form.

The four steps

Step 0. Reduce the grammar. Remove null productions, then unit productions, then useless symbols, by chapters 45 to 47. Without this the conversion cannot succeed, because the form forbids exactly those two kinds of production.

Step 1. Add a fresh start symbol if needed. If the old start symbol appears on any right hand side, add S0 to S with S0 fresh. This is needed only for the S to epsilon exception, and it costs one production.

munotes.in237

Chomsky Normal Form

Step 2. Replace the terminals inside long right hand sides. For every terminal a appearing in a production whose right hand side has two or more symbols, introduce a fresh variable, conventionally called Xa or Ta, with the single production Xa to a, and replace every such occurrence of a by Xa.

After this step every right hand side is either a single terminal or a string of variables.

Step 3. Break up the long right hand sides. For every production A to B1 B2 ... Bk with k at least 3, introduce fresh variables and split from the right:

A->B1 Y1

Y1->B2 Y2

...

Y(k minus 2)->B(k minus 1) Bk

A production with k symbols becomes k minus 1 productions and needs k minus 2 fresh variables. Splitting from the left works equally well and gives a different but equally correct answer.

Worked example 1

S->ASA | aB
A->B | S
B->b | ε

Accepts: a, aa, ab, ba, aba, bab, abb, aab

Rejects: ε, b, bb, bbb, bbbb

Step 0, reduce.

Nullable: B, from B to epsilon; A, from A to B; and S? S to ASA has A, S, A, and S is not yet known nullable, so not; S to aB has the terminal a, so not. So nullable is {A, B}.

Removing null productions: S to ASA gives S to ASA, S to SA, S to AS, S to S. S to aB gives S to aB and S to a. A to B gives A to B; A to S gives A to S. B to b stands, B to epsilon goes.

S->ASA | SA | AS | S | aB | a
A->B | S
B->b

Accepts: a, aa, ab, ba, aba, bab, abb, aab

Rejects: ε, b, bb, bbb, bbbb

Now the unit productions: S to S, A to B, A to S. Unit pairs beyond the reflexive: (S, S) is reflexive anyway, (A, B), (A, S). The replacement gives A the non unit productions of B, which is b, and of S, which are ASA, SA, AS, aB, a.

S->ASA | SA | AS | aB | a
A->b | ASA | SA | AS | aB | a
B->b

Accepts: a, aa, ab, ba, aba, bab, abb, aab

Rejects: ε, b, bb, bbb, bbbb

L(AB3) = L(AB2)

Useless symbols: all three variables are generating and all three reachable, so nothing goes.

Step 1. S appears on right hand sides, so a fresh start symbol would be needed if the empty string were in the language. It is not, so this step is skipped and recorded as skipped.

munotes.in238

Chomsky Normal Form

Step 2, terminals in long right hand sides. The terminal a appears in S to aB and A to aB, both of length 2. Add Xa to a and replace:

S->ASA | SA | AS | Xa B | a
A->b | ASA | SA | AS | Xa B | a
B->b
Xa->a

Accepts: a, aa, ab, ba, aba, bab, abb, aab

Rejects: ε, b, bb, bbb, bbbb

Note that S to a and A to a are left alone: they are already in the form, being a single terminal.

Step 3, break up the long ones. Only ASA is longer than two symbols. It appears in S and in A, and one fresh variable serves both:

S->A Y1 and Y1->S A

S->A Y1 | SA | AS | Xa B | a
A->b | A Y1 | SA | AS | Xa B | a
B->b
Xa->a
Y1->SA

Accepts: a, aa, ab, ba, aba, bab, abb, aab

Rejects: ε, b, bb, bbb, bbbb

AB5 = CNF(AB3)

That single claim is the whole verification, and the checker does two things with it: it inspects every production of AB5 against the three permitted forms, and it compares the language of AB5 with the language of AB3 over every string up to its bound. So the answer is in the form and generates the right language, both decided.

Check a few productions by eye as well, against the table of permitted forms.

ProductionShapePermitted?
S to A Y1two variablesyes
S to SAtwo variablesyes
S to Xa Btwo variablesyes
S to aone terminalyes
B to bone terminalyes
Xa to aone terminalyes
Y1 to SAtwo variablesyes

Every production is one of the two permitted shapes, and none is a unit production or longer than two symbols.

Worked example 2, with a long production

S->aXbXc
X->d | e

Accepts: adbdc, aebec, adbec, aebdc

Rejects: ε, abc, adbc, dadbdc

Step 0. No null, unit or useless anything. Nothing to do.

Step 2, terminals. The right hand side aXbXc has three terminals in it. Add Xa to a, Xb to b, Xc to c:

S->Xa X Xb X Xc

Step 3, break up. Five symbols, so four productions and three fresh variables, splitting from the right:

S->Xa Y1
Y1->X Y2
Y2->Xb Y3
Y3->X Xc
X->d | e
Xa->a
Xb->b
Xc->c
munotes.in239

Chomsky Normal Form

Accepts: adbdc, aebec, adbec, aebdc

Rejects: ε, abc, adbc, dadbdc

AB7 = CNF(AB6)

Five symbols became four productions with three fresh variables, exactly as the count said: k minus 1 productions and k minus 2 new variables, with k equal to 5.

The cost

A production of k symbols with t terminals in it becomes k minus 1 productions and adds up to k minus 2 fresh variables, plus one fresh variable per distinct terminal used in a long right hand side. So the grammar grows roughly in proportion to the total length of its right hand sides, which is a linear cost and the reason the form is practical.

The number of productions never becomes more than a constant times the original grammar's total size. That bound is what makes CYK's cost in chapter 51 a polynomial in the length of the string.

Distinctions

In Chomsky normal formNot
A to BCyes
A to ayes
S to epsilon, S on no right hand sideyes, the one exception
A to aBno, mixes a terminal and a variable
A to BCDno, three symbols
A to Bno, a unit production
A to epsilon for A not the start symbolno
Step 2Step 3
Removesterminals from long right hand sidesright hand sides longer than two
Fresh variablesone per terminalk minus 2 per production of length k
Must comebefore step 3after step 2

What it does NOT mean

A to AA is not forbidden. Two variables, and they may be the same one.

A to a is not a unit production. a is a terminal. Unit means variable to variable.

The conversion does not work on an unreduced grammar. Null and unit productions are forbidden by the form, so chapters 46 and 47 come first.

The form is not unique. Splitting from the left instead of the right gives a different grammar in the form, generating the same language. Both are correct.

The empty string needs the exception. A grammar in the strict form with no S to epsilon cannot generate epsilon, so a language containing it needs the one permitted null production on a start symbol that appears on no right hand side.

Chomsky normal form does not remove ambiguity. It changes the shape of the trees, not their number. An ambiguous grammar converts to an ambiguous grammar.

Quick revision

  • The form: every production is A to BC with both variables, or A to a with a a terminal, plus S to epsilon when the

empty string is in the language and S is on no right hand side.

munotes.in240

Chomsky Normal Form

  • Four steps: reduce the grammar; add a fresh start symbol if needed; replace terminals inside long right hand sides

by fresh variables; then split long right hand sides two symbols at a time.

  • Step 2 before step 3, or the split would leave a terminal in a two symbol production.
  • A production of k symbols gives k minus 1 productions and k minus 2 fresh variables.
  • Every interior vertex of a tree then has two children or one leaf child, which fixes the number of derivation steps

by the length of the string. That is what CYK and the CFG pumping lemma both need.

  • The form is not unique and does not remove ambiguity.

Test yourself

1. Give the permitted forms of a production, including the exception. A to BC with B and C variables; A to a with a a terminal; and S to epsilon, only for the start symbol and only when S appears on no right hand side.

2. Convert S->aAB, A->b, B->c to Chomsky normal form. The terminal a sits in a three symbol right hand side, so add Xa to a, giving S to Xa A B. Then split: S to Xa Y1 and Y1 to A B. With A to b and B to c, which are already in the form, the answer is S->Xa Y1, Y1->AB, A->b, B->c, Xa->a.

3. Why must the grammar be reduced first? Because the form forbids unit productions and null productions other than the one on the start symbol, so a grammar still carrying them cannot be in the form however the right hand sides are split.

4. How many productions and fresh variables does a right hand side of 6 variables become? Five productions and four fresh variables, being k minus 1 and k minus 2 with k equal to 6.

5. Is A->AA in Chomsky normal form? Is A->aa? A to AA is, since it is two variables and they may coincide. A to aa is not, since it is two terminals; it becomes A to Xa Xa with Xa to a.

6. What property of derivation trees does the form guarantee, and which two later results need it? Every interior vertex has exactly two children or a single leaf child, so a tree for a string of length n is binary with n leaves and n minus 1 branching vertices. The CYK algorithm of chapter 51 and the pumping lemma for context free languages of chapter 50 both depend on it.

Contents This chapter on its own page

munotes.in241

Chapter Forty-Nine

Greibach Normal Form

Syllabus topic Module 1, "Context Free Languages: Normal Forms"

In one line

A grammar is in Greibach normal form when every production starts with a terminal and has nothing but variables after it, so every step of a derivation produces exactly one terminal.

In the wording a student can write in an examination: a context free grammar is in Greibach normal form if every production has the form A to a alpha, where a is a terminal and alpha is a string of variables, possibly empty, together with the single production S to epsilon when the empty string is in the language and S appears on no right hand side. Every context free language is generated by a grammar in Greibach normal form.

Why this form and not the other

Because of one property, and it is the property chapter 58 needs.

Every step of a derivation produces exactly one terminal. So a derivation of a string of length n takes exactly n steps, no more and no fewer.

Compare with Chomsky normal form, where a derivation of a string of length n takes 2n minus 1 steps, some of them producing no terminal at all. Both are bounded, and only Greibach's bound is tight.

That tightness is what makes the grammar to pushdown automaton construction of chapter 58 terminate. That construction keeps the outstanding variables on a stack and consumes one input symbol per move. With Greibach normal form every move consumes a symbol, so the machine cannot loop; with any other form it could push and pop for ever without reading.

The form, precisely

Every production must be one of:

FormMeaning
A to aone terminal, nothing after it
A to a B1 B2 ... Bkone terminal, then any number of variables
S to epsilononly for the start symbol, and only if S appears on no right hand side

So three things are forbidden: a production starting with a variable; a terminal anywhere but first; and a unit production, which starts with a variable.

Note that the number of variables after the terminal is unrestricted, unlike Chomsky normal form's exactly two. So Greibach normal form is not a special case of Chomsky's, nor the other way round.

Left recursion, and why it must go first

A variable A is left recursive when A derives a string beginning with A itself. The simplest case is a production A to A alpha, which is called immediate left recursion.

Left recursion is fatal to this form, and the reason is immediate: a production A to A alpha starts with a variable, and substituting A's productions into it gives another production starting with whatever A starts with, which is A again. The substitution never terminates.

So left recursion is removed first, and the removal is a construction worth knowing on its own.

munotes.in242

Greibach Normal Form

Removing immediate left recursion

Suppose A has productions

A->A a1 | A a2 | ... | A am | b1 | b2 | ... | bn

where the a's are the tails of the left recursive productions and the b's are the productions not starting with A, of which there must be at least one or A generates nothing.

Introduce a fresh variable A prime and replace them by

A->b1 | b2 | ... | bn | b1 A' | b2 A' | ... | bn A'

A'->a1 | a2 | ... | am | a1 A' | a2 A' | ... | am A'

Read it as a change of direction. The original grammar built a string by adding tails on the right while recursing on the left. The new one begins with one of the non recursive productions and then adds tails on the right, recursing on the right. The set of strings is the same and the recursion now moves the other way.

Worked, on its own

A->Ab | Ac | a | d

Accepts: a, d, ab, ac, db, abc, dcb

Rejects: ε, b, c, ba, ca, bd

The tails are b and c; the non recursive productions are a and d. So:

A->a | d | aA' | dA'
A'->b | c | bA' | cA'

Accepts: a, d, ab, ac, db, abc, dcb

Rejects: ε, b, c, ba, ca, bd

L(AC2) = L(AC1)

AC2 = GNF(AC1)

The language is unchanged, which the checker decides, and the result happens to be in Greibach normal form already, because every right hand side begins with a terminal and holds nothing but variables after it. That is not a coincidence for this grammar and it is not general.

The conversion

Step 0. Reduce the grammar, by chapters 45 to 47, and put it in Chomsky normal form by chapter 48. Starting from Chomsky normal form is not compulsory but it makes every later step mechanical, because then every production is two variables or one terminal.

Step 1. Order the variables A1 to An, with the start symbol first.

Step 2. Remove left recursion, working through the variables in order. For each Ai in turn, first substitute the productions of A1 to A(i minus 1) into any production of Ai beginning with one of them, repeatedly, until no production of Ai begins with an earlier variable; then remove Ai's immediate left recursion by the construction above.

After this step, every production of Ai begins with a terminal or with a later variable.

munotes.in243

Greibach Normal Form

Step 3. Substitute backwards. Working from An down to A1, replace a leading later variable by its productions, which by step 2 all begin with terminals. Then every production of every variable begins with a terminal.

Step 4. Do the same for the fresh primed variables, whose productions may still begin with a variable.

Worked example 1

S->AB
A->aA | a
B->b

Accepts: ab, aab, aaab

Rejects: ε, a, b, ba, abb, aba

What needs fixing. A to aA and A to a begin with terminals and hold only variables after. B to b likewise. Only S to AB is wrong: it begins with the variable A.

There is no left recursion, since no production of S begins with S and none of A begins with A: A to aA begins with the terminal a.

Substitute A into S to AB. A's productions are aA and a, so:

S->aAB | aB
A->aA | a
B->b

Accepts: ab, aab, aaab

Rejects: ε, a, b, ba, abb, aba

L(AC4) = L(AC3)

AC4 = GNF(AC3)

Every production now begins with a terminal and holds only variables after it, and the language is unchanged. Both facts are decided by the checker.

Worked example 2, with left recursion

E->E+T | T
T->i

Accepts: i, i+i, i+i+i

Rejects: ε, +, i+, +i, ii

This is the addition half of the arithmetic grammar, and E is immediately left recursive.

Step 2, remove the left recursion of E. The tail is +T and the non recursive production is T. So with a fresh E prime:

E->T | T E'

E'->+T | +T E'

E prime's productions now begin with the terminal plus, which is right. E's still begin with the variable T.

Step 3, substitute T into E. T's only production is T to i, so:

E->iE' | i
E'->+TE' | +T
T->i

Accepts: i, i+i, i+i+i

Rejects: ε, +, i+, +i, ii

L(AC6) = L(AC5)

AC6 = GNF(AC5)

Every production begins with a terminal. E to iE prime and E to i; E prime to +TE prime and E prime to +T; T to i. The language is unchanged, and the form is verified.

Worked example 3, where a terminal has to be given a variable

The case students meet and do not expect.

S->(S)S | ε

Accepts: ε, (), (()), ()(), (()())

Rejects: (, ), )(, ((), ())

The production S to (S)S begins with a terminal, which is right, and then has S, a terminal close bracket, and S. That middle terminal is forbidden by the form.

munotes.in244

Greibach Normal Form

The remedy is chapter 48's step 2 in another costume: give the offending terminal a variable of its own.

R->)

and replace the occurrence, then deal with the empty string by a fresh start symbol as chapter 46 prescribes. The nullable S also has to be removed, which gives the four versions of the production:

S0->S | ε
S->(SRS | (RS | (SR | (R
R->)

Accepts: ε, (), (()), ()(), (()())

Rejects: (, ), )(, ((), ())

L(AC8) = L(AC7)

Every production of S begins with the terminal open bracket and holds only variables after it, and R to close bracket is one terminal. The only production not in the form is the S0 to S left by the fresh start symbol, which the standard statement of the form allows for the start symbol alongside S0 to epsilon, and which can be removed by substituting S's productions into it if a question insists.

The two forms compared

Chomsky normal formGreibach normal form
ProductionsA to BC, or A to aA to a followed by variables
Terminals per production0 or 1exactly 1
Position of the terminalalonefirst
Variables per production2 or 0any number, including 0
Steps in a derivation of length n2n minus 1exactly n
Left recursionpermittedmust be removed
Needed forCYK, chapter 51; the CFG pumping lemma, chapter 50the pushdown construction, chapter 58
Conversion difficultymechanicalneeds left recursion removal and two substitution passes

Distinctions

Immediate left recursionIndirect left recursion
Looks likeA to A alphaA to B alpha and B derives A something
Removed bythe A prime construction directlysubstituting the earlier variables first, then the construction
Found byreading one productionfollowing derivations

What it does NOT mean

Greibach normal form is not a special case of Chomsky's. The number of variables after the terminal is unrestricted, so A to aBCD is in Greibach form and not in Chomsky form.

A production may have no variables after the terminal. A to a is in the form.

Removing left recursion does not change the language. It changes the direction the recursion runs, and the worked example checks that the two languages are equal.

A terminal in the middle is not allowed. It must be given a variable of its own, as in the bracket example.

The conversion does not need Chomsky normal form first. Starting there makes it mechanical, and the examples above start from grammars that are not in it.

Quick revision

  • The form: every production is a terminal followed by any number of variables, plus S to epsilon when the empty string
munotes.in245

Greibach Normal Form

is in the language and S is on no right hand side.

  • The property: every derivation step produces exactly one terminal, so a string of length n takes exactly n steps.

That is what makes chapter 58's pushdown construction terminate.

  • Left recursion must go first. Immediate left recursion A to A alpha with non recursive productions b is removed by A

to b and A to b A prime, with A prime to alpha and A prime to alpha A prime.

  • Then order the variables, substitute earlier ones out of the front of later ones, remove each immediate left

recursion, and substitute backwards so that every production begins with a terminal.

  • A terminal anywhere but first gets a variable of its own.
  • Not a special case of Chomsky normal form: any number of variables may follow the terminal.

Test yourself

1. Give the permitted form of a production, with the exception. A to a alpha where a is a terminal and alpha is a string of variables, possibly empty; plus S to epsilon for the start symbol when the empty string is in the language and S is on no right hand side.

2. Which property does this form have that Chomsky normal form does not, and what needs it? Every derivation step produces exactly one terminal, so a derivation of a string of length n has exactly n steps. The grammar to pushdown automaton construction of chapter 58 needs it, so that every move of the machine consumes an input symbol and the machine cannot loop.

3. Remove the left recursion from A->Aa | b. A to b and A to b A prime, with A prime to a and A prime to a A prime. The recursion now runs to the right.

4. Why must left recursion be removed before the substitutions? Because a production A to A alpha begins with A, and substituting A's productions into it produces another production beginning with whatever A begins with, which is A again. The substitution never terminates.

5. Convert S->AB, A->a, B->b to Greibach normal form. S to AB begins with a variable, so substitute A to a, giving S to aB. With A to a and B to b already in the form, the answer is S->aB, A->a, B->b, and A is now unreachable and may be dropped.

6. Is A->aBCD in Greibach normal form? Is it in Chomsky normal form? It is in Greibach normal form, since the terminal comes first and only variables follow. It is not in Chomsky normal form, which permits exactly two variables or exactly one terminal.

Contents This chapter on its own page

munotes.in246

Chapter Fifty

The Pumping Lemma for Context Free Languages

Syllabus topic Module 1, "Context Free Languages: Pumping Lemma for CFG"

In one line

Every long enough string of a context free language can be cut into five pieces, and the second and fourth can be repeated together any number of times without leaving the language.

In the wording a student can write in an examination: let L be a context free language. Then there exists a constant n such that every string z in L with length at least n can be written as z equal to uvwxy, where the length of vwx is at most n, the length of vx is at least 1, and u v to the i w x to the i y is in L for every i at least 0.

Why five pieces rather than three

Because a context free machine has a stack, and a stack matches one thing against another.

The regular pumping lemma of chapter 36 came from a state repeating on a path, which gave one loop and hence one piece to repeat. This one comes from a variable repeating on a path in a derivation tree, and a variable's subtree has material on both sides of the repeat, so the repeat produces two pieces, one on each side, and they grow together.

That is also why the lemma is powerful enough to defeat a language like a to the n b to the n c to the n: the two pieces can only be in two of the three blocks, so the third gets left behind.

The statement, clause by clause

There exists a constant n, depending on L, which you do not choose.

Every z in L with length at least n. Shorter strings are exempt.

z equals uvwxy. Five pieces, in order. The first and last may be empty.

The length of vwx is at most n. So the three middle pieces together lie in a window of at most n symbols. This is the condition that limits where v and x can be, and it is the one that does the work.

The length of vx is at least 1. So v and x are not both empty. One of them may be, and a proof must handle that.

u v to the i w x to the i y is in L for every i at least 0. The two pieces are repeated the same number of times, together, and i equal to 0 deletes both.

The proof

Take a grammar G for L in Chomsky normal form, by chapter 48, with k variables.

Why Chomsky normal form. Because in that form every interior vertex of a derivation tree has exactly two children, so a tree of height h has at most 2 to the (h minus 1) leaves. Turn that round: a tree with more than 2 to the (h minus

munotes.in247

The Pumping Lemma for Context Free Languages

  1. leaves has height more than h. That is the bound the proof needs and no other form gives it.

Choose n equal to 2 to the k. Then any string z in L of length at least n has a derivation tree with at least 2 to the k leaves, so its height is greater than k, so it has a path with more than k interior vertices on it.

Pigeonhole. Those interior vertices are labelled with variables, and there are only k variables. So some variable repeats on that path. Take the lowest two occurrences of a repeated variable, and call it A.

Read off the five pieces. The lower occurrence of A has a subtree yielding some string w. The upper occurrence has a subtree yielding some string vwx, where v is the part to the left of the lower subtree and x the part to the right. And the rest of the tree, outside the upper A's subtree, yields u on the left and y on the right. So z is uvwxy.

The three conditions.

The length of vwx is at most n, because the path from the upper A downward has at most k plus 1 interior vertices, the two A among them, so the upper A's subtree has height at most k plus 1 and therefore at most 2 to the k leaves.

The length of vx is at least 1, because the upper occurrence of A has two children, and the lower A lies under one of them, so the other contributes at least one leaf, and every leaf in Chomsky normal form carries a terminal. This is where the normal form is needed a second time: without it the upper A could have a single child and v and x could both be empty.

The pumping. A derives vAx, because that is what the piece of the tree between the two occurrences does, and A derives w, because that is the lower subtree. So A derives v to the i, then A, then x to the i, for any i, by repeating that piece; and then A derives w. Putting u and y back:

S derives u A y, and A derives v to the i then w then x to the i

so u v to the i w x to the i y is in L for every i at least 0, the case i equal to 0 coming from A deriving w directly.

That is the lemma.

Using it: the shape of an argument

The four move game of chapter 37 applies unchanged, with one extra complication at move 3.

munotes.in248

The Pumping Lemma for Context Free Languages

MoveWhoseWhat happens
1the adversarychooses n
2youchoose z in L, of length at least n, in terms of n
3the adversarychooses uvwxy with vwx at most n long and vx at least 1 long
4youchoose i, and must land outside L

Move 3 is harder than it was. In chapter 37 the condition on uv confined v to the first n symbols of z, and a well chosen z made every splitting the same. Here vwx may sit anywhere in z, as long as it is a window of at most n symbols. So a proof must consider every position the window can take, and that is what the case analysis is for.

The standard way to control it. Choose z with three or more blocks, each of length n or more. Then the window vwx, being at most n long, can touch at most two of the blocks. That gives a small number of cases, and in each one at least one block is untouched by the pumping and falls out of step with the others.

Language 1: a to the n b to the n c to the n

Claim. The language of a to the m, b to the m, c to the m, for m at least 0, is not context free.

Proof. Suppose it is, with constant n.

Choose z equal to a to the n, then b to the n, then c to the n. It is in the language and its length is 3n.

Any splitting. z is uvwxy with the length of vwx at most n and vx non empty. The three blocks are each n long, so a window of at most n symbols cannot reach from the a block into the c block: they are n symbols apart. So vwx touches at most two of the three blocks, and at least one block is untouched.

Choose i equal to 2. Pumping increases the count of whichever symbols appear in v and x, and vx is non empty so at least one count increases. The untouched block's count does not change. So the three counts are no longer all equal, and the pumped string is not in the language.

Written out by cases, which is what a full answer does:

Where vwx sitsWhat pumping changesWhy it fails
inside the a blockthe a count onlya count exceeds b and c
inside the b blockthe b count onlyb count exceeds a and c
inside the c blockthe c count onlyc count exceeds a and b
across a and bthe a and b countsc count is left behind
across b and cthe b and c countsa count is left behind
munotes.in249

The Pumping Lemma for Context Free Languages

Five cases, and the sixth, across a and c, cannot arise because the window is too short. Every case gives a string outside the language.

Contradiction, so the language is not context free.

That completes chapter 26's hierarchy. This language is context sensitive, and chapter 23's grammar for it is type 1, so the containment of the context free languages in the context sensitive ones is proper. The witness is this language and the proof is this one.

Language 2: w followed by w

Claim. The language of strings of the form ww, where w is any string over {a, b}, is not context free.

Proof. Suppose it is, with constant n.

Choose z equal to a to the n, then b to the n, then a to the n, then b to the n. That is w followed by w with w equal to a to the n followed by b to the n, so z is in the language, and its length is 4n.

Any splitting. vwx is at most n long, so it lies within two adjacent blocks.

Choose i equal to 0, which deletes v and x. The result is shorter than z by between 1 and n symbols, and it must still be of the form ww for the language to hold. Its first half and second half must agree.

Deleting from within the first two blocks shortens the first half and leaves the second half untouched, or shortens it unevenly, and in every case the two halves no longer agree symbol for symbol: the deletion removes a from the front half without removing the matching a from the back half, or removes b likewise. The same argument applies to a deletion from the last two blocks. And a deletion straddling the middle removes b from the end of the first half and a from the start of the second, which destroys the pattern at the join.

Contradiction, so the language is not context free.

Why this language matters. It is the language of a repeated block, and it is the standard example of something a compiler cannot check with a context free grammar: "this identifier has been declared before use" is a condition of exactly this shape, which is why type checking is not part of parsing.

Language 3: a to the m squared

Claim. The language of a to the m squared, the strings of a whose length is a perfect square, is not context free.

Proof. Suppose it is, with constant n.

Choose z equal to a to the n squared, whose length is n squared, at least n.

munotes.in250

The Pumping Lemma for Context Free Languages

Any splitting. Every symbol is a, so vx is a to the p for some p with 1 at most p at most n.

Choose i equal to 2. The pumped string has length n squared plus p, and

n squared < n squared plus p and n squared plus p at most n squared plus n < n squared plus 2n plus 1

so the length lies strictly between n squared and the next perfect square, and is therefore not a perfect square.

Contradiction. Note that this is chapter 37's proof with vx in place of v, and the bound on p comes from the length of vwx being at most n. The same language is therefore neither regular nor context free, by nearly the same argument.

The two lemmas compared

Regular, chapter 36Context free, here
Piecesthree, uvwfive, uvwxy
Repeatedvv and x, together, the same number of times
Window conditionuv at most nvwx at most n
Non empty conditionv at least 1vx at least 1
Where the window can bethe first n symbolsanywhere
Comes froma state repeating on a runa variable repeating on a path in a tree
Needsthe pigeonhole principlethe pigeonhole principle and Chomsky normal form
Number of cases in a proofusually oneusually several, one per window position

What it does NOT mean

It does not say every splitting pumps. Some splitting does, so a proof must beat every legal one.

v and x are not repeated independently. The same i for both. A proof that pumps only v has used a different and false lemma.

The window can be anywhere. Unlike chapter 36, there is no condition placing it at the front, which is why the case analysis is needed.

vx at least 1 does not mean both are non empty. Either may be empty, and a proof must allow it.

Satisfying the lemma does not make a language context free. The implication runs one way, exactly as in chapter 36, and there is no context free analogue of Myhill Nerode that settles the question both ways.

Quick revision

  • Statement: if L is context free there is an n such that every z in L of length at least n splits as uvwxy with vwx at

most n long, vx at least 1 long, and u v to the i w x to the i y in L for every i at least 0.

  • Proof: take a grammar in Chomsky normal form with k variables and set n to 2 to the k. A string that long has a tree
munotes.in251

The Pumping Lemma for Context Free Languages

of height more than k, so a variable repeats on a path; the piece between the two occurrences gives v and x, and the lower subtree gives w.

  • Chomsky normal form is needed twice: for the height bound, and to guarantee vx is non empty.
  • Using it: choose z with three or more blocks each at least n long, so the window of at most n symbols touches at most

two blocks and at least one block is left behind.

  • a to the m b to the m c to the m is not context free, which makes chapter 26's third containment proper.
  • ww is not context free, which is why a compiler cannot check declaration before use while parsing.
  • Perfect squares are neither regular nor context free, by the same arithmetic both times.

Test yourself

1. State the lemma with all three conditions. If L is context free there is a constant n such that every z in L of length at least n can be written uvwxy with the length of vwx at most n, the length of vx at least 1, and u v to the i w x to the i y in L for every i at least 0.

2. Why five pieces, and where do v and x come from? Because the repeat is of a VARIABLE on a path in a derivation tree, and the piece of tree between the two occurrences has material on both sides of the lower subtree. v is what lies to its left and x what lies to its right.

3. Where is Chomsky normal form used in the proof, and for what? Twice. For the height bound, since two children per vertex makes a tree with more than 2 to the (h minus 1) leaves taller than h; and to guarantee that vx is non empty, since the upper repeated variable has two children and the one not containing the lower occurrence must contribute a leaf.

4. Prove that a to the m b to the m c to the m is not context free. Suppose it is with constant n and take z equal to a to the n b to the n c to the n. Since vwx is at most n long it touches at most two of the three blocks. Pumping to i equal to 2 raises the counts of the symbols in vx, which is non empty, and leaves at least one block's count unchanged, so the three counts differ and the string leaves the language.

5. Why is a case analysis needed here and not in chapter 37? Because the window vwx may sit anywhere in the string, whereas chapter 36's condition confined v to the first n symbols. Choosing z with three long blocks keeps the number of cases small.

munotes.in252

The Pumping Lemma for Context Free Languages

6. Can this lemma prove a language context free? No. Like chapter 36's it is a one way implication, and unlike the regular case there is no context free analogue of the Myhill Nerode theorem that would settle the question in both directions.

Contents This chapter on its own page

munotes.in253

Chapter Fifty-One

Closure Properties and Decision Problems of the Context Free Languages

Syllabus topic Module 1, "Context Free Languages: Context-free Languages"

In one line

The context free languages are closed under union, concatenation and closure, and are NOT closed under intersection or complement, and membership and emptiness are decidable while equivalence is not.

In the wording a student can write in an examination: the context free languages are closed under union, concatenation, Kleene closure, reversal, and intersection with a regular language, and are not closed under intersection or complement. The membership, emptiness and finiteness problems are decidable for context free languages, and the equivalence, containment, ambiguity and emptiness of intersection problems are undecidable.

The closures that hold

Each is proved by a construction on grammars, and each is three lines.

Union. Let G1 and G2 generate L1 and L2, with disjoint variable sets, which can always be arranged by renaming. Add a fresh start symbol S with

S->S1 | S2

where S1 and S2 are the two old start symbols. A derivation commits to one grammar at the first step and stays there, so the language is the union.

Concatenation. With the same disjoint variable sets, add a fresh start symbol with

S->S1 S2

Every derivation produces a string from L1 followed by one from L2.

Kleene closure. Add a fresh start symbol with

S->S1 S | ε

so S derives any number of copies of S1, including none.

Reversal. Reverse every right hand side of the grammar. A derivation of w becomes a derivation of w reversed.

Intersection with a REGULAR language. This one is not a grammar construction and is the most useful of the five. Take a pushdown automaton for the context free language and a finite automaton for the regular one, and build the product: states are pairs, the stack is the pushdown automaton's, and both components move on each input symbol. The result is a pushdown automaton, so the intersection is context free. Chapter 52 defines the machine this needs.

That last closure is what makes chapter 38's proof technique work in this class too, and the worked example below uses it.

Worked: union and concatenation

S1->aS1b | ε

Accepts: ε, ab, aabb

Rejects: a, b, ba, aab

S2->cS2 | c

Accepts: c, cc, ccc

Rejects: ε

The rejection list holds one string because this grammar's only terminal is c and it accepts every non empty string of c, so the empty string is the only thing over its own alphabet left to reject.

Their union:

S->S1 | S2
S1->aS1b | ε
S2->cS2 | c

Accepts: ε, ab, aabb, c, cc, ccc

Rejects: a, b, ba, abc, cab

Their concatenation:

S->S1 S2
S1->aS1b | ε
S2->cS2 | c

Accepts: c, cc, abc, abcc, aabbc

munotes.in254

Closure Properties and Decision Problems of the Context Free Languages

Rejects: ε, ab, a, ca, cab, aabb

Both constructions are run, so the claim lists demonstrate the closures rather than illustrating them. Notice that the concatenation no longer contains the empty string, because S2 always produces at least one c.

The closures that FAIL

This is the section that distinguishes this class from the regular one, and MU's questions turn on it.

Intersection: NOT closed

The counterexample, and it is the standard one.

S->XY
X->aXb | ε
Y->cY | ε

Accepts: ε, ab, c, abc, aabbc, abcc

Rejects: a, b, ba, ac, cab, aabc

This generates a to the i, b to the i, c to the j, with i and j independent. It is context free, as the grammar shows.

S->AZ
A->aA | ε
Z->bZc | ε

Accepts: ε, a, bc, abc, aabc, abbcc

Rejects: b, c, cb, acb, abac

This generates a to the i, b to the j, c to the j, again with i and j independent, and is context free.

Their intersection is the set of strings that are both a to the i b to the i c to the j and a to the i b to the j c to the j. The first forces the a count to equal the b count; the second forces the b count to equal the c count. So the intersection is

a to the m then b to the m then c to the m

which chapter 50 proved is not context free.

So two context free languages have an intersection that is not context free. The class is not closed under intersection.

Complement: NOT closed

The proof is a corollary and is worth learning in this form, because it is one line.

Suppose the class were closed under complement. It is closed under union, as shown above. Then by De Morgan's law, which chapter 2 gave,

L1 intersect L2 = (L1' union L2')'

so the class would be closed under intersection. It is not. So it is not closed under complement.

That argument is the same one chapter 38 used to get intersection from complement and union, run backwards. It is worth noticing that the same identity proves closure in the regular case and non closure here.

And a direct witness exists too. The complement of the language ww of chapter 50 is context free, while ww itself is not, which shows the failure the other way round: a context free language whose complement is not.

munotes.in255

Closure Properties and Decision Problems of the Context Free Languages

Difference: NOT closed

Immediate: L1 minus L2 is L1 intersected with the complement of L2, and the class fails both. A weaker statement does hold and is useful: a context free language minus a regular one is context free, because the complement of a regular language is regular and intersection with a regular language is closed.

The ten operations, for both classes

OperationRegularContext free
unionyesyes
concatenationyesyes
Kleene closureyesyes
reversalyesyes
intersection with a regular languageyesyes
intersectionyesno
complementyesno
differenceyesno
prefix, suffix, substringyesyes

Three rows differ, and they are the three that involve saying "not" or "both". A student who remembers only that much has the useful part.

The decision problems

The same table for the five questions of chapter 39, and here three answers change.

ProblemRegularContext free
membershipyesyes, by CYK
emptinessyesyes
finitenessyesyes
equivalenceyesno
containmentyesno

And three more questions that did not arise for regular languages:

ProblemContext free
is the grammar ambiguous?no
is the intersection of two of them empty?no
is the language all of Sigma star?no

The undecidable ones are all proved in chapter 77, by reduction from the Post correspondence problem. What matters here is knowing which side of the line each question falls on.

Emptiness, which is decidable

The algorithm is chapter 45's generating test. L(G) is empty exactly when the start symbol is not generating, that is, when S derives no terminal string. That test is a closure and terminates, so emptiness is decidable.

Finiteness, which is decidable

Reduce the grammar, put it in Chomsky normal form, and look for a variable A that derives a string properly containing A. If there is one, the pumping argument of chapter 50 gives infinitely many strings; if not, every derivation tree has bounded height and the language is finite.

Membership, by CYK

The algorithm, named after Cocke, Younger and Kasami, and it is the reason chapter 48 exists.

It needs the grammar in Chomsky normal form. Then it fills a triangular table: for every substring of the input, it records which variables derive exactly that substring.

  1. For each single symbol of the input, record every variable A with the production A to that terminal.
  2. For each longer substring, in increasing order of length, and for each way of splitting it into two non empty halves,

record every variable A having a production A to BC where B is recorded for the left half and C for the right.

  1. The string is in the language exactly when the start symbol is recorded for the whole input.

It terminates because the table has one cell per substring, and there are about half n squared of them for an input of length n, each taking at most n splits, so the whole thing is a polynomial in n. That polynomial cost is what makes membership practical and is why the normal form was worth converting to.

munotes.in256

Closure Properties and Decision Problems of the Context Free Languages

CYK worked

Take the grammar

S->AB | BC
A->BA | a
B->CC | b
C->AB | a

Accepts: baaba, bab, ab, ba, aaa, bbab

Rejects: ε, a, b, aa, bb, aba, abab

and the input baaba, of length 5. The table has one row per substring length, and the cell for a substring holds every variable that derives exactly that substring.

Row 1, the single symbols. A to a and C to a, so a gives A and C. B to b, so b gives B.

SubstringPositionsVariables
b1B
a2A, C
a3A, C
b4B
a5A, C

Row 2, the pairs. Each pair splits one way only. Look for a production whose right hand side is a variable from the left cell followed by one from the right cell. The productions with two variables are S to AB, S to BC, A to BA, B to CC and C to AB.

SubstringPositionsVariables
ba1 to 2A, S
aa2 to 3B
ab3 to 4C, S
ba4 to 5A, S

Take the first cell. The left half b gives B and the right half a gives A and C, so the candidate pairs are BA and BC. A to BA matches, and S to BC matches. So the cell is A and S.

Take the second. Both halves give A and C, so the candidate pairs are AA, AC, CA and CC. Only CC is a right hand side, of B to CC. So the cell is B.

Row 3, the triples. Each has two splittings, one symbol with two and two with one, and both must be tried.

SubstringPositionsVariables
baa1 to 3none
aab2 to 4B
aba3 to 5B

The first cell is empty, and it is worth seeing why, because an empty cell is a normal and useful outcome. Splitting baa as b with aa gives B on the left and B on the right, and BB is nobody's right hand side. Splitting it as ba with a gives A and S on the left and A and C on the right, so the candidate pairs are AA, AC, SA and SC, and none of the five two variable productions has any of them. So no variable derives baa.

munotes.in257

Closure Properties and Decision Problems of the Context Free Languages

Row 4, the four symbol substrings. Three splittings each.

SubstringPositionsVariables
baab1 to 4none
aaba2 to 5A, C, S

baab is empty for the same kind of reason, and one of its three splittings is dead already: splitting it as baa with b uses the empty cell from row 3, so it contributes nothing.

aaba is the interesting one. Splitting as a with aba gives A and C on the left and B on the right, so AB matches S to AB and C to AB, giving S and C. Splitting as aa with ba gives B on the left and A and S on the right, so BA matches A to BA, giving A. Splitting as aab with a gives B on the left and A and C on the right, so BA gives A and BC gives S. The cell is therefore A, C and S.

Row 5, the whole string. Four splittings.

SubstringPositionsVariables
baaba1 to 5A, C, S

Splitting as b with aaba gives B on the left and A, C, S on the right: BA matches, giving A, and BC matches, giving S. Splitting as ba with aba gives A, S on the left and B on the right: AB matches, giving S and C. The other two splittings use the empty cells for baa and baab and contribute nothing.

The cell contains S, so baaba is in the language, and the checker's claim list above reports the same answer from its own run of the algorithm.

The practical point. CYK is the general parser for context free grammars and it always works, in time proportional to the cube of the input length. Real compilers use faster methods on restricted grammar classes, which is what the LL and LR families of chapter 43 are, and fall back on nothing: they insist on a grammar in the restricted class.

Distinctions

RegularContext free
Closed under intersectionyesno, and the counterexample is the standard pair
Closed under complementyesno, and it follows from intersection by De Morgan
Closed under intersection with a regular languageyesyes, and this is the useful one
Equivalence decidableyesno
Membership decidableyes, by running the machineyes, by CYK, in cubic time
ClosedDecidable
Is aboutbuilding a new languageanswering a question
Failure meansthe operation escapes the classno algorithm exists
The two are related bythe emptiness of an intersection being undecidable, which needs the non closure

What it does NOT mean

Not closed under intersection does not mean intersections are never context free. Often they are. It means some intersection is not, and one counterexample settles the class.

munotes.in258

Closure Properties and Decision Problems of the Context Free Languages

Intersection with a regular language is closed. The failure is only when both languages are context free. This is the single most useful thing in the chapter.

Undecidable equivalence does not mean you can never tell two grammars apart. It means no algorithm works for every pair. Particular pairs are often easy, and the checker behind this book decides bounded comparisons for exactly that reason, and says the bound.

CYK does not need the grammar unambiguous. It records variables, not trees, so ambiguity costs it nothing.

Chapter 38's table does not carry over. Three of its rows are false here, and assuming otherwise is the commonest error in Module 1's last topic.

Quick revision

  • Closed: union, concatenation, Kleene closure, reversal, intersection with a REGULAR language, prefix, suffix,

substring.

  • Not closed: intersection, complement, difference.
  • The intersection counterexample: a to the i b to the i c to the j intersected with a to the i b to the j c to the j is

a to the m b to the m c to the m, which chapter 50 proved not context free.

  • Complement follows from intersection by De Morgan, in one line.
  • Decidable: membership by CYK in cubic time, emptiness by the generating test, finiteness by looking for a variable

deriving a string properly containing itself.

  • Undecidable: equivalence, containment, ambiguity, emptiness of an intersection, and whether the language is all of

Sigma star. All proved in chapter 77.

  • CYK needs Chomsky normal form, fills a triangular table of substrings, and accepts when the start symbol reaches the

whole string.

Test yourself

1. Give the union construction for two context free grammars. Rename so the variable sets are disjoint, then add a fresh start symbol S with the productions S to S1 and S to S2, where S1 and S2 are the two old start symbols.

2. Prove the class is not closed under intersection. Take a to the i b to the i c to the j and a to the i b to the j c to the j, both context free with easy grammars. Their intersection forces the a, b and c counts all equal, giving a to the m b to the m c to the m, which chapter 50 proved is not context free.

3. Deduce that the class is not closed under complement. If it were, then since it is closed under union, De Morgan's law would make it closed under intersection, and it is not.

4. Which intersection IS closed, and why is that useful? The intersection of a context free language with a regular one. It is useful because it lets a language be narrowed before the pumping lemma is applied, exactly as chapter 38 narrowed a regular language with a regular one.

munotes.in259

Closure Properties and Decision Problems of the Context Free Languages

5. What does CYK need, what does it fill in, and what does it cost? A grammar in Chomsky normal form. It fills a triangular table recording, for each substring of the input, which variables derive exactly that substring. The cost is proportional to the cube of the input length.

6. Name three undecidable questions about context free grammars. Whether two grammars generate the same language; whether one grammar is ambiguous; and whether the intersection of two context free languages is empty. All three are proved in chapter 77 by reduction from the Post correspondence problem.

Contents This chapter on its own page

munotes.in260

Module II

Pushdown Automata, Linear Bound Automata, Turing Machines, and Computability and Complexity

munotes.in

Chapter Fifty-Two

The Pushdown Automaton

Syllabus topic Module 2, "Pushdown Automata: Definitions"

In one line

A pushdown automaton is a finite automaton with a stack, and the stack is exactly the memory a finite automaton was missing.

In the wording a student can write in an examination: a pushdown automaton is a seven tuple P = (Q, Sigma, Gamma, delta, q0, Z0, F) where Q is a finite set of states, Sigma the input alphabet, Gamma the stack alphabet, q0 the initial state, Z0 the initial stack symbol, F the set of final states, and delta a map from Q times (Sigma union {epsilon}) times Gamma to finite subsets of Q times Gamma star.

What was missing, exactly

Chapter 37 proved that no finite automaton accepts the strings a to the n followed by b to the n. The proof was the pigeonhole principle: a machine with m states, reading m or more a, must repeat a state, and then it cannot tell how many a it has seen.

So the fault is not a shortage of cleverness. It is that the machine's memory is a fixed number of states and the job needs a count that grows.

A stack is the smallest thing that fixes it. Push one symbol per a, pop one per b, and the machine has counted without needing a state per count. And a stack is not merely enough for this job: chapters 58 and 59 prove it is exactly the right amount of memory for the whole context free class.

Why a stack and not a counter

Because a stack does more than count, and chapter 50 shows it does less than a tape.

More than a counter. A counter could match a to the n against b to the n. It could not recognise a palindrome, which needs the first half remembered in order and read back in reverse. A stack gives that for free: push the first half, then pop and compare, and the last in first out order is exactly the reversal.

Less than a tape. Only the top of the stack is visible, and reading it destroys it. So the machine can match ONE nesting, and chapter 50 proves it cannot match two, which is why a to the n b to the n c to the n is out of reach.

That is the whole shape of this class: nesting yes, two independent counts no.

The seven parts

P = (Q, Sigma, Gamma, delta, q0, Z0, F)

PartWhat it isNew here?
Qa finite set of statesno
Sigmathe input alphabetno
Gammathe STACK alphabetyes
deltathe transition functionchanged
q0the initial stateno
Z0the initial stack symbolyes
Fthe set of final statesno
munotes.in261

The Pushdown Automaton

Gamma is a separate alphabet. The symbols pushed on the stack need not be input symbols, and usually are not. It is normal to push a capital A to record that an a was read, and Gamma is where such symbols live.

Z0 is on the stack before anything happens. The stack is never empty at the start, and Z0 is what marks the bottom. Its job is to let the machine know when the stack is back to where it began, which is how the last example of this chapter detects a balanced string.

delta takes three things and gives a set.

delta(q, a, X) = a set of pairs (p, gamma)

Read it: in state q, reading input symbol a, with X on top of the stack, the machine may move to state p and replace X by the string gamma. The three parts of the input to delta and the two of the output are worth naming one at a time.

The input symbol may be epsilon. Then the machine moves without reading, exactly as chapter 15's empty moves let a finite automaton do. A pushdown automaton uses them constantly, and the last move of most machines in this book is an epsilon move.

The stack symbol is always read, and it is always the top one, and it is always removed. There is no move that leaves the stack untouched without naming its top.

gamma is what replaces the top symbol, written top first.

gamma isEffectCalled
epsilonX is removed and nothing replaces ita pop
XX is removed and put backno change
YXX is removed and Y then X are put backa push of Y
YZXtwo symbols pushed above Xa double push

So there is no separate push and pop instruction: there is one instruction that replaces the top symbol with a string, and push and pop are the cases where that string is longer or shorter.

The answer is a SET, so a pushdown automaton is nondeterministic by definition, exactly as chapter 14's NFA was. Chapter 57 shows that unlike chapter 16's case, the nondeterminism cannot be removed.

How this book writes one

A transition table for a pushdown automaton would need three dimensions, and on a phone it would be unreadable. So this book prints the transition list, which is what textbooks print and what an examination answer writes.

start q0
stack Z
final q2
q0, a, Z -> q0, A Z
q0, a, A -> q0, A A
q0, b, A -> q1, ε
q1, b, A -> q1, ε
q1, ε, Z -> q2, Z
q0, ε, Z -> q2, Z
munotes.in262

The Pushdown Automaton

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

Read the six moves as a description of a plan.

Line 1. In q0, reading a, with the bottom marker on top: push an A above it. This is the first a.

Line 2. In q0, reading a, with an A on top: push another A. Each a adds one A.

Line 3. In q0, reading b, with an A on top: pop it, and go to q1. This is the first b, and moving to q1 records that the b block has begun so that no further a will be accepted.

Line 4. In q1, reading b, with an A on top: pop it. Each b removes one A.

Line 5. In q1, reading nothing, with the bottom marker on top: go to q2, the final state. The marker being on top means every A has been popped, so the counts matched.

Line 6. The same from q0, which handles the empty input: no a were read, so the marker is still on top and the string is accepted.

Why it rejects what it rejects. On aab the machine pops one A and reaches q1 with an A still on the stack and no input left; line 5 needs the marker on top, so there is no move and the run dies. On abab the machine is in q1 after the first b, and q1 has no move on a, so the run dies. On ba there is no move for b with the marker on top.

The checker runs all ten of those strings against the six moves, so the account just given is a reading of something that has been executed.

And the grammar it matches

S->aSb | ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

L(BA1) = L(BA2)

Two productions and six moves for the same language, and the equality is decided by the checker over every string up to its bound. Chapters 58 and 59 turn that coincidence into a theorem in both directions.

A second machine: the stack as a reverser

The job a counter could not do.

start p0
stack Z
final p2
p0, a, Z -> p0, A Z
p0, b, Z -> p0, B Z
p0, a, A -> p0, A A
p0, a, B -> p0, A B
p0, b, A -> p0, B A
p0, b, B -> p0, B B
p0, c, Z -> p1, Z
p0, c, A -> p1, A
p0, c, B -> p1, B
p1, a, A -> p1, ε
p1, b, B -> p1, ε
p1, ε, Z -> p2, Z
munotes.in263

The Pushdown Automaton

Accepts: c, aca, bcb, abcba, aabcbaa

Rejects: ε, a, ac, ca, abc, acb, abcab

This accepts w, then c, then w reversed: the odd length palindromes with a marker in the middle. The first six moves push a record of each symbol read. The middle three cross to p1 on the marker c, whatever is on top, leaving the stack alone. The next two pop one symbol per input symbol, and only when they match: p1 on a needs an A on top, so an a in the second half must correspond to an a in the first. The last move accepts when the marker is back on top.

Note what the stack has done. It was filled with the first half in order, and popped in reverse order, so it compared the second half against the reverse of the first without any state remembering anything. That is the whole reason a stack is the right memory for this class.

The marker c is doing real work, and chapter 56 is about the machine for palindromes without one, where the machine must guess where the middle is and nondeterminism becomes essential.

Distinctions

Finite automatonPushdown automaton
Memorythe current statethe state and a stack of unbounded height
Partsfiveseven, adding Gamma and Z0
delta readsa state and a symbola state, a symbol or epsilon, and the top of the stack
delta givesa state, or a set of thema set of (state, replacement string) pairs
Can match one pair of countsnoyes
Can match twonono
Nondeterminism removableyes, chapter 16no, chapter 57
gammaMeans
εpop
X, the same symbolleave the stack unchanged
YXpush Y
YZXpush Z then Y

What it does NOT mean

There is no separate push and pop. One instruction replaces the top symbol with a string, and push and pop are the cases where that string grows or shrinks it.

The stack top is always consumed. Every move names it and removes it; leaving it in place means writing it back.

The stack is not the input. Gamma is a separate alphabet, and pushing a copy of an input symbol is a choice the designer makes, not a rule.

An empty stack is not an error in itself, but in most machines there is then no move, because every move needs a top symbol. Z0 exists so that the stack is never empty by accident.

A pushdown automaton is nondeterministic by definition. The deterministic kind is a restriction, defined in chapter 57, and it is strictly weaker.

munotes.in264

The Pushdown Automaton

Unbounded stack does not mean unbounded power. Chapter 50 already proved that a to the n b to the n c to the n is out of reach, and this machine is what that proof was about.

Quick revision

  • A pushdown automaton is (Q, Sigma, Gamma, delta, q0, Z0, F): a finite automaton plus a stack alphabet and a bottom

marker.

  • delta(q, a, X) is a set of pairs (p, gamma): in state q, reading a or nothing, with X on top, move to p and replace

X by gamma, written top first.

  • gamma equal to epsilon pops, gamma equal to X leaves the stack alone, gamma equal to YX pushes Y.
  • The input symbol may be epsilon; the stack symbol never is.
  • Z0 marks the bottom, so the machine can tell when the stack is back where it started.
  • The stack counts, and because it is last in first out it also reverses, which a counter cannot do.
  • Nondeterministic by definition, and chapter 57 shows the nondeterminism cannot be removed.

Test yourself

1. Name the seven parts, and say which two are new. Q, Sigma, Gamma, delta, q0, Z0 and F. Gamma, the stack alphabet, and Z0, the initial stack symbol, are the two that a finite automaton does not have.

2. Give the three arguments of delta and the shape of its value. A state, an input symbol or epsilon, and the symbol on top of the stack. Its value is a set of pairs, each a next state and a string to put in place of the top symbol.

3. What does a move with gamma equal to epsilon do? With gamma equal to AX where X was on top? The first pops the top symbol and puts nothing back. The second removes X and puts A then X back, which is a push of A.

4. Why can a stack recognise palindromes when a counter cannot? Because it is last in first out, so the first half pushed comes back off in reverse order, which is exactly what has to be compared with the second half. A counter records only how many, not which or in what order.

5. What is Z0 for? It marks the bottom of the stack, so the machine can recognise that everything it pushed has been popped. Without it there would be no way to distinguish an empty stack from one holding pushed symbols, and every move needs a top symbol to read.

6. Is the definition deterministic? No. delta returns a SET, and epsilon moves are allowed, so the machine may have several possible runs on one input. Chapter 57 defines the deterministic restriction and shows it accepts strictly fewer languages.

Contents This chapter on its own page

munotes.in265

Chapter Fifty-Three

Instantaneous Descriptions and Moves

Syllabus topic Module 2, "Pushdown Automata: Definitions"

In one line

An instantaneous description is a snapshot of everything a pushdown automaton knows at one moment, and a move takes one snapshot to the next.

In the wording a student can write in an examination: an instantaneous description of a pushdown automaton is a triple (q, w, alpha) where q is the current state, w the portion of the input not yet read, and alpha the contents of the stack with its top symbol written first. If delta(q, a, X) contains (p, gamma) then (q, aw, X beta) moves to (p, w, gamma beta), written with the turnstile. A sequence of such moves is a run, and the reflexive transitive closure of the move relation is written with a starred turnstile.

Why a three part snapshot

Because a pushdown automaton knows exactly three things, and a description that omits one of them cannot say what happens next.

The state. Which of the finitely many conditions it is in.

The unread input. Everything already read is gone and cannot influence anything; everything not yet read is what remains to be decided.

The stack. All of it, though only the top is readable.

Nothing else exists. There is no position counter beyond the unread input, no history and no scratch memory. So the triple is complete: two runs that reach the same triple are indistinguishable from then on, which is the same observation chapter 11 made about a finite automaton's state and is what makes the subject tractable.

Writing one

(q, w, alpha)

The stack is written top first. So (q, ab, AAZ) means the stack has two A above the bottom marker Z, and the top A is the one a move will read. Writing it the other way round is a common slip and it inverts every push.

The unread input is written whole, not as a position.

The empty stack is written as the empty string. A machine whose stack has been emptied has alpha equal to epsilon, and then no move naming a top symbol can apply.

This book writes a run as a table of configurations, one per line, in the form state | input left | stack, which is the same triple with the brackets dropped so it fits a phone.

The move relation

One move, written with a turnstile, which this book writes as |- because the turnstile glyph is above the range the typography allows:

if delta(q, a, X) contains (p, gamma) then (q, a w, X beta) |- (p, w, gamma beta)

Read the two sides against each other. On the left the state is q, the input begins with a, and the stack begins with X. On the right the state is p, the a has been consumed, and X has been replaced by gamma. Everything after a in the input and everything under X in the stack is copied across untouched, which is what w and beta are doing.

munotes.in266

Instantaneous Descriptions and Moves

An epsilon move is the same with a equal to epsilon, so the input is not shortened:

if delta(q, ε, X) contains (p, gamma) then (q, w, X beta) |- (p, w, gamma beta)

Zero or more moves is written with a starred turnstile and is the reflexive transitive closure of chapter 4, so every description reaches itself in zero moves.

The two facts that make runs easy to reason about

Both are worth stating because proofs in chapters 58 and 59 use them.

The unread input can be extended. If (q, w, alpha) reaches (p, w prime, beta), then (q, w z, alpha) reaches (p, w prime z, beta) for any z. The machine cannot see past the input it is reading, so appending more input changes nothing about the moves already made.

The stack can be extended underneath. If (q, w, alpha) reaches (p, w prime, beta), then (q, w, alpha gamma) reaches (p, w prime, beta gamma) for any gamma, provided the run never empties the stack down into gamma. The machine cannot see below the top, so symbols further down are inert until reached.

Those two are the pushdown analogue of chapter 11's identity, and they are the reason a sub computation can be reasoned about in isolation.

Worked: a run, printed and re-executed

The machine of chapter 52, for a to the n followed by b to the n.

start q0
stack Z
final q2
q0, a, Z -> q0, A Z
q0, a, A -> q0, A A
q0, b, A -> q1, ε
q1, b, A -> q1, ε
q1, ε, Z -> q2, Z
q0, ε, Z -> q2, Z

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

The run on aabb, as a sequence of configurations:

q0 | aabb | Z
q0 | abb | A Z
q0 | bb | A A Z
q1 | b | A Z
q1 | ε | Z
q2 | ε | Z

Six configurations, so five moves. Read them against the transition list:

MoveRule usedWhat changed
1q0, a, Z to q0, A Zread a, pushed A
2q0, a, A to q0, A Aread a, pushed A
3q0, b, A to q1, εread b, popped A, went to q1
4q1, b, A to q1, εread b, popped A
5q1, ε, Z to q2, Zread nothing, went to the final state
munotes.in267

Instantaneous Descriptions and Moves

The last configuration has no input left and is in q2, which is final, so aabb is accepted.

The checker re-executes every step of that trace against the six rules above, checks that each consumes one input symbol or none, and checks that the last configuration accepts. A trace with a wrong stack, a wrong state or a skipped step fails the build.

The run on aaabbb, the same machine, one symbol longer in each block:

q0 | aaabbb | Z
q0 | aabbb | A Z
q0 | abbb | A A Z
q0 | bbb | A A A Z
q1 | bb | A A Z
q1 | b | A Z
q1 | ε | Z
q2 | ε | Z

Eight configurations, seven moves, and the stack rises to three A and comes back down. Setting a run out like this is what MU's question asks for.

A run that fails, and what failing looks like

There is no separate reject move. A run fails by having no move available, and the description at that point is the answer to what went wrong.

Take aab on the same machine. The run reaches

(q1, ε, A Z)

after reading a, a, b: two A were pushed and one popped. There is no input left, so only an epsilon move could apply, and the only epsilon move from q1 needs Z on top. The top is A. So no move exists, the run is over, and since q1 is not final the string is not accepted by this run. There is no other run, so aab is rejected.

Take abab. After a, b the machine is in q1 with the marker on top. The next symbol is a, and q1 has no move on a at all. The run dies, and abab is rejected.

Reading a failure correctly is a skill an examination tests, and the form of the answer is always the same: give the description reached, and say which of the three parts makes every rule inapplicable.

A longer run: the marked palindromes

The second machine of chapter 52, which uses the stack as a reverser.

start p0
stack Z
final p2
p0, a, Z -> p0, A Z
p0, b, Z -> p0, B Z
p0, a, A -> p0, A A
p0, a, B -> p0, A B
p0, b, A -> p0, B A
p0, b, B -> p0, B B
p0, c, Z -> p1, Z
p0, c, A -> p1, A
p0, c, B -> p1, B
p1, a, A -> p1, ε
p1, b, B -> p1, ε
p1, ε, Z -> p2, Z
munotes.in268

Instantaneous Descriptions and Moves

Accepts: c, aca, bcb, abcba, aabcbaa

Rejects: ε, a, ac, ca, abc, acb, abcab

The run on abcba:

p0 | abcba | Z
p0 | bcba | A Z
p0 | cba | B A Z
p1 | ba | B A Z
p1 | a | A Z
p1 | ε | Z
p2 | ε | Z

Seven configurations, six moves. Read the third: the input symbol is c, the top of the stack is B, and the rule p0, c, B to p1, B crosses to the second half without changing the stack, which is what leaving gamma equal to the symbol read means.

Then the two pops each match: reading b with B on top, and reading a with A on top. Had the second half been ab rather than ba, the fourth move would have been reading a with B on top, and there is no such rule, so the run would have died. That is the machine checking the reversal.

Distinctions

A configurationA move
Isa snapshot, three partsa step from one snapshot to the next
Written(q, w, alpha)with a turnstile between two of them
The stack is writtentop first
Zero of themreaches itselfthe starred turnstile
A run that rejectsA run that fails
Meansended somewhere not acceptinghad no move available
In a nondeterministic machineother runs may still succeedthe same
How to report itgive the final descriptiongive the description and say why no rule applies

What it does NOT mean

The stack is not written bottom first. Top first, always, or every push reads backwards.

A configuration does not record what has been read. Only what has not. The read part cannot influence anything and is therefore not part of the machine's knowledge.

A failed run does not mean the string is rejected, in general. The machine is nondeterministic, so rejection means every run fails, and acceptance means one succeeds. Chapter 52's two machines happen to be deterministic in effect, which is why a single run settles them.

An epsilon move still takes a move. It consumes no input and it is a step, and a trace that skips one does not match the machine.

A move always consumes the top of the stack. Writing it back is what "no change" means, and there is no rule that ignores the stack.

Quick revision

  • A configuration, or instantaneous description, is (q, w, alpha): the state, the unread input, and the stack written

top first.

  • A move: if delta(q, a, X) contains (p, gamma) then (q, aw, X beta) reaches (p, w, gamma beta). Everything after a
munotes.in269

Instantaneous Descriptions and Moves

and everything under X is copied unchanged.

  • An epsilon move is the same with no input consumed. Zero or more moves is the starred turnstile and is reflexive.
  • Extending the unread input, or the stack underneath, does not change the moves already possible. Both facts are

used in chapters 58 and 59.

  • A run is written one configuration per line. There is no reject move: a run fails by having no rule apply, and the

answer to what went wrong is the description reached.

  • The last configuration accepts when the input is exhausted and the state is final.

Test yourself

1. Give the three parts of an instantaneous description and say which end of the stack is written first. The current state, the unread portion of the input, and the whole stack with its TOP symbol written first.

2. Write the move rule in full. If delta(q, a, X) contains (p, gamma), then (q, a w, X beta) moves to (p, w, gamma beta) for every w and beta: the a is consumed, X is replaced by gamma, and the rest of the input and the rest of the stack are unchanged.

3. For chapter 52's machine, give the run on ab. (q0, ab, Z), then (q0, b, AZ) by pushing on the a, then (q1, epsilon, Z) by popping on the b, then (q2, epsilon, Z) by the epsilon move. The last is final with no input left, so ab is accepted.

4. The machine reaches (q1, epsilon, AZ). What happens, and why? Nothing. There is no input left so only an epsilon move could apply, and the only one from q1 requires Z on top while A is on top. The run has no continuation, and since q1 is not final this run does not accept.

5. Why does extending the stack underneath not change a run? Because only the top symbol is ever read, so symbols below are inert until they are reached. The proviso is that the run must not pop down into the added part, and with that the same sequence of moves is available.

6. What does gamma equal to the symbol just read from the stack mean? That the stack is unchanged: the top symbol is removed and immediately put back. It is how a machine reads the top without disturbing it, and chapter 52's crossing moves use it.

Contents This chapter on its own page

munotes.in270

Chapter Fifty-Four

Acceptance by a PDA: by Final State and by Empty Stack

Syllabus topic Module 2, "Pushdown Automata: Acceptance by PDA"

In one line

A pushdown automaton may be said to accept when it ends in a final state, or when it empties its stack, and the two definitions accept the same class of languages but not the same language from the same machine.

In the wording a student can write in an examination: the language accepted by final state, written L(P), is the set of strings w such that (q0, w, Z0) reaches (p, epsilon, alpha) for some p in F and some alpha. The language accepted by empty stack, written N(P), is the set of w such that (q0, w, Z0) reaches (q, epsilon, epsilon) for some q. For every P there is a P prime with N(P prime) equal to L(P), and for every P there is a P double prime with L(P double prime) equal to N(P), so the two definitions characterise the same class of languages.

Why there are two definitions at all

Because the two are natural in different places, and both are in the literature.

Final state is the definition inherited from chapter 9. A finite automaton accepts by ending in a final state, so a machine that is a finite automaton with an extra component ought to as well. It is the definition MU's papers use most and the one this book uses by default.

Empty stack is the definition the construction of chapter 58 produces. That construction turns a grammar into a machine whose stack holds the variables still to be expanded, and the natural moment to accept is when there are none left, which is when the stack is empty. Forcing that construction to use final states would mean adding machinery that does nothing.

So the two definitions exist because two different jobs each have an obvious answer, and the theorem below says the choice does not matter for the class of languages.

The two definitions, precisely

Write the starting description as (q0, w, Z0) in both cases.

Acceptance by final state.

L(P) = { w : (q0, w, Z0) reaches (p, ε, alpha) with p in F, for any alpha }

The input must be exhausted and the state must be final. The stack may hold anything at all, including everything that was ever pushed. F matters; the stack does not.

Acceptance by empty stack.

N(P) = { w : (q0, w, Z0) reaches (q, ε, ε) for any q }

The input must be exhausted and the stack must be empty. The state does not matter, and such a machine usually has no final states at all, F being empty.

Note what each ignores. That is where the trap is.

The trap: one machine, two languages

Take a machine and read it under both definitions, and the answers usually differ. Here is the smallest useful case.

munotes.in271

Acceptance by a PDA: by Final State and by Empty Stack

start s0
stack Z
final s1
s0, a, Z -> s0, A Z
s0, a, A -> s0, A A
s0, b, A -> s1, ε
s1, b, A -> s1, ε

Accepts: ab, aabb, aab, aaabb

Rejects: ε, a, b, ba, abb, abab

Under final state, this accepts a string exactly when it has at least one a, then at least one b, and no more b than a: the machine must reach s1, which needs one b, and every b must find an A to pop. So aab is accepted, with one A left on the stack, because acceptance by final state does not care what the stack holds.

Under empty stack the same machine would accept nothing at all, because Z is never popped by any rule. Add one rule and it becomes interesting; the point here is that the two readings of one machine are genuinely different questions, and aab is the witness: accepted by final state, rejected by empty stack.

So a question must say which definition it means, and an answer should say which it is using. MU's papers use final state unless they say otherwise.

From final state to empty stack

The construction. Given P accepting L by final state, build P prime accepting the same language by empty stack.

  1. Add a fresh bottom marker X0, below P's Z0, and a fresh start state s that pushes Z0 on top of X0 and hands over

to q0. This protects P from emptying its own stack by accident.

  1. Add a fresh state e, the eraser, with a rule for every stack symbol: in e, reading nothing, pop whatever is on

top. So e empties the stack.

  1. From every final state of P, add an epsilon move to e, whatever is on top.

Then a run of P that ends in a final state can slip into e and empty the stack, so P prime accepts by empty stack exactly what P accepted by final state. And a run that did not reach a final state cannot reach e, and cannot pop X0, so it cannot empty the stack.

Step 1 is the one that matters. Without X0, a run of P that happened to empty its own stack in the middle would be accepted by P prime even though P rejected it. The fresh marker makes that impossible.

Worked

Take the machine of chapter 52, which accepts a to the n b to the n by final state, and convert it.

start s
stack X
final q2
s, ε, X -> q0, Z X
q0, a, Z -> q0, A Z
q0, a, A -> q0, A A
q0, b, A -> q1, ε
q1, b, A -> q1, ε
q1, ε, Z -> q2, Z
q0, ε, Z -> q2, Z
q2, ε, Z -> e, ε
q2, ε, A -> e, ε
q2, ε, X -> e, ε
e, ε, Z -> e, ε
e, ε, A -> e, ε
e, ε, X -> e, ε
accept by empty stack
munotes.in272

Acceptance by a PDA: by Final State and by Empty Stack

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

S->aSb | ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

L(BC2) = L(BC3)

The machine has no final states in the sense that matters: the declaration says acceptance is by empty stack, and the checker runs it under that definition. The three moves from q2 into e, and the three loops in e, are the eraser, and the first line is the fresh marker being laid down.

From empty stack to final state

The construction, and it is the mirror image.

  1. Add a fresh bottom marker X0 below Z0, laid down by a fresh start state exactly as before.
  2. Add one fresh final state f.
  3. Add, from every state of P, an epsilon move that pops X0 and goes to f.

Then a run of P that empties P's own stack will have X0 on top, and only then, so it can move to f. A run that has not emptied P's stack cannot see X0 and cannot reach f.

The fresh marker is doing the detection here, not the protection. X0 becoming visible is precisely the event "P's stack is empty", and without it there would be no way to observe that event, because a machine cannot test for an empty stack: every move needs a top symbol to read.

That last sentence is worth keeping. A pushdown automaton cannot ask whether its stack is empty. The bottom marker is how the question is answered, and it is why Z0 is part of the definition.

Worked

start t0
stack Z
accept by empty stack
t0, a, Z -> t0, A Z
t0, a, A -> t0, A A
t0, b, A -> t1, ε
t1, b, A -> t1, ε
t1, ε, Z -> t1, ε
t0, ε, Z -> t0, ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

That machine accepts a to the n b to the n by empty stack: the last two rules pop the bottom marker once the A are gone. Converting it to final state acceptance:

munotes.in273

Acceptance by a PDA: by Final State and by Empty Stack

start u
stack X
final f
u, ε, X -> t0, Z X
t0, a, Z -> t0, A Z
t0, a, A -> t0, A A
t0, b, A -> t1, ε
t1, b, A -> t1, ε
t1, ε, Z -> t1, ε
t0, ε, Z -> t0, ε
t0, ε, X -> f, ε
t1, ε, X -> f, ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

L(BC5) = L(BC4)

The two extra rules at the end are the detection: whichever state the run is in when X becomes visible, it moves to f. And X becomes visible exactly when the original machine's stack is empty.

The theorem

Statement. A language is accepted by some pushdown automaton by final state if and only if it is accepted by some pushdown automaton by empty stack.

Proof. The two constructions above, one in each direction. Each preserves the language, and the four worked machines above check that they do, over every string up to the checker's bound.

So the class is the same. What differs is which machine, and the constructions cost two states and one stack symbol each way.

Distinctions

By final stateBy empty stack
WrittenL(P)N(P)
Requiresinput exhausted, state in Finput exhausted, stack empty
Ignoresthe stackthe state
F is usuallynon emptyempty
Natural fora machine designed by handthe grammar construction of chapter 58
The same machine under bothusually two different languages
The fresh marker X0, converting to empty stackconverting to final state
Its jobprotection, so P cannot empty the stack by accidentdetection, so emptiness becomes visible
Where it goesbelow Z0below Z0
Laid down bya fresh start statea fresh start state

What it does NOT mean

Acceptance by final state does not require an empty stack. The stack may hold anything, and aab above is accepted with an A left on it.

Acceptance by empty stack does not require a final state. Such machines usually have none.

The two definitions do not agree on a given machine. They agree on the CLASS of languages, which is a different statement and is what the theorem says.

A machine cannot test for an empty stack. Every move reads a top symbol, so an empty stack offers no move at all. The bottom marker is how the question is asked.

The conversions are not relabelling. Each needs a fresh bottom marker and a fresh state, and leaving the marker out breaks the first conversion silently.

munotes.in274

Acceptance by a PDA: by Final State and by Empty Stack

Quick revision

  • L(P), by final state: the input is exhausted and the state is in F. The stack is ignored.
  • N(P), by empty stack: the input is exhausted and the stack is empty. The state is ignored, and F is usually empty.
  • The same machine under the two definitions usually accepts two different languages, so a question must say which.
  • Final state to empty stack: add a marker X0 below Z0 from a fresh start state, add an eraser state reached by an

epsilon move from every final state, and let it pop everything. The marker protects P from emptying its own stack.

  • Empty stack to final state: add the same marker, and from every state an epsilon move that pops X0 into one fresh

final state. The marker makes emptiness observable.

  • Both directions preserve the language, so the two definitions characterise the same class.
  • A pushdown automaton cannot test for an empty stack, which is why Z0 is in the definition at all.

Test yourself

1. Give both definitions, and say what each ignores. By final state: the input is exhausted and the state is in F, whatever the stack holds. By empty stack: the input is exhausted and the stack is empty, whatever the state is.

2. Why does the same machine usually accept two different languages under the two definitions? Because each ignores what the other requires. A run may end in a final state with symbols still on the stack, or empty the stack in a state that is not final, and only one of the two definitions accepts in each case.

3. What is the fresh bottom marker for in each conversion? Converting to empty stack it is protection: without it a run of the original machine that emptied its own stack in the middle would be accepted wrongly. Converting to final state it is detection: the marker becoming visible is exactly the event that the original stack is empty.

4. Why can a pushdown automaton not simply test whether its stack is empty? Because every move reads the top symbol, so with an empty stack no move applies at all. The machine cannot observe the condition; it can only arrange for a marker to become visible when it holds.

5. State the theorem relating the two. A language is accepted by final state by some pushdown automaton exactly when it is accepted by empty stack by some pushdown automaton, so the two definitions characterise the same class of languages.

6. Under acceptance by final state, is a run that ends in a final state with a full stack accepting? Yes. The stack is ignored by that definition entirely.

Contents This chapter on its own page

munotes.in275

Chapter Fifty-Five

Designing a PDA: the Balanced Languages

Syllabus topic Module 2, "Pushdown Automata: Acceptance by PDA"

In one line

To design a pushdown automaton, decide what the stack will hold, then write one move per situation the machine can be in.

In the wording a student can write in an examination: the design of a pushdown automaton proceeds by identifying what the stack must record, choosing a stack alphabet to record it, dividing the input into phases and giving each phase a state, and writing a transition for every combination of state, input symbol and stack top that can arise, with epsilon transitions used to change phase and to accept.

The method

Step 1. Decide what the stack holds. This is the whole design. Ask: what has to be remembered, and does it grow without bound? Whatever grows goes on the stack, one symbol per unit.

Step 2. Divide the input into phases, and give each a state. Most of these machines have two phases, pushing and popping, and so two states plus one to accept in. A phase change is where an epsilon move or a distinguishing symbol goes.

Step 3. Write a move for every (state, input symbol, stack top) that can arise. Not every combination exists, and the ones that do not are what makes the machine reject.

Step 4. Arrange to accept. By final state, reach the final state on an epsilon move when the stack is back to the bottom marker. By empty stack, pop the marker.

Step 5. Test the boundaries, then check. The empty string, the shortest accepted string, a string with one too many of something, and a string in the wrong order.

The one rule that governs everything

A move must name the top of the stack, and there is no way to test for an empty stack. Chapter 54 said it and it is what shapes every machine below: the bottom marker Z is not decoration, it is how the machine knows it has popped everything it pushed.

Machine 1: a to the n followed by b to the n

MU's own question, in a and b.

Step 1. The number of a must be remembered, and it grows without bound. So push one A per a.

Step 2. Two phases: reading a, reading b. Two states q0 and q1, plus q2 to accept in.

start q0
stack Z
final q2
q0, a, Z -> q0, A Z
q0, a, A -> q0, A A
q0, b, A -> q1, ε
q1, b, A -> q1, ε
q1, ε, Z -> q2, Z
q0, ε, Z -> q2, Z

Accepts: ε, ab, aabb, aaabbb, aaaabbbb

Rejects: a, b, ba, aab, abb, abab, aabbb

S->aSb | ε
munotes.in276

Designing a PDA: the Balanced Languages

Accepts: ε, ab, aabb, aaabbb, aaaabbbb

Rejects: a, b, ba, aab, abb, abab, aabbb

L(BD1) = L(BD2)

Why each rejection happens, which is what an examiner asks about:

StringWhy rejected
aends in q0 with A on the stack, and the only epsilon move needs Z
aabends in q1 with an A left, same reason
abbafter the first b the machine is in q1 with Z on top, and q1 has no b move on Z
ababq1 has no move on a at all
baq0 has no b move with Z on top

Each is a missing move, and the missing moves are the design.

Machine 2: a to the n followed by b to the 2n

The same shape with one change, and the change is the point.

Step 1. Each a now has to account for two b. So push two A per a.

start r0
stack Z
final r2
r0, a, Z -> r0, A A Z
r0, a, A -> r0, A A A
r0, b, A -> r1, ε
r1, b, A -> r1, ε
r1, ε, Z -> r2, Z
r0, ε, Z -> r2, Z

Accepts: ε, abb, aabbbb, aaabbbbbb

Rejects: a, b, ab, ba, abbb, aabbb, aabb

S->aSbb | ε

Accepts: ε, abb, aabbbb, aaabbbbbb

Rejects: a, b, ab, ba, abbb, aabbb, aabb

L(BD3) = L(BD4)

Two rules changed: the pushes became double. That is the whole design change, and it is worth noticing how small it is, because it shows what the stack is really doing. It is not counting a; it is counting obligations, and each a creates two.

The general pattern. For a to the n followed by b to the kn, push k symbols per a. For a to the 2n followed by b to the n, push one A per pair of a, which needs a state to remember whether the current a is the first or second of its pair.

Machine 3: the balanced brackets

The language every compiler checks, and the one a finite automaton cannot.

Step 1. The stack holds one symbol per unmatched open bracket. Its height is the current nesting depth.

Step 2. One phase only, because open and close brackets interleave freely. So one state does the work.

start p0
stack Z
final p1
p0, (, Z -> p0, A Z
p0, (, A -> p0, A A
p0, ), A -> p0, ε
p0, ε, Z -> p1, Z

Accepts: ε, (), (()), ()(), (()()), ((()))

munotes.in277

Designing a PDA: the Balanced Languages

Rejects: (, ), )(, ((), ()), ()(

S->(S)S | ε

Accepts: ε, (), (()), ()(), (()()), ((()))

Rejects: (, ), )(, ((), ()), ()(

L(BD5) = L(BD6)

Four moves. Read the rejections: )( fails because the first symbol is a close bracket with Z on top and there is no such move, which is the machine noticing a close bracket with nothing to match. (() fails because the run ends with an A on the stack, so the epsilon move to p1 cannot fire, which is the machine noticing an unclosed bracket.

That is the whole of bracket matching, and it is worth comparing with the finite automaton that cannot do it: the stack height is the nesting depth, it is unbounded, and a fixed number of states cannot hold an unbounded number.

Machine 4: equal numbers of a and b, in any order

The language chapter 25 needed three grammar rules for.

Step 1. What has to be remembered is the difference between the counts, and which way round it is. The difference grows without bound, so it goes on the stack, one symbol per unit of imbalance. And the stack alphabet needs two symbols, A for an excess of a and B for an excess of b, because the imbalance has a direction.

Step 2. One phase, because the symbols interleave.

Step 3. Six situations: each input symbol against each of the three possible stack tops.

start n0
stack Z
final n1
n0, a, Z -> n0, A Z
n0, a, A -> n0, A A
n0, a, B -> n0, ε
n0, b, Z -> n0, B Z
n0, b, B -> n0, B B
n0, b, A -> n0, ε
n0, ε, Z -> n1, Z

Accepts: ε, ab, ba, aabb, abab, baab, abba

Rejects: a, b, aab, abb, aba, aaabb

S->aSb | bSa | SS | ε

Accepts: ε, ab, ba, aabb, abab, baab, abba

Rejects: a, b, aab, abb, aba, aaabb

L(BD7) = L(BD8)

The design in one sentence. Reading a symbol either cancels an opposite symbol on the stack or adds to the pile of its own kind, and the string is balanced exactly when the pile is empty at the end.

The two cancelling rules are the interesting ones. n0, a, B -> n0, ε says: an a arriving when there is an excess of b cancels one b. n0, b, A -> n0, ε is its mirror. Without them the machine would push seven symbols for abababa instead of one, and would never come back to Z.

munotes.in278

Designing a PDA: the Balanced Languages

Compare it with machine 1. There the a all came first, so one phase pushed and another popped and two states were needed. Here the order is free, so pushing and popping happen in the same state and the stack alphabet, not the state, records which way the imbalance runs. That is the general trade: what varies without bound goes on the stack, what is one of finitely many conditions goes in a state.

What all four have in common

The stack alphabet is chosen, not inherited. Machine 1 pushes A for an a; machine 4 pushes A or B for a direction. Neither pushes the input symbol itself, and doing so would work in machine 1 and would confuse machine 4.

The bottom marker is load bearing in all four. Every one of them accepts by an epsilon move that requires Z on top, which is the only way the machine can know the pile is empty.

Every rejection is a missing move. There is no rule that rejects; a run dies for want of a rule, and the table of rejections above is a table of situations the designer left unhandled on purpose.

Every one is proved against a grammar. Four machines, four grammars, four exact comparisons, so none of these is offered on the strength of looking right.

Distinctions

JobWhat the stack holdsStates needed
a to the n then b to the none symbol per atwo, plus one to accept in
a to the n then b to the 2ntwo symbols per atwo, plus one
balanced bracketsone symbol per unmatched openone, plus one
equal counts in any orderthe imbalance, with its directionone, plus one
Goes in a stateGoes on the stack
Becauseit is one of finitely many conditionsit grows without bound
Examplewhich phase of the input we are inhow many a are outstanding
Getting it wronga machine with infinitely many statesa machine that pushes what it did not need to

What it does NOT mean

The stack does not have to hold input symbols. Gamma is a separate alphabet and the designer chooses it.

The number of states does not grow with the input. If a design needs it to, the thing that grows belongs on the stack.

A missing move is not an omission. It is how the machine rejects, and every one of the rejections above is a deliberate gap.

The bottom marker is not optional. Without it the machine cannot detect that it has popped everything, because it cannot test for an empty stack.

munotes.in279

Designing a PDA: the Balanced Languages

A machine that works on your examples is not finished. All four above are compared with a grammar over every string up to the checker's bound, which is what makes them trustworthy.

Quick revision

  • Five steps: decide what the stack holds; divide the input into phases and give each a state; write a move for every

situation that can arise; arrange to accept; test the boundaries.

  • What grows without bound goes on the stack; what is one of finitely many conditions goes in a state.
  • a to the n b to the kn: push k symbols per a. The stack counts obligations, not symbols.
  • Brackets: one symbol per unmatched open bracket, one state, and the stack height is the nesting depth.
  • Equal counts in any order: the stack holds the IMBALANCE, with a symbol per direction, and an arriving symbol

either cancels an opposite or adds to its own pile.

  • Every rejection is a missing move; there is no rule that rejects.
  • Accept by an epsilon move requiring the bottom marker on top, which is the only way to know the pile is empty.

Test yourself

1. Design a PDA for a to the n followed by b to the 3n. The same shape as machine 2 with three pushes: q0 on a with Z pushes AAAZ, q0 on a with A pushes AAAA, then pop one A per b in a second state, and accept by an epsilon move when Z is back on top.

2. What does the stack hold in the balanced bracket machine, and what is its height? One symbol per open bracket not yet matched. Its height is the current nesting depth of the string read so far.

3. Why does the equal counts machine need two stack symbols? Because the imbalance has a direction: there may be an excess of a or an excess of b, and one symbol could not tell the two apart, so an arriving symbol would not know whether to cancel or to add.

4. Why is there no rule that rejects? Because a pushdown automaton rejects by having no move available. The designer leaves a situation unhandled and the run dies there, so the rejections are the gaps in the transition list.

5. Which of these goes in a state and which on the stack: which half of the input we are in; how many a are outstanding? The phase goes in a state, since there are two of them. The count goes on the stack, since it grows without bound.

6. Design a PDA for a to the 2n followed by b to the n. Push one A for every SECOND a, which needs two states in the first phase to remember whether the current a is the first or the second of its pair. Then pop one A per b in a third state, and accept by an epsilon move on the bottom marker.

Contents This chapter on its own page

munotes.in280

Chapter Fifty-Six

Designing a PDA: Palindromes, and Counting Two Things at Once

Syllabus topic Module 2, "Pushdown Automata: Acceptance by PDA"

In one line

A palindrome with no marker in the middle forces the machine to guess where the middle is, and guessing is exactly what nondeterminism is for.

In the wording a student can write in an examination: the language of palindromes over an alphabet of two or more symbols is accepted by a nondeterministic pushdown automaton which pushes the symbols read, changes phase by an epsilon transition at a nondeterministically chosen point, and then pops one symbol per input symbol, requiring a match. No deterministic pushdown automaton accepts it.

Why the marker mattered

Chapter 52's second machine accepted w, then c, then w reversed. The c told it when to stop pushing and start popping, and the machine was deterministic: at every point the next move was forced.

Take the marker away and the machine has a genuine problem. Reading abba, when it has read ab it cannot know whether it is halfway through a four symbol palindrome or a quarter of the way through an eight symbol one. Nothing in the input says. And it cannot look ahead, and it cannot go back.

Nondeterminism is the answer, and here it is not a convenience. Chapter 16 showed that a nondeterministic finite automaton can always be made deterministic, so nondeterminism there was only ever a shorthand. Chapter 57 proves that is false for pushdown automata, and this language is the witness. So this chapter is the one that shows the two models really are different.

Machine 1: the even length palindromes

Step 1. The stack holds the first half, one symbol per input symbol, so that it comes back off reversed.

Step 2. Two phases: pushing and popping. The change of phase happens on an epsilon move, which is where the guess lives: the machine may change phase at any point, and the definition of acceptance says the string is in the language if some choice works.

start q0
stack Z
final q2
q0, a, Z -> q0, A Z
q0, b, Z -> q0, B Z
q0, a, A -> q0, A A
q0, a, B -> q0, A B
q0, b, A -> q0, B A
q0, b, B -> q0, B B
q0, ε, Z -> q1, Z
q0, ε, A -> q1, A
q0, ε, B -> q1, B
q1, a, A -> q1, ε
q1, b, B -> q1, ε
q1, ε, Z -> q2, Z

Accepts: ε, aa, bb, abba, baab, aabbaa, abaaba

Rejects: a, b, ab, ba, aba, abab, aabb

Read the three epsilon moves in the middle. They are the guess. From q0, whatever is on top, the machine may decide the first half is over and cross to q1. It may do this at any point, including immediately, which is what accepts the empty string.

munotes.in281

Designing a PDA: Palindromes, and Counting Two Things at Once

Read the two popping moves. q1, a, A -> q1, ε pops only when the symbol read matches the symbol recorded. If they do not match there is no move and that run dies. So a run survives to the end only if the second half is the reverse of the first half as the guess divided them.

Why aabb is rejected, which is the case worth walking through. Every guess point is tried by the definition, and each fails:

Guess afterStack at the guessWhat happens next
nothingZreads a in q1 with Z on top: no move
aA Zreads a with A on top, pops; then b with Z on top: no move
aaA A Zreads b with A on top: no move
aabB A A Zreads b with B on top, pops; ends with A A Z, not Z: no accept
aabbB B A A Zno input left, Z not on top: no accept

Five possible guesses, five failures, so aabb is rejected. The checker tries all of them, which is what running a nondeterministic machine means.

And the matching grammar:

S->aSa | bSb | ε

Accepts: ε, aa, bb, abba, baab, aabbaa, abaaba

Rejects: a, b, ab, ba, aba, abab, aabb

L(BE1) = L(BE2)

Machine 2: all the palindromes, odd length as well

The odd ones have a middle symbol belonging to neither half, so the machine needs a guess that consumes one symbol as well as one that does not.

start q0
stack Z
final q2
q0, a, Z -> q0, A Z
q0, b, Z -> q0, B Z
q0, a, A -> q0, A A
q0, a, B -> q0, A B
q0, b, A -> q0, B A
q0, b, B -> q0, B B
q0, ε, Z -> q1, Z
q0, ε, A -> q1, A
q0, ε, B -> q1, B
q0, a, Z -> q1, Z
q0, b, Z -> q1, Z
q0, a, A -> q1, A
q0, a, B -> q1, B
q0, b, A -> q1, A
q0, b, B -> q1, B
q1, a, A -> q1, ε
q1, b, B -> q1, ε
q1, ε, Z -> q2, Z

Accepts: ε, a, b, aa, aba, abba, ababa, aabbaa

Rejects: ab, ba, aab, abab, aabb, abb

S->aSa | bSb | a | b | ε
munotes.in282

Designing a PDA: Palindromes, and Counting Two Things at Once

Accepts: ε, a, b, aa, aba, abba, ababa, aabbaa

Rejects: ab, ba, aab, abab, aabb, abb

L(BE3) = L(BE4)

Six new moves, and they are the second kind of guess: read one symbol, discard it, and change phase. That symbol is the middle of an odd palindrome. Notice that the machine now has two moves available from q0 on a with A on top, one that pushes and one that crosses, which is nondeterminism in its plainest form: two moves, same situation.

Compare the two grammars. BE2 has two recursive rules and one terminating rule; BE4 adds two more terminating rules, one per middle symbol. The machine's two kinds of guess correspond exactly to the grammar's two kinds of terminating rule, which is the correspondence chapter 58 will make into a construction.

Machine 3: one count against the sum of two others

The other thing the stack can do, and the boundary of what it can do.

The language. a to the n, then b to the m, then c to the n plus m, with n and m independent. So the number of c must equal the number of a and b put together.

Step 1. The stack holds one symbol per obligation, and an a and a b each create one c's worth. So push A for every a and for every b, and pop one per c. The stack alphabet needs only one symbol, because the two kinds of obligation are identical.

Step 2. Three phases: a, then b, then c. Three states plus one to accept in.

start s0
stack Z
final s3
s0, a, Z -> s0, A Z
s0, a, A -> s0, A A
s0, b, Z -> s1, A Z
s0, b, A -> s1, A A
s1, b, A -> s1, A A
s0, c, A -> s2, ε
s1, c, A -> s2, ε
s2, c, A -> s2, ε
s2, ε, Z -> s3, Z
s0, ε, Z -> s3, Z

Accepts: ε, ac, bc, abcc, aabccc, abbccc

Rejects: a, b, c, ab, ca, abc, abccc, cab, aabcccb

Read the three phases. In s0 the machine pushes an A for each a. The first b moves it to s1 and pushes an A for that b as well, and s1 pushes one per b thereafter. The first c moves it to s2 and pops, and s2 pops one per c. The epsilon move from s2 accepts when the bottom marker is back on top, and the one from s0 accepts the empty string.

munotes.in283

Designing a PDA: Palindromes, and Counting Two Things at Once

Why the block order is enforced for free. There is no move for an a in s1 or s2, and none for a b in s2, so a string whose blocks are out of order has no run at all. The last rejection above, aabcccb, fails for exactly that reason rather than for a wrong count.

Why the machine is deterministic in effect. At every point the next move is forced by the state, the symbol and the stack top, and no two rules overlap. Chapter 57 defines that properly, and the contrast with machines 1 and 2 is the point: this language needs no guessing, because the block boundaries are visible in the input.

The two things a pushdown automaton cannot do

Both were proved in chapter 50 and both are worth naming here, because a design question sometimes asks for one.

Two independent matchings. The language a to the n, b to the n, c to the n needs the a count matched against the b count and then against the c count. The stack can do one matching: the symbols pushed for the a are consumed by the b, and by the time the c arrive the record is gone. There is no second stack.

A repeated block. The language ww needs the first half remembered in order and compared in the same order. The stack gives the first half back reversed, which is why palindromes work and ww does not. That is a precise and satisfying reason: the machine's one memory is last in first out, so it is exactly suited to reversal and exactly unsuited to repetition.

The general test to apply to a design question. Ask what must be remembered and in what order it must come back. If it comes back reversed, a stack does it. If it comes back in the same order, or if two records must be kept at once, it does not, and the machine you need is in chapter 62.

Distinctions

With a marker, chapter 52Without a marker, here
The middle isannounced by the inputguessed
The machine isdeterministicnecessarily nondeterministic
Number of moves in the same situationonetwo
Can a deterministic machine do ityesno, chapter 57
PalindromesA repeated block, ww
The second half must matchthe first half REVERSEDthe first half in ORDER
A stack gives backthe reversethe reverse
Accepted by a PDAyesno
Goes on the stackCannot
one count matched against anotheryes
one count matched against a sum of twoyes, both contribute the same symbol
two counts matched independentlyno, there is one stack
munotes.in284

Designing a PDA: Palindromes, and Counting Two Things at Once

What it does NOT mean

Nondeterminism is not the machine choosing well. Chapter 14 said it: acceptance is defined as some run succeeding, and the definition needs no agent. Every guess point is a branch, and the string is accepted when one branch survives.

A machine with two moves in one situation is not broken. It is nondeterministic, which the definition permits.

Palindromes being accepted does not mean ww is. The two look alike and differ in exactly the property a stack has.

A pushdown automaton cannot be given a second stack. A machine with two stacks is as powerful as a Turing machine, which chapter 68 notes, so it is a different model and not a variant of this one.

Quick revision

  • A palindrome with no marker forces a guess at the middle, and the guess is an epsilon move that changes phase.

Acceptance is defined as some choice working.

  • Even length palindromes: push the first half, guess, then pop with a match required. Odd length needs a second

kind of guess that consumes the middle symbol.

  • This is the language proving nondeterminism cannot be removed from a pushdown automaton, which chapter 57 sets

out.

  • One count against the sum of two others is fine: both contribute the same stack symbol.
  • Two independent matchings are not, because there is one stack: a to the n b to the n c to the n is out of reach.
  • A repeated block ww is not, because a stack returns the first half REVERSED, which suits palindromes exactly and

repetition not at all.

  • The design test: what must be remembered, and in what order must it come back?

Test yourself

1. Why does a palindrome with no middle marker need a nondeterministic machine? Because the machine must stop pushing and start popping at the middle, and nothing in the input marks where the middle is. It cannot look ahead, so it must be allowed to change phase at any point, with acceptance defined as some choice working.

2. What are the two kinds of guess in the machine for all palindromes? An epsilon move that changes phase without consuming anything, which handles even lengths, and a move that consumes one symbol and discards it, which handles the middle symbol of an odd length palindrome.

3. Why can a PDA accept palindromes and not ww? Because a stack returns what was pushed in reverse order. A palindrome's second half is the reverse of its first, so the comparison matches the stack's behaviour exactly; ww's second half is the first in the same order, which the stack cannot supply.

4. Design the stack contents for a to the n b to the m c to the n plus m. Push one symbol for every a and one for every b, using the same stack symbol for both, since an a and a b each create one obligation for a c. Then pop one per c, and accept when the bottom marker is back on top.

munotes.in285

Designing a PDA: Palindromes, and Counting Two Things at Once

5. Why can a PDA not accept a to the n b to the n c to the n? Because matching the a against the b consumes the record of the a, and there is nothing left to match the c against. One stack supports one matching.

6. A design needs two things remembered at once, each growing without bound. What does that tell you? That a pushdown automaton will not do it, and that the machine needed is the Turing machine of chapter 62, whose tape is not restricted to last in first out access.

Contents This chapter on its own page

munotes.in286

Chapter Fifty-Seven

The Deterministic Pushdown Automaton

Syllabus topic Module 2, "Pushdown Automata: Definitions"

In one line

A deterministic pushdown automaton never has a choice of move, and it accepts strictly fewer languages than a nondeterministic one.

In the wording a student can write in an examination: a pushdown automaton is deterministic if for every state q, every input symbol a and every stack symbol X, the set delta(q, a, X) has at most one element, and whenever delta(q, epsilon, X) is non empty, delta(q, a, X) is empty for every input symbol a. A language accepted by a deterministic pushdown automaton is called a deterministic context free language, and the class of these is strictly smaller than the class of context free languages.

The two conditions, and why the second is needed

Condition one: at most one move per situation. For every state, input symbol and stack top, delta gives at most one pair. That is the obvious condition and it corresponds exactly to chapter 9's requirement on a DFA.

Condition two: an epsilon move excludes every reading move from the same situation. If delta(q, epsilon, X) is non empty, then delta(q, a, X) must be empty for every a.

The second condition is the one students omit, and without it the machine still has a choice. Suppose in state q with X on top there is both a move on a and an epsilon move. The machine could read the a, or it could take the epsilon move and read the a later from somewhere else. That is two different runs on one input, which is exactly what determinism forbids.

Note what is still allowed. A deterministic machine may have epsilon moves; it may not have one in a situation where it also has a reading move. And it may be incompletely specified, with no move at all in some situation, exactly as chapter 14's incomplete finite automaton was. A missing move is not a choice.

Checking a machine

Three questions, in order.

  1. Does any (state, symbol, stack top) have two moves? If so, not deterministic.
  2. Does any (state, stack top) have both an epsilon move and a reading move? If so, not deterministic.
  3. Otherwise it is deterministic.

Worked, on the machines of the last two chapters

start q0
stack Z
final q2
q0, a, Z -> q0, A Z
q0, a, A -> q0, A A
q0, b, A -> q1, ε
q1, b, A -> q1, ε
q1, ε, Z -> q2, Z
q0, ε, Z -> q2, Z

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abb, abab

Is it deterministic? Question 1: every triple appears at most once in the list, so yes. Question 2: the epsilon moves are from (q1, Z) and (q0, Z), and neither of those pairs has a reading move, because the only reading moves on Z are... none: q0, a, Z is a reading move from (q0, Z)! So question 2 fails: from q0 with Z on top the machine may read an a, or take the epsilon move to q2.

munotes.in287

The Deterministic Pushdown Automaton

So this machine is NOT deterministic, and it is worth seeing why that does not change what it accepts. The epsilon move to q2 is only useful when the input is exhausted; if there is an a to read, the branch that takes the epsilon move dies immediately, because q2 has no move on a. So the language is the same. But the machine as written violates condition two, and a question asking whether it is deterministic has the answer no.

Condition two is easy to violate by accident, and that is the lesson of this example: a machine that accepts the right language may still fail the definition, and the failure is in one pair, (q0, Z), which carries both a reading move and an epsilon move.

A machine that IS deterministic

The language here is a to the n followed by b to the n with n at least one, and dropping the empty string is what removes the offending epsilon move from the start state.

start d0
stack Z
final d3
d0, a, Z -> d1, A Z
d1, a, A -> d1, A A
d1, b, A -> d2, ε
d2, b, A -> d2, ε
d2, ε, Z -> d3, Z

Accepts: ab, aabb, aaabbb, aaaabbbb

Rejects: ε, a, b, ba, aab, abb, abab

Check it against the three questions. One: no triple appears twice in the list. Two: the only epsilon move is from (d2, Z), and the only reading move from d2 is on A, a different stack top, so that pair carries no reading move. Every other pair has reading moves only. Three: therefore deterministic.

Notice how the first a is separated out. In d0 the machine must read an a, and there is no epsilon move to compete with it, which is exactly the fault repaired. The cost is the empty string, which this language no longer contains.

S->aSb | ab

Accepts: ab, aabb, aaabbb, aaaabbbb

Rejects: ε, a, b, ba, aab, abb, abab

L(BF2) = L(BF3)

So a to the n b to the n with n at least one is a deterministic context free language, and the machine above is the proof.

The theorem: nondeterminism cannot be removed

Statement. There is a context free language accepted by no deterministic pushdown automaton. The even length palindromes over {a, b} are one.

munotes.in288

The Deterministic Pushdown Automaton

Why the proof is not given in full. It is not hard to believe and it is long to write, and it is outside what MU asks. The idea is worth having.

The idea. A deterministic machine has exactly one run on each input, so at every moment its entire knowledge is one configuration. Feed it a to the n, then b, then b, then a to the n, which is a palindrome. Feed it also a to the n, then b, then b, then a to the m for m not equal to n, which is not. A deterministic machine must accept the first and reject the second, and by the time it has read the middle it must already have committed to treating the b b as the middle. Now construct a longer palindrome in which that commitment is wrong, and the single run is stuck with it. Nondeterminism escapes this because it commits every way at once.

What follows, and it is the sentence to remember:

deterministic context free is a PROPER subset of context free

Compare chapter 16, where the corresponding containment for finite automata is an equality. The two models diverge here, and this is where.

The deterministic context free languages are closed under complement

A striking consequence, and an examinable one.

Chapter 51 proved the context free languages are NOT closed under complement. The deterministic ones are, and the construction is roughly chapter 38's: swap the accepting and non accepting states of the deterministic machine. Making that precise is fiddly, because a deterministic pushdown automaton may fail to read the whole input by getting stuck or by looping on epsilon moves, and both cases have to be handled before the swap works. But it can be done.

Why this matters. It gives a way to prove a language is not deterministic context free without any of the machinery above:

If L is context free and the complement of L is not, then L is not deterministic context free.

That is a one line test, and it works. The language of strings that are NOT of the form ww is context free; the strings that ARE of that form are not; so the first is context free with a non context free complement, and is therefore not deterministic.

What the class is good for

Every parser in every compiler is a deterministic pushdown automaton.

The reason is cost. A nondeterministic machine must in effect explore many runs, and the general algorithm for that is CYK from chapter 51, which costs the cube of the input length. A deterministic machine has one run, and it costs the length of the input. For a compiler reading a million line program the difference is not academic.

munotes.in289

The Deterministic Pushdown Automaton

So programming languages are designed so that their grammars are deterministic, and the LL and LR families mentioned in chapter 43 are exactly the grammar classes for which a deterministic machine can be built mechanically. A language feature that would make the grammar nondeterministic is usually changed rather than accepted.

Distinctions

Finite automataPushdown automata
Nondeterministic versionsame class, chapter 16STRICTLY larger class
Witness that they differnone existsthe even length palindromes
Closed under complementboth areonly the deterministic ones are
Practical consequencedeterminise and rundesign the language to be deterministic
Condition oneCondition two
Saysat most one move per (state, symbol, stack top)an epsilon move rules out reading moves from the same (state, stack top)
Violated bytwo rules with the same left sidean epsilon rule sitting beside a reading rule
Commonly forgottennoyes

What it does NOT mean

Deterministic does not mean no epsilon moves. It means no epsilon move in a situation that also has a reading move.

A missing move is not nondeterminism. An incompletely specified machine is deterministic; it simply rejects by running out of moves.

Deterministic context free is not the same as unambiguous. Every deterministic context free language has an unambiguous grammar, and the converse fails: there are unambiguous context free languages that no deterministic machine accepts.

A language, not a machine, is deterministic context free. The property is about whether SOME deterministic machine exists, so writing a nondeterministic machine for a language proves nothing about the class it is in.

The class is not small. It contains every programming language's syntax, and most languages a student meets.

Quick revision

  • Two conditions: at most one move per state, symbol and stack top; and an epsilon move from a (state, stack top)

forbids every reading move from the same pair.

  • Epsilon moves are allowed; missing moves are allowed; a choice is not.
  • The deterministic context free languages are a PROPER subset of the context free ones, and the even length

palindromes are the witness. This is where pushdown automata part company with finite automata.

  • The deterministic class IS closed under complement, unlike the full context free class, which gives a one line

test: a context free language whose complement is not context free cannot be deterministic.

  • Every compiler's parser is a deterministic pushdown automaton, because one run costs the length of the input and

the general case costs its cube.

  • Deterministic context free implies unambiguous, and not the other way round.

Test yourself

1. State both conditions for determinism. For every state, input symbol and stack symbol, delta has at most one element; and whenever delta on epsilon is non empty for a state and stack symbol, delta on every input symbol is empty for that same state and stack symbol.

munotes.in290

The Deterministic Pushdown Automaton

2. Which condition is usually forgotten, and what goes wrong without it? The second. Without it a machine may both read a symbol and take an epsilon move in the same situation, which is two runs on one input, so the machine is not deterministic even though no two rules share a left hand side.

3. Give the language that proves nondeterminism cannot be removed. The even length palindromes over an alphabet of at least two symbols. They are context free and no deterministic pushdown automaton accepts them.

4. How does this differ from the finite automaton case? For finite automata the subset construction of chapter 16 removes nondeterminism entirely, so the two classes are equal. For pushdown automata the deterministic class is strictly smaller.

5. Give the one line test for a language not being deterministic context free. If the language is context free and its complement is not, it cannot be deterministic, since the deterministic class is closed under complement.

6. Why do programming languages have deterministic grammars? Because a deterministic machine has one run and costs the length of the input, while the general context free case costs the cube of it, and a compiler reading a large program cannot afford the difference. Language features that would break determinism are normally redesigned.

Contents This chapter on its own page

munotes.in291

Chapter Fifty-Eight

From a Context Free Grammar to a Pushdown Automaton

Syllabus topic Module 2, "Pushdown Automata: PDA and CFG"

In one line

Put the sentential form on the stack, expand the variable on top by a production, and match a terminal on top against the input. That is the standard construction from a grammar to a PDA, and it is half of the pair of constructions that prove the two devices describe the same languages.

In the wording a student can write in an examination: for every context free grammar G there is a pushdown automaton P with N(P) equal to L(G). P has a single state q, its stack alphabet is the variables and terminals of G, its initial stack symbol is the start symbol, it accepts by empty stack, and its transitions are, for every production A to alpha, the move (q, epsilon, A) to (q, alpha), and for every terminal a, the move (q, a, a) to (q, epsilon).

The idea

Chapter 29 said it: in a leftmost derivation the outstanding variables are dealt with last created, first expanded, which is a stack.

Make that literal. The stack holds the part of the sentential form that has not yet been matched against the input, with its leftmost symbol on top. Then two things can happen, and they are the two kinds of move.

The top is a variable. Nothing in the input corresponds to it yet, so expand it: replace it by the right hand side of one of its productions, without reading any input. That is an epsilon move, and it is nondeterministic when the variable has several productions.

The top is a terminal. Then the derivation has committed to that terminal appearing next in the string, so the input must supply it. Read one input symbol and require it to match, popping the terminal.

When the stack is empty, the whole sentential form has been matched, and the input should be exhausted too. Hence acceptance by empty stack.

The construction

Given G = (V, T, P, S):

Part of the machineValue
statesone, call it q
input alphabetT
stack alphabetV union T
initial stack symbolS
final statesnone; acceptance is by empty stack

and two families of moves:

for every production A -> alpha: (q, ε, A) -> (q, alpha)

for every terminal a: (q, a, a) -> (q, ε)

One state. The machine's entire memory is its stack, and it needs no state at all beyond the one the definition insists on. That is the clearest possible statement of what a context free grammar is: a stack and nothing else.

alpha is written top first, which for a right hand side means leftmost symbol on top, so that the leftmost symbol is the next one dealt with.

munotes.in292

From a Context Free Grammar to a Pushdown Automaton

Acceptance is by empty stack because the natural moment to accept is when nothing is outstanding. Chapter 54 converts this to final state acceptance if a question wants it.

Worked example 1

S->aSb | ab

Accepts: ab, aabb, aaabbb

Rejects: ε, a, b, ba, aab, abab

The machine. Two productions give two epsilon moves; two terminals give two matching moves.

start q
stack S
accept by empty stack
q, ε, S -> q, a S b
q, ε, S -> q, a b
q, a, a -> q, ε
q, b, b -> q, ε

Accepts: ab, aabb, aaabbb

Rejects: ε, a, b, ba, aab, abab

L(BG2) = L(BG1)

Four moves, one state, and the machine accepts exactly what the grammar generates, which the checker decides.

A run on aabb, with the stack written top first:

q | aabb | S
q | aabb | a S b
q | abb | S b
q | abb | a b b
q | bb | b b
q | b | b
q | ε | ε

Seven configurations, six moves, and they alternate in a recognisable pattern. Move 1 expands S by the first production, putting a, S, b on the stack. Move 2 matches the a on top against the a in the input. Move 3 expands the S by the second production. Moves 4, 5 and 6 match the three remaining terminals.

Compare the stack contents with a leftmost derivation of the same string:

S
aSb
aabb

The derivation's sentential forms are aSb and aabb. The machine's stacks, with the already matched prefix put back in front, are exactly those forms. The stack is the unmatched tail of the leftmost derivation, which is the construction's whole content.

Worked example 2, with several variables

S->AB
A->aA | a
B->bB | b

Accepts: ab, aab, abb, aabb, aaabbb

Rejects: ε, a, b, ba, bab

The machine. Five productions give five epsilon moves; two terminals give two matching moves.

start q
stack S
accept by empty stack
q, ε, S -> q, A B
q, ε, A -> q, a A
q, ε, A -> q, a
q, ε, B -> q, b B
q, ε, B -> q, b
q, a, a -> q, ε
q, b, b -> q, ε

Accepts: ab, aab, abb, aabb, aaabbb

Rejects: ε, a, b, ba, bab

L(BG4) = L(BG3)

Seven moves for five productions and two terminals, and the count is the rule: one move per production, one per terminal. A grammar with p productions over t terminals gives a machine with p plus t moves, and always one state.

munotes.in293

From a Context Free Grammar to a Pushdown Automaton

A run on aab:

q | aab | S
q | aab | A B
q | aab | a A B
q | ab | A B
q | ab | a B
q | b | B
q | b | b
q | ε | ε

Eight configurations. Notice the alternation: expand, match, expand, match. And notice that at the third configuration the machine chose A to aA rather than A to a. That was a guess, and on this input it was the right one; the run that guessed A to a would have reached q | ab | B and then had to match an a against a b, which has no move.

Why the machine must be nondeterministic

Because a variable with several productions gives several epsilon moves from the same situation, and nothing tells the machine which to take.

That is not a defect of the construction; it is the construction being faithful. A grammar with a choice of productions genuinely has a choice, and a parser has to resolve it somehow. Chapter 57 said how the practical world resolves it: restrict the grammars so that the choice is determined by the next input symbol, which is what LL parsing is.

Two consequences worth stating.

The machine is deterministic only by accident. BG2 above has two moves from (q, epsilon, S) and is nondeterministic.

And the machine may loop. If the grammar is left recursive, with A deriving A something, then the machine can expand A to A something for ever without reading any input. The run never fails and never accepts; it just continues. That is why Greibach normal form matters, and it is the next section.

Greibach normal form, and why it is the right preparation

Chapter 49 built a form in which every production is a terminal followed by variables. Apply this chapter's construction to such a grammar and something good happens.

Every expansion puts a terminal on top. The right hand side is a terminal then variables, and the terminal goes on top, so the very next move must be a matching move, which reads an input symbol.

So every two moves consume one input symbol, and a run on a string of length n takes exactly 2n moves. It cannot loop, and it cannot run for ever.

That is what chapter 49 promised, and it is worth doing once.

S->aSb | ab

Its Greibach normal form needs the middle b given a variable, by chapter 49's third example:

S->aSR | aR
R->b

Accepts: ab, aabb, aaabbb

munotes.in294

From a Context Free Grammar to a Pushdown Automaton

Rejects: ε, a, b, ba, aab, abab

L(BG6) = L(BG5)

BG6 = GNF(BG5)

start q
stack S
accept by empty stack
q, ε, S -> q, a S R
q, ε, S -> q, a R
q, ε, R -> q, b
q, a, a -> q, ε
q, b, b -> q, ε

Accepts: ab, aabb, aaabbb

Rejects: ε, a, b, ba, aab, abab

L(BG7) = L(BG5)

Now every expansion of S puts an a on top, and the next move must read an a. The machine still guesses which production to use, but it can no longer make progress without consuming input, so every run terminates in at most 2n moves on an input of length n.

The standard version of this construction goes further and folds the matching move into the expansion, giving a machine whose every move reads an input symbol. It is the same idea with one fewer move per step and the same one state.

Distinctions

The top of the stack is a variablea terminal
The moveexpand by a productionmatch against the input
Input consumednoneone symbol
Nondeterministicwhen the variable has several productionsnever
Corresponds toa derivation stepa symbol of the string being confirmed
An arbitrary grammarGreibach normal form
Moves per input symbolunboundedexactly two
Can loopyes, if left recursiveno
Number of moves on a string of length nunboundedexactly 2n

What it does NOT mean

The machine is not a parser. It is a recogniser, and it guesses. A parser resolves the guesses, which needs either a restricted grammar or the CYK algorithm of chapter 51.

One state is not a simplification. It is the construction's real content: a context free grammar needs a stack and nothing else.

Acceptance by empty stack is not optional here. It is what "nothing outstanding" means. Chapter 54 converts if a question wants final states.

Left recursion is not an error in the grammar. It is an error in this machine, which may loop on it, and chapter 49's conversion is the remedy.

The number of moves is not the number of productions. It is the number of productions plus the number of terminals.

Quick revision

  • One state. Stack alphabet is the variables and the terminals. Initial stack symbol is S. Acceptance by empty

stack. No final states.

  • One epsilon move per production, (q, epsilon, A) to (q, alpha) with alpha top first; one matching move per

terminal, (q, a, a) to (q, epsilon).

  • The stack holds the unmatched tail of a leftmost derivation, leftmost symbol on top.
  • Expand when the top is a variable, match when it is a terminal, accept when the stack empties.
  • The machine is nondeterministic whenever a variable has several productions, which is faithful to the grammar.
  • On a left recursive grammar it can loop. In Greibach normal form every expansion puts a terminal on top, so every
munotes.in295

From a Context Free Grammar to a Pushdown Automaton

two moves read one symbol and a run on a string of length n takes exactly 2n moves.

Test yourself

1. Give the two families of moves. For each production A to alpha, an epsilon move from (q, epsilon, A) to (q, alpha) with alpha written top first; and for each terminal a, a matching move from (q, a, a) to (q, epsilon).

2. How many states does the machine have, and what does that tell you? One. It tells you that a context free grammar needs no finite memory beyond its stack, which is the exact content of the correspondence in chapter 29.

3. Build the machine for S->aS | b. One state q, initial stack symbol S, acceptance by empty stack, and four moves: (q, eps, S) to (q, aS); (q, eps, S) to (q, b); (q, a, a) to (q, eps); (q, b, b) to (q, eps).

4. What is on the stack part way through a run, and how does it relate to a derivation? The part of the current sentential form that has not yet been matched against the input, leftmost symbol on top. Putting the already matched prefix back in front of it gives exactly a sentential form of a leftmost derivation.

5. Why can the machine loop, and what prevents it? Because a left recursive grammar lets the machine expand a variable into itself for ever without reading any input. Converting the grammar to Greibach normal form prevents it, since every expansion then puts a terminal on top and the next move must read.

6. How many moves does a grammar with 6 productions over 3 terminals give? Nine: one per production and one per terminal, and one state.

Contents This chapter on its own page

munotes.in296

Chapter Fifty-Nine

From a Pushdown Automaton to a Context Free Grammar

Syllabus topic Module 2, "Pushdown Automata: PDA and CFG"

In one line

Make a variable for every claim of the form "the machine goes from this state to that one, popping exactly this stack symbol", and let the productions say how such a journey is made. That is the construction from a PDA to a grammar, and it is the harder half of the pair.

In the wording a student can write in an examination: given a pushdown automaton P accepting by empty stack, define a grammar whose variables are triples written [q X p], meaning that P can go from state q with X on top of the stack to state p having removed exactly that X and everything it pushed above it. The start symbol derives [q0 Z0 p] for every state p, and each move of P gives a family of productions. Then L(G) equals N(P).

Why a triple, and why it is not obvious

The difficulty is that a grammar variable stands for a set of strings, and a machine's state does not. A state is a moment; what has to be captured is a whole stretch of a run.

So the right thing for a variable to stand for is not a state but a journey, and the journey that matters is the one a stack symbol has. A symbol is pushed, things happen above it, and eventually it is popped. Everything between its being on top and its being removed is one self contained piece of the run, because chapter 53's second fact says the machine cannot see below it.

That piece is what a variable will stand for, and it needs three things to be named: where the machine was when the symbol came to the top, which symbol it is, and where the machine is when it is removed. Hence a triple.

[q X p] derives exactly the strings w such that (q, w, X) reaches (p, ε, ε)

Read it: starting in q with just X on the stack, the machine can consume w and end in p with the stack empty.

The construction

Let P be (Q, Sigma, Gamma, delta, q0, Z0, empty), accepting by empty stack.

The variables. One for every triple [q X p] with q and p in Q and X in Gamma, plus a fresh start symbol S.

The terminals. Sigma.

The start productions. For every state p:

S -> [q0 Z0 p]

Read it: the machine starts in q0 with Z0 on the stack, and must remove it, ending in some state, and any state will do because acceptance is by empty stack.

The move productions. For each move (q, a, X) to (r, Y1 Y2 ... Yk), with a a symbol or epsilon:

munotes.in297

From a Pushdown Automaton to a Context Free Grammar

If k is 0, the move pops X and pushes nothing, so the journey of X ends here:

[q X r] -> a

If k is at least 1, the move replaces X by k symbols, so the journey of X is made of the journeys of those k symbols, one after another. Choose intermediate states p1 to p(k minus 1) and a final state p, and write:

[q X p] -> a [r Y1 p1] [p1 Y2 p2] ... [p(k-1) Yk p]

One production for every choice of the intermediate states, which is why the construction produces so many. With n states and a push of k symbols, one move gives n to the k productions.

Why the intermediate states are guessed. The grammar has to describe every possible run, and the state the machine is in when Y1 is finally popped determines where the journey of Y2 begins. The grammar does not know it, so it writes a production for every possibility, and the ones that describe impossible runs turn out to be useless symbols and are removed by chapter 45.

Worked example

Take a machine accepting a to the n followed by b to the n, with n at least 1, by empty stack. Two states, two stack symbols.

start p
stack Z
accept by empty stack
p, a, Z -> p, A Z
p, a, A -> p, A A
p, b, A -> q, ε
q, b, A -> q, ε
q, ε, Z -> q, ε

Accepts: ab, aabb, aaabbb, aaaabbbb

Rejects: ε, a, b, ba, aab, abb, abab

The variables. Two states and two stack symbols give eight triples, plus S. Not all eight will survive.

The productions, move by move.

Move p, b, A -> q, ε. Pushes nothing, so k is 0 and the production is

[pAq] -> b

Move q, b, A -> q, ε. Likewise:

[qAq] -> b

Move q, ε, Z -> q, ε. Likewise, with a equal to epsilon:

[qZq] -> ε

Move p, a, A -> p, A A. Pushes two symbols, so k is 2 and one intermediate state is chosen, and one final state, giving four productions:

[pAp] -> a [pAp] [pAp]

[pAp] -> a [pAq] [qAp]

[pAq] -> a [pAp] [pAq]

[pAq] -> a [pAq] [qAq]

Move p, a, Z -> p, A Z. Also k equal to 2, with Y1 equal to A and Y2 equal to Z, giving four more:

[pZp] -> a [pAp] [pZp]

[pZp] -> a [pAq] [qZp]

[pZq] -> a [pAp] [pZq]

[pZq] -> a [pAq] [qZq]

The start productions. Two states, so two:

munotes.in298

From a Pushdown Automaton to a Context Free Grammar

S -> [pZp]

S -> [pZq]

The grammar in full, thirteen productions over seven variables:

vars [pAp] [pAq] [pZp] [pZq] [qAq] [qZq]
S -> [pZp] | [pZq]
[pAp] -> a [pAp] [pAp] | a [pAq] [qAp]
[pAq] -> a [pAp] [pAq] | a [pAq] [qAq] | b
[pZp] -> a [pAp] [pZp] | a [pAq] [qZp]
[pZq] -> a [pAp] [pZq] | a [pAq] [qZq]
[qAq] -> b
[qZq] -> ε

Accepts: ab, aabb, aaabbb, aaaabbbb

Rejects: ε, a, b, ba, aab, abb, abab

L(BH2) = L(BH1)

The grammar accepts exactly what the machine does, which the checker decides.

Now simplify it, by chapter 45. Which variables are generating?

[pAq] is, by [pAq] to b. [qAq] is, by [qAq] to b. [qZq] is, by the epsilon production. [pZq] is, by [pZq] to a [pAq] [qZq], since both of those are generating.

[pAp], [pZp] and [qAp] are not. Look at [pAp]: both its productions mention [pAp] or [qAp], and [qAp] has no production at all. So neither can ever produce a terminal string.

Delete them and the productions mentioning them:

vars [pAq] [pZq] [qAq] [qZq]
S -> [pZq]
[pAq] -> a [pAq] [qAq] | b
[pZq] -> a [pAq] [qZq]
[qAq] -> b
[qZq] -> ε

Accepts: ab, aabb, aaabbb, aaaabbbb

Rejects: ε, a, b, ba, aab, abb, abab

L(BH3) = L(BH1)

Five productions instead of thirteen, four variables instead of seven, and the same language. Most of what the construction produces is useless and the simplification of chapter 45 is not optional in practice.

And it is now readable. [qZq] derives the empty string, [qAq] derives a single b, and [pAq] derives a to the n followed by b to the n with n at least one, which is the language itself. The triples were opaque when written down and turn out to mean something after all.

The theorem, and the pair with chapter 58

Statement. A language is context free if and only if it is accepted by some pushdown automaton.

Proof. Chapter 58 turns a grammar into a machine; this chapter turns a machine into a grammar. Both preserve the language, and chapter 54 converts between the two acceptance conditions, so every combination is covered.

That is the second row of chapter 29's table, proved. And it is worth noticing which of the four rows are now fully proved in this book: the regular row by chapter 40, and this one by chapters 58 and 59.

The cost

QuantityValue
variablesthe number of states squared, times the stack alphabet size, plus 1
productions from a move pushing k symbolsthe number of states to the power k
after simplificationusually a small fraction of that
munotes.in299

From a Pushdown Automaton to a Context Free Grammar

With 3 states and a move pushing 2 symbols, one move gives 9 productions. With 5 states, 25. The construction is correct and it is not practical for a large machine, which is why compilers go the other way, from grammar to machine.

Distinctions

Chapter 58, grammar to machineThis chapter, machine to grammar
Result hasone state, p plus t movesmany variables, many productions
Sizelinear in the grammarpolynomial, and large
Acceptanceby empty stackassumes empty stack
Used in practiceyes, this is what a parser generator doesno
Difficultymechanicalneeds the triple idea
A state of the machineA variable of the grammar
Stands fora momenta whole journey: from one state to another, popping one symbol
Number of themthe state countstates squared times stack symbols

What it does NOT mean

A variable is not a state. It is a triple, and the triple names a stretch of a run rather than a moment.

The intermediate states are not known. They are guessed, one production per guess, and the wrong guesses become useless symbols.

The construction does not assume the machine is deterministic. It describes every run, which is what the grammar has to do.

The raw output is not the answer a question wants. Simplify it by chapter 45, and say that you have.

Acceptance must be by empty stack. If the machine accepts by final state, convert it first by chapter 54.

Quick revision

  • Variables are triples [q X p], meaning the machine goes from q with X on top to p having removed exactly that X

and everything pushed above it.

  • Start productions: S to [q0 Z0 p] for every state p.
  • A move (q, a, X) to (r, epsilon) gives [q X r] to a.
  • A move (q, a, X) to (r, Y1 ... Yk) gives, for every choice of intermediate states, [q X p] to a followed by the k

triples chaining through them.

  • One move pushing k symbols gives n to the k productions with n states, so the output is large and mostly useless.
  • Simplify by chapter 45; the worked example drops from thirteen productions to five.
  • With chapter 58 this proves the second row of chapter 29's table: context free exactly matches pushdown automata.

Test yourself

1. What does the variable [q X p] stand for? The set of strings the machine can consume while going from state q, with X on top of the stack, to state p, having removed exactly that X and everything pushed above it in the meantime.

munotes.in300

From a Pushdown Automaton to a Context Free Grammar

2. Give the production a move that pops without pushing produces. A move (q, a, X) to (r, epsilon) gives [q X r] to a, with a being the empty string if the move consumes nothing.

3. Why is one production written per choice of intermediate state? Because the state the machine is in when the first pushed symbol is finally popped determines where the next symbol's journey begins, and the grammar cannot know which state that will be, so it must allow all of them.

4. How many productions does a move pushing 2 symbols give, in a machine with 4 states? Sixteen, being 4 squared: one for each choice of the single intermediate state and the final state.

5. Why are so many of the variables useless? Because most triples describe journeys the machine cannot actually make, and a variable for an impossible journey has no way to derive a terminal string. Chapter 45's generating test removes them.

6. What do chapters 58 and 59 prove together? That a language is context free exactly when some pushdown automaton accepts it, which is the second row of the correspondence table of chapter 29.

Contents This chapter on its own page

munotes.in301

Chapter Sixty

The Linear Bounded Automaton Model

Syllabus topic Module 2, "Linear Bound Automata: The Linear Bound Automata Model"

In one line

A linear bounded automaton is a Turing machine that may not use more tape than its input occupies.

In the wording a student can write in an examination: a linear bounded automaton is a nondeterministic Turing machine whose tape head is confined to the portion of the tape holding the input, bounded at each end by an end marker which the machine may read but not overwrite and may not move beyond. Formally it is an eight tuple (Q, Sigma, Gamma, delta, q0, <, >, F) where the two extra components are the left and right end markers.

Where it sits, and why it exists

Chapter 50 proved a to the n b to the n c to the n is not context free, and chapter 23 gave a grammar for it, which chapter 26 classified as type 1. So there is a class of languages above the context free ones, defined by a kind of grammar, and the four row table of chapter 29 promised a machine for it.

This is that machine, and the restriction defining it is a space bound.

Why space and not something else. Chapter 29 gave the reason: a type 1 production never shortens the string, so in a derivation of a string of length n no sentential form is ever longer than n. The derivation therefore fits in the space the string occupies, and a machine that simulates the derivation needs exactly that much room. The restriction on the machine is the restriction on the grammar, read on the tape instead of on the page.

The definition

An eight tuple, and the two new parts are the markers.

PartWhat it is
Qa finite set of states
Sigmathe input alphabet
Gammathe tape alphabet, containing Sigma and the two markers
deltathe transition function, nondeterministic
q0the initial state
the left markerwritten at the cell before the input
the right markerwritten at the cell after the input
Fthe set of final states

The tape at the start holds the left marker, then the input, then the right marker, and the head is on the first symbol of the input.

The two rules the markers enforce. The machine may read a marker and act on it, but it may not overwrite either of them, and it may not move past either of them. So the reachable cells are exactly the ones the input occupies, plus the two the markers are in.

Everything else is chapter 62's Turing machine, which the next block defines properly: a move reads the cell under the head, writes a symbol into it, moves the head one cell left or right, and changes state. The only difference is the confinement.

munotes.in302

The Linear Bounded Automaton Model

It is nondeterministic by definition, and that matters: see the open question below.

Why "linear" bounded

Because the usual definition allows a little more room than the input, and the amount is a constant multiple.

The tape alphabet may be larger than the input alphabet. So a cell can hold more than one input symbol's worth of information: with a tape alphabet of pairs, one cell can hold two tracks, and with triples, three. A machine allowed k tracks has effectively k times the input's length of working space, and k is a constant fixed by the machine.

So the space available is c times n for a constant c, which is linear in n, and that is where the name comes from. Confining the head to exactly n cells and allowing a richer alphabet are the same thing up to a constant, and every treatment allows one or the other.

What linear rules out. Space that grows faster than the input: n squared cells, or unbounded cells. A machine with unbounded tape is chapter 62's Turing machine and is strictly stronger.

What the bound buys and costs

It buys decidability of membership. Chapter 61 proves it: with a bounded tape there are finitely many configurations, so a run that goes on long enough must repeat one, and a machine that detects the repeat can stop. That is the property the full Turing machine lacks and it is the whole difference between chapter 61 and chapter 74.

It costs the ability to compute freely. A machine that needs scratch space beyond its input cannot have it, and many natural computations do.

The bargain is the same one chapter 26 described for grammars, in the other direction: tighten the restriction, lose power, gain decidability.

The machine as a worked object

A linear bounded automaton is a Turing machine with a fence, and this book's Turing machine block begins in chapter

  1. Rather than print a transition table here for a machine whose notation has not yet been defined, this chapter

describes what one does and chapter 61 gives the correspondence. The shape of a machine for a to the n b to the n c to the n is worth having in words, because MU's note question can be answered with it.

The task. Decide whether the input is a block of a, then a block of b, then a block of c, with the three blocks of equal length.

The method, in the implementation description style chapter 65 will define.

  1. Sweep left to right and check the input is of the form a star, b star, c star. If not, reject. This uses no extra
munotes.in303

The Linear Bounded Automaton Model

space: it is a finite automaton's job carried out by the states.

  1. Return to the left end.
  2. Repeat: find the leftmost unmarked a and mark it; find the leftmost unmarked b and mark it; find the leftmost

unmarked c and mark it. If any of the three cannot be found while the others could, reject.

  1. When no unmarked a, b or c remains, accept.

Why it stays inside the bound. The only writing is the marking, and marking replaces a symbol in a cell it already occupies. The tape alphabet gains three symbols, the marked forms of a, b and c, and no cell is ever added. So the head never needs to leave the input, and the machine is a linear bounded automaton.

Why a pushdown automaton could not do this. It has no way to mark and come back. Its single stack gives access to one end only, and the marking method depends on crossing the input repeatedly, which needs a head that can move both ways over cells that stay where they are.

That contrast is the clearest statement of what the extra power is: a stack is read at one end and destroyed as it is read; a tape can be revisited.

The open question

This is worth knowing because it is unusual for a syllabus to contain one, and because a student who states it correctly has said something true that most textbooks state carelessly.

Is a deterministic linear bounded automaton as powerful as a nondeterministic one?

It is not known. Immerman's 1988 paper, which settled the other question about this machine, records the position in its own closing section: "We still do not know whether nondeterministic space is equal to deterministic space". Compare the three rows around it:

ModelDeterministic equals nondeterministic?
finite automatonyes, chapter 16
pushdown automatonNO, chapter 57
linear bounded automatonopen
Turing machineyes, chapter 69

Two settled yes, one settled no, and one open, which is why the row is worth a sentence rather than an assumption.

A second question about this machine WAS settled, and late: whether the context sensitive languages are closed under complement. It was raised by Kuroda in 1964, in the same paper that introduced this machine, and answered yes twenty four years later by Neil Immerman, in "Nondeterministic Space is Closed Under Complementation", SIAM Journal of Computing volume 17 number 5 (1988), pages 935 to 938. His paper's own opening says it settles "a question raised by Kuroda in 1964". Chapter 61 states the result, and the point for a student is that a textbook printed before 1988 will say the question is open.

Distinctions

Pushdown automatonLinear bounded automatonTuring machine
Memorya stacka tape as long as the inputan unbounded tape
Accessone end only, destructiveanywhere within the bound, revisitableanywhere
Grammartype 2type 1type 0
Membership decidableyesyesno
Deterministic equals nondeterministicnoopenyes
munotes.in304

The Linear Bounded Automaton Model

The end markersThe input
May be readyesyes
May be overwrittennoyes
May be moved pastnonot applicable
Their jobto make the bound enforceableto be decided

What it does NOT mean

Linear bounded does not mean the tape is exactly the input length. A richer tape alphabet gives a constant multiple of it, and the two formulations are equivalent.

It is not a pushdown automaton with a bigger stack. The access pattern is different: a tape can be revisited and a stack cannot, and that is where the extra power comes from.

It is not a Turing machine with a time limit. The bound is on SPACE. A linear bounded automaton may run for a very long time, and chapter 61 uses exactly that fact.

The markers are not part of the input. They are written by the definition before the machine starts, and the input is what lies between them.

Deterministic and nondeterministic are not known to be equivalent here. Saying they are is a common textbook slip; the question is open.

Quick revision

  • A linear bounded automaton is a nondeterministic Turing machine whose head may not leave the portion of tape the

input occupies, fenced by two end markers it may read but not overwrite or pass.

  • Eight tuple: the Turing machine's seven parts with the two markers added.
  • "Linear" because a richer tape alphabet gives a constant multiple of the input's length, which is the same

restriction up to a constant.

  • The bound comes from the grammar: a type 1 production never shortens, so a derivation of a string of length n

never exceeds n symbols.

  • It buys decidable membership, because a bounded tape has finitely many configurations. It costs free scratch

space.

  • It can mark cells and revisit them, which is what a stack cannot do and what makes a to the n b to the n c to the

n reachable.

  • Whether the deterministic version is as strong is not known; Immerman's 1988 paper records it as open.
  • Closure of the context sensitive languages under complement was raised by Kuroda in 1964 and settled yes by

Immerman in 1988, so a textbook older than that will call it open.

Test yourself

1. Define a linear bounded automaton. A nondeterministic Turing machine whose tape head is confined between two end markers written immediately before and after the input, which it may read but may neither overwrite nor move past, so that it uses no more tape than the input occupies.

munotes.in305

The Linear Bounded Automaton Model

2. Why is it called linear rather than exactly bounded? Because the tape alphabet may be larger than the input alphabet, so one cell can hold several symbols' worth of information, giving a constant multiple of the input's length. Up to that constant the two formulations are the same.

3. Where does the bound come from, in terms of grammars? From the type 1 restriction that no production shortens the string: a derivation of a string of length n never produces a sentential form longer than n, so the whole derivation fits in the space the string occupies.

4. What can this machine do that a pushdown automaton cannot, and why? Mark cells and come back to them. A stack is read at one end and destroyed as it is read, while a tape stays where it is and can be crossed repeatedly, which is what matching three counts needs.

5. Which question about this machine is open, and which one was settled late? Whether a deterministic linear bounded automaton is as powerful as a nondeterministic one is not known; Immerman's 1988 paper records it as open in its closing section. Whether the context sensitive languages are closed under complement was raised by Kuroda in 1964 and answered yes by Immerman in 1988.

6. Is the restriction on time or on space? On space. The machine may run for a very long time; it may not use more tape than the input occupies, and chapter 61 uses the long running to prove membership decidable anyway.

Contents This chapter on its own page

munotes.in306

Chapter Sixty-One

Linear Bounded Automata and the Context Sensitive Languages

Syllabus topic Module 2, "Linear Bound Automata: Linear Bound Automata and Languages."

In one line

A language is context sensitive exactly when some linear bounded automaton accepts it, and unlike the class below it, membership is decidable and the class is closed under complement.

In the wording a student can write in an examination: the class of languages accepted by nondeterministic linear bounded automata is exactly the class of context sensitive languages, that is, the type 1 languages of the Chomsky hierarchy. The membership problem for this class is decidable, the class is closed under complement, and it is properly contained in the recursive languages.

The equivalence

Statement. A language is generated by a type 1 grammar if and only if it is accepted by a linear bounded automaton.

Kuroda proved it in 1964, in the paper that introduced the machine, and Immerman's 1988 paper states the result in that form: "Kuroda showed in 1964 that CSL = NSPACE[n]", citing Kuroda's "Classes of Languages and Linear-Bounded Automata", Information and Control 7 (1964), pages 207 to 233. NSPACE[n] is the class of languages a nondeterministic machine decides in space linear in the input, which is this machine.

From a grammar to a machine

Given a type 1 grammar G and an input w of length n, the machine works backwards, from the string towards the start symbol.

  1. Write w on the tape, between the markers.
  2. Repeat: choose nondeterministically a position and a production alpha to beta of G whose right hand side beta

matches the tape there, and replace that occurrence by alpha.

  1. If the tape ever holds just the start symbol, accept.

Why it stays inside the bound. Every production of a type 1 grammar has its right hand side at least as long as its left, so replacing beta by alpha never makes the tape longer. It may make it shorter, which leaves blank cells, and those are inside the original input's space. So the head never needs to go beyond the markers.

That is the whole reason the bound is the right one: it is the grammar's non shortening restriction, read backwards.

Why it is nondeterministic. The machine must guess which production to undo and where. That is the same guessing as chapter 58's expansion, in reverse, and it is why the machine's definition has nondeterminism in it.

From a machine to a grammar

The other direction builds a grammar whose derivation simulates a run of the machine. Its variables carry pairs: the original input symbol and the current tape symbol, so that the derivation can remember what the input was while rewriting what the tape holds. Productions move a state marker along and rewrite cells exactly as the machine's moves do, and the non shortening restriction is satisfied because the machine never uses more cells than it started with.

munotes.in307

Linear Bounded Automata and the Context Sensitive Languages

The construction is long and MU does not set it. What she does set is the statement and the reason, which is the paragraph above about the bound being the restriction read backwards.

Membership is decidable, and the argument is worth knowing

Theorem. Given a linear bounded automaton M and a string w, it is decidable whether M accepts w.

Proof. Count the configurations. A configuration of M on an input of length n consists of the state, the head position and the tape contents. With q states, a tape alphabet of size g and n cells:

PartNumber of possibilities
the stateq
the head positionn plus 2, counting the two marker cells
the tape contentsg to the n

so the total is q times (n plus 2) times g to the n, which is finite.

Now run M on w, exploring its nondeterministic branches, and keep a record of the configurations seen. If a branch accepts, accept. If every branch has either halted or reached a configuration already seen on that branch, reject: a repeated configuration means the branch is in a loop and will never do anything new.

Since there are finitely many configurations, no branch can go on for ever without repeating, so the procedure always terminates.

This is the argument the full Turing machine cannot make, and the difference is one word: finite. An unbounded tape has infinitely many configurations, so a machine may run for ever without repeating one, and chapter 74 proves that no method can tell in advance whether it will. The bound is what turns that into a decision procedure, and it is the single most important consequence of the restriction.

The cost. The number of configurations is exponential in n, so the procedure is decidable and hopeless in practice. Decidable does not mean cheap, as chapter 39 said.

Closure under complement, settled late

Chapter 51 proved the context free languages are not closed under complement. The context sensitive languages are, and the history is worth a sentence because a student's textbook may predate the answer.

Kuroda raised the question in 1964, in the paper that introduced the machine. It was answered yes by Neil Immerman in "Nondeterministic Space is Closed Under Complementation", SIAM Journal of Computing volume 17 number 5 (1988), pages 935 to 938. The paper's own opening states that it settles "a question raised by Kuroda in 1964", and its main theorem is more general than the LBA case: nondeterministic space is closed under complementation for any space bound at least logarithmic, and closure of the context sensitive languages follows at once.

munotes.in308

Linear Bounded Automata and the Context Sensitive Languages

Twenty four years, and a result that changed a row of the table every textbook prints.

The whole picture, with this row filled in

ClassGrammarMachineMembershipClosed under complementClosed under intersection
regulartype 3finite automatondecidableyesyes
context freetype 2pushdown automatondecidablenono
context sensitivetype 1linear bounded automatondecidableyes, since 1988yes
recursivenonealways halting Turing machinedecidableyesyes
recursively enumerabletype 0Turing machineundecidablenoyes

Two things to read off it. The context free row is the odd one out for closure, and the recursively enumerable row is the odd one out for decidability. Everything in between behaves well.

And the containments are all proper, with the witnesses this book has proved:

StepWitnessProved in
regular inside context freea to the n b to the nchapters 21 and 37
context free inside context sensitivea to the n b to the n c to the nchapter 50
context sensitive inside recursiveexists, by a diagonal argumentchapter 7's technique
recursive inside recursively enumerablethe halting languagechapter 74

The third is the one this book states without working: a diagonal argument over the context sensitive languages, which can be enumerated because their grammars can, produces a recursive language outside the class. Chapter 7 gave the technique and chapter 74 gives the model for writing such an argument out.

What a linear bounded automaton can do that a pushdown automaton cannot

Three examples, and each is a language this book has already proved out of reach of the class below.

a to the n b to the n c to the n. Chapter 50 proved it not context free. Chapter 60 described the marking machine for it, which uses no space beyond the input.

ww. Chapter 50 proved it not context free. A linear bounded automaton finds the middle by marking from both ends alternately, then compares the two halves symbol by symbol, crossing back and forth. Every cell it uses was already there.

Strings whose length is a perfect square. Chapter 37 proved it not regular and chapter 50 proved it not context free. A linear bounded automaton can compute a square on a second track of its own tape and compare, all within a constant multiple of the input's length.

The common feature is revisiting. Each needs the input read more than once, or marked and returned to, which a stack cannot do and a bounded tape can.

Distinctions

Pushdown automatonLinear bounded automaton
Memorya stack, one ended and destructivea tape, revisitable, bounded by the input
Classcontext freecontext sensitive
Closed under complementnoyes
Closed under intersectionnoyes
Membershipdecidable, cubic by CYKdecidable, exponential
Deterministic equals nondeterministicnonot known
munotes.in309

Linear Bounded Automata and the Context Sensitive Languages

Why membership is decidable hereWhy not for a Turing machine
Configurationsfinitely many, so a branch must repeatinfinitely many, so it need not
The procedurerun and detect a repeatno such procedure exists, chapter 74

What it does NOT mean

Decidable does not mean practical. The configuration count is exponential in the input's length, so the procedure is a proof and not an algorithm anybody runs.

Closed under complement does not mean the complement is easy to find. It is an existence statement about the class.

The equivalence is with the NONDETERMINISTIC machine. Whether the deterministic version gives the same class is not known, which chapter 60 recorded from Immerman's own closing section.

Context sensitive does not mean the rules inspect context. Chapter 26 said it: the defining restriction is that no production shortens the string, and the context formulation is an equivalent way of saying the same thing.

A textbook older than 1988 will get the complement row wrong. That is not the textbook's fault and it is worth knowing before an examination.

Quick revision

  • Kuroda 1964: the context sensitive languages are exactly the languages a linear bounded automaton accepts.
  • Grammar to machine: write the string, and repeatedly undo a production, accepting if the start symbol is reached.

It stays in the bound because no type 1 production shortens, so undoing one never lengthens the tape.

  • Membership is decidable: the number of configurations is q times (n plus 2) times g to the n, which is finite, so

a branch that has not halted must eventually repeat a configuration and can be cut off.

  • That argument is unavailable to a Turing machine because an unbounded tape has infinitely many configurations,

which is the whole difference between this chapter and chapter 74.

  • Closure under complement: raised by Kuroda in 1964, settled yes by Immerman in 1988.
  • Three languages this machine reaches and a pushdown automaton does not: a to the n b to the n c to the n, ww, and

the perfect square lengths. The common feature is revisiting the input.

Test yourself

1. State the equivalence and say who proved it and when. A language is context sensitive exactly when a linear bounded automaton accepts it. Kuroda proved it in 1964, in "Classes of Languages and Linear-Bounded Automata", Information and Control volume 7, pages 207 to 233.

2. Describe the machine built from a type 1 grammar, and say why it stays in the bound. Write the input on the tape and repeatedly guess a production and a position where its right hand side occurs, replacing that occurrence by the left hand side; accept if the tape reduces to the start symbol. It stays in the bound because no type 1 production shortens, so undoing one never lengthens the tape.

munotes.in310

Linear Bounded Automata and the Context Sensitive Languages

3. Prove that membership is decidable for this class. With q states, a tape alphabet of size g and an input of length n there are q times (n plus 2) times g to the n configurations, which is finite. Run the machine and cut off any branch that repeats a configuration, since it is looping. Every branch must halt or repeat, so the procedure terminates.

4. Why does that argument fail for a Turing machine? Because an unbounded tape gives infinitely many configurations, so a run may go for ever without repeating one, and there is no bound after which a repeat is guaranteed.

5. Which closure result about this class was settled late, and when? Closure under complement. Kuroda raised the question in 1964 and Immerman answered it yes in 1988, in "Nondeterministic Space is Closed Under Complementation".

6. Name two languages a linear bounded automaton accepts that a pushdown automaton does not, and say what they have in common. a to the n b to the n c to the n, and ww. Both need the input revisited, by marking and returning or by crossing back and forth, which a one ended destructive stack cannot do and a bounded tape can.

Contents This chapter on its own page

munotes.in311

Chapter Sixty-Two

The Turing Machine

Syllabus topic Module 2, "Turing Machines: Turing Machine Definition"

In one line

A Turing machine is a finite control with a head that reads and writes on an endless tape and moves one cell at a time.

In the wording a student can write in an examination: a Turing machine is a seven tuple M = (Q, Sigma, Gamma, delta, q0, B, F) where Q is a finite set of states, Sigma the input alphabet, Gamma the tape alphabet containing Sigma and the blank symbol B, delta a partial function from Q times Gamma to Q times Gamma times {L, R}, q0 the initial state and F the set of final states. A move reads the symbol under the head, writes a symbol in its place, moves the head one cell left or right, and changes state.

What Turing was actually doing

Not designing a computer. The machine was a tool invented for a proof, and the proof was about mathematics.

David Hilbert had asked whether there is a definite method that, given any statement of first order logic, decides whether it is provable. The German name for the question is the Entscheidungsproblem, the decision problem, and it is the word in Turing's title. To answer it, Turing had first to say what a definite method IS, because until that is pinned down "no method exists" cannot be proved about anything.

His answer was the machine. And having defined it he showed there are things it cannot do, and hence that the answer to Hilbert's question is no.

So the machine arrived as a definition of "method", and only later as a model of a computer. Chapter 72 is about why that definition is believed to be the right one.

The definition

M = (Q, Sigma, Gamma, delta, q0, B, F)

PartWhat it isNote
Qa finite set of statesas before
Sigmathe input alphabetas before
Gammathe TAPE alphabetcontains Sigma, and the blank
deltathe transition functionnew shape, below
q0the initial stateas before
Bthe blank symbolin Gamma, not in Sigma
Fthe set of final statesas before

The transition function is where everything changed.

delta(q, X) = (p, Y, D)

Read it: in state q, with X under the head, go to state p, write Y in that cell, and move the head one cell in direction D, which is L or R. Some treatments allow a third direction S for staying put, and it adds no power because staying can be simulated by moving right then left.

Three things a move does at once, and each is new or changed.

It writes. The cell's contents are replaced. This is the first machine in the book that can change what it reads, and it is where the extra power comes from.

munotes.in312

The Turing Machine

It moves either way. A finite automaton and a pushdown automaton read left to right and never go back. This head can return.

It may be undefined. delta is a partial function. Where it is undefined the machine halts, and chapter 64 says what that means for acceptance.

The tape

Endless in both directions, or endless to the right with a fixed left end; the two versions accept the same languages and chapter 68 shows why. This book uses the one sided tape, which is what most textbooks draw and what chapter 60's linear bounded automaton is a restriction of.

At the start the tape holds the input, one symbol per cell, and blanks everywhere else. The head is on the leftmost input symbol and the machine is in q0.

The blank is not a symbol of the input alphabet, which is why B is listed separately. If it were, the machine could not tell where the input ended.

How much stronger is it

Can go back over the inputCan writeMemory
finite automatonnonothe state
pushdown automatonnothe stack, at one end, destructivelythe state and the stack
linear bounded automatonyesyes, within the input's spacethe state and n cells
Turing machineyesyes, anywherethe state and unbounded cells

Reading down, the difference at each step is one restriction lifted. And the last row is the end of the road: no model anyone has proposed reaches past it, which is the content of chapter 72.

The first machine, built and run

The task. Accept the strings 0 to the n followed by 1 to the n, for n at least 0. Chapter 52 built a pushdown automaton for this; here is the Turing machine, which works differently and shows what the tape buys.

The method. Repeatedly cross off the leftmost 0 and the leftmost 1, until neither remains.

  1. In q0, if the cell holds a 0, write X over it and move right, into q1.
  2. In q1, scan right past any 0 and any Y, until a 1 is found; write Y over it and move left, into q2.
  3. In q2, scan left past any 0 and any Y, until the X is found; move right, into q0. The head is now on the

leftmost 0 that is still uncrossed.

  1. In q0, if the cell holds a Y rather than a 0, every 0 has been crossed off, so move into q3 and check that only

Y remain before the blank.

  1. In q0, if the cell holds a blank at once, the input was empty, and empty is accepted.
munotes.in313

The Turing Machine

start q0
blank B
accept qa
q0, 0 -> q1, X, R
q0, Y -> q3, Y, R
q0, B -> qa, B, S
q1, 0 -> q1, 0, R
q1, Y -> q1, Y, R
q1, 1 -> q2, Y, L
q2, 0 -> q2, 0, L
q2, Y -> q2, Y, L
q2, X -> q0, X, R
q3, Y -> q3, Y, R
q3, B -> qa, B, S

Accepts: ε, 01, 0011, 000111

Rejects: 0, 1, 10, 001, 011, 0101, 0110

Twelve moves, five states plus the accepting one, and a tape alphabet of five symbols: 0, 1, X, Y and the blank.

The run on 0011, with the state written immediately before the cell the head is on:

q0 0011
X q1 011
X0 q1 11
X q2 0Y1
q2 X0Y1
X q0 0Y1
XX q1 Y1
XXY q1 1
XX q2 YY
X q2 XYY
XX q0 YY
XXY q3 Y
XXYY q3 B
XXYY qa B

Fourteen configurations, thirteen moves, for an input of four symbols. That ratio is worth noticing: a Turing machine crossing back and forth does far more work than the pushdown automaton of chapter 52, which read each symbol once. The tape buys power and costs time, which is the whole subject of chapters 78 onward.

The checker re-executes every one of those thirteen moves against the twelve rules above.

Read four of the configurations against the method. The second, X q1 011, is after step 1: the first 0 has become X and the head has moved right. The fourth, X q2 0Y1, is after the first 1 became Y and the head turned round. The sixth, X q0 0Y1, is after scanning back to the X and stepping right, so the head is on the leftmost surviving 0. And the last, XXYY qa B, is the accepting halt with every symbol crossed off.

Halting

A Turing machine halts when it reaches a state and symbol for which delta is undefined, or when it reaches a halting state that the definition names.

This book writes machines with a named accepting state and, where useful, a named rejecting one, and treats an undefined move as a rejection. That is one of several conventions in the literature and it is the one chapter 64 sets out properly.

A machine need not halt. It may move back and forth for ever, and nothing in the definition forbids it. That single possibility is what separates this model from every earlier one in the book, and chapters 70 and 74 are about it.

munotes.in314

The Turing Machine

Distinctions

A finite automaton's moveA Turing machine's move
Readsthe next input symbolthe cell under the head
Writesnothinga symbol into that cell
Movesforward, always oneleft or right, one cell
May be undefinedmeans rejectmeans halt
Number of moves on an input of length nexactly nany number, or none ever
SigmaGamma
Holdsthe input symbolsthose, the blank, and any working symbols
Appears on the tape at the startyesthe rest is blank
May the blank be in itnoyes, and that is the point

What it does NOT mean

It is not a computer. It is a definition of what a method is, and Turing wrote it to answer a question about logic.

The tape is not infinite in the sense of being all there at once. It is unbounded: however far the head goes, there is more, and at any moment only finitely many cells are non blank.

The head does not skip. One cell per move, and a machine that wants to reach cell 100 takes at least 100 moves.

An undefined move is not an error. It is how the machine halts, and chapter 64 makes that precise.

Halting is not guaranteed. A machine may run for ever, and no earlier machine in this book could.

Writing is not optional power. Without it the machine is a two way finite automaton, which accepts only the regular languages, so the writing is where the strength is.

Quick revision

  • Seven tuple: (Q, Sigma, Gamma, delta, q0, B, F), with Gamma the tape alphabet containing Sigma and the blank.
  • delta(q, X) is (p, Y, D): change state, write Y over X, move one cell left or right. It is PARTIAL, and where it

is undefined the machine halts.

  • The tape starts with the input and blanks elsewhere, head on the leftmost input symbol.
  • Three things are new: it writes, it moves both ways, and it may fail to halt.
  • The standard first machine crosses off matching symbols from the two ends, which needs the head to return, which

is exactly what earlier machines could not do.

  • A configuration is written with the state immediately before the cell the head is on.
  • Turing's paper was written to answer Hilbert's Entscheidungsproblem, and the machine was the tool, not the goal.

Test yourself

1. Give the seven parts and say what is new in delta. Q, Sigma, Gamma, delta, q0, B and F. delta maps a state and a TAPE symbol to a state, a symbol to write and a direction, so a move writes as well as reads and may go either way; and it is partial, so it may be undefined, which is how the machine halts.

munotes.in315

The Turing Machine

2. Why is the blank not in the input alphabet? Because the machine has to be able to tell where the input ends, and if a blank could appear inside the input there would be no way to distinguish the end of the input from a blank within it.

3. Describe the method the 0 to the n 1 to the n machine uses. Cross off the leftmost 0 with an X, scan right to the leftmost 1 and cross it off with a Y, scan back left to the X, step right, and repeat. When no 0 remains, check that only Y lie before the blank, and accept.

4. How many moves does that machine take on 0011, and how does that compare with a pushdown automaton? Thirteen, against four for a pushdown automaton reading each symbol once. The tape buys power and costs time, which is what the complexity chapters measure.

5. What happens when delta is undefined for the current state and symbol? The machine halts. In this book's convention an undefined move is a rejection, and chapter 64 sets out the alternatives.

6. Name the one behaviour this machine has that no earlier machine in the book has. It may fail to halt. Every finite and pushdown automaton finishes when the input runs out; a Turing machine may move for ever, and chapters 70 and 74 are about the consequences.

Contents This chapter on its own page

munotes.in316

Chapter Sixty-Three

Representations of a Turing Machine

Syllabus topic Module 2, "Turing Machines: Representations"

In one line

A Turing machine can be written as a set of quintuples, as a transition table, or as a diagram, and a run can be written as a sequence of instantaneous descriptions.

In the wording a student can write in an examination: a Turing machine may be represented by its transition function written as a set of quintuples (q, X, p, Y, D), by a transition table whose rows are states and whose columns are tape symbols with each entry giving the next state, the symbol written and the direction, or by a transition diagram whose vertices are states and whose edges are labelled X slash Y comma D. The progress of the machine on an input is represented by a sequence of instantaneous descriptions.

The four representations

RepresentationWhat it isBest for
quintuplesa list of five part tuplesproofs, and stating the definition
transition tablea grid, states by tape symbolsdesigning, checking, and implementing
transition diagramstates as circles, moves as labelled arrowsseeing a loop
instantaneous descriptionsa sequence of snapshots of one runshowing what happens on one input

A student must be able to convert between any two, and MU's question about the table is the commonest.

The quintuple

(q, X, p, Y, D)

means: in state q, reading X, go to state p, write Y, and move in direction D. It is delta written as a tuple rather than as an equation, and it is the same information:

delta(q, X) = (p, Y, D)

Some books write quadruples instead, separating writing from moving into two kinds of instruction: (q, X, Y, p) writes, and (q, X, D, p) moves. That model is equivalent and takes twice as many instructions, and MU's syllabus uses the quintuple, which is what this book uses.

The count. A machine with q states and g tape symbols has at most q times g quintuples, one per cell of the table, and usually fewer because delta is partial.

The transition table

The rows are states, the columns are tape symbols, and each entry holds the three things a move does.

The convention this book uses, and it is the one MU's question expects:

  • the top left cell names the machine;
  • the second column is the state;
  • the first column carries the markers: start for the initial state, accept for the accepting state, reject

for a rejecting one if the machine has one;

  • an empty divider column follows the state, which is what draws the double rule;
  • the remaining columns are the tape symbols, the blank among them;
  • a cell holds state, write, direction, and a dash means no move, which is a halt.

Here is the machine of chapter 62 written as a table.

munotes.in317

Representations of a Turing Machine

CB1State01XYB
startq0q1,X,R--q3,Y,Rqa,B,S
q1q1,0,Rq2,Y,L-q1,Y,R-
q2q2,0,L-q0,X,Rq2,Y,L-
q3---q3,Y,Rqa,B,S
acceptqa-----

Accepts: ε, 01, 0011, 000111

Rejects: 0, 1, 10, 001, 011, 0101, 0110

Twenty five cells, twelve of them filled, and the twelve are the machine's twelve quintuples. The checker reads this table as the machine and runs it, so a wrong cell here stops the book being built.

Reading a row. The row for q1 says: on a 0, stay in q1, write the 0 back, move right; on a 1, go to q2, write Y, move left; on a Y, stay in q1, write Y back, move right; on X or a blank, halt.

The dashes are the design. The dash for q1 on a blank is the machine noticing that it scanned to the end without finding a 1, which means there were more 0 than 1, and halting in a non accepting state is the rejection.

Converting the table to quintuples is reading it cell by cell: (q0, 0, q1, X, R), (q0, Y, q3, Y, R) and so on, twelve of them. Converting quintuples to a table is filling a grid. Neither needs thought, and both are asked.

The transition diagram

Circles for states, an arrow from nowhere into the start state, a double circle for an accepting state, and one arrow per quintuple labelled

X / Y, D

meaning: reading X, write Y, move D. Where several quintuples join the same pair of states, one arrow carries all their labels separated by semicolons.

The diagram for the machine above has five circles and twelve arrows, three of them loops: q1 on 0 and on Y, and q2 on 0 and on Y. Those loops are the scanning, and they are what a diagram shows better than a table: the machine sweeping right in q1 and left in q2 is two loops with an arrow between them, which is visible at a glance and is not visible in the grid.

That is the general division of labour. The table cannot hide a missing cell and the diagram cannot hide a cycle, which is chapter 10's point about finite automata, unchanged.

Instantaneous descriptions

A configuration of a Turing machine needs the tape, the head position and the state, and the standard notation folds all three into one string: write the tape, and insert the state immediately before the cell the head is on.

X X q0 Y Y

munotes.in318

Representations of a Turing Machine

means the tape holds X X Y Y, the machine is in q0, and the head is on the first Y. Blanks beyond the written part are not shown.

Why the state goes inside the tape. Because it fixes the head position without a separate number, so a configuration is one string, which is what chapter 73's universal machine needs when it encodes a configuration on a tape of its own.

A move is written with a turnstile, as for a pushdown automaton, and this book writes it |-.

The whole run of chapter 62's machine on 0011:

q0 0011
X q1 011
X0 q1 11
X q2 0Y1
q2 X0Y1
X q0 0Y1
XX q1 Y1
XXY q1 1
XX q2 YY
X q2 XYY
XX q0 YY
XXY q3 Y
XXYY q3 B
XXYY qa B

Fourteen descriptions, thirteen moves, and the checker re-executes every one against the table above.

Two things to notice in the run. At the fifth description the machine is in q2 at the very left end, and at the sixth it has stepped right into q0: that is the return to the leftmost uncrossed symbol. And at the last description the head is on a blank beyond the written tape, which is normal and is how the machine knows the input is finished.

Converting between the four

FromToHow
quintuplestableone per cell; the empty cells are the undefined moves
tablequintuplesread each filled cell as (row, column, contents)
tablediagramone arrow per filled cell, labelled column slash write, direction
diagramtableone filled cell per arrow
anya runstart with the state before the first input symbol and apply moves

The count is the check. After converting to a table, count the filled cells and compare with the number of quintuples. They must agree, and a mismatch means a move was dropped, which is the commonest error in this exercise.

Distinctions

Transition tableTransition diagram
Shows a missing moveas an empty cell, visiblynot at all
Shows a scanning loopas two cells that name their own rowat a glance
Good fordesigning, checking, implementingexplaining, spotting a cycle
Sizestates times symbols cellsone arrow per move
A finite automaton's tableA Turing machine's table
Columnsinput symbolsTAPE symbols, including the blank
A cell holdsthe next statethe next state, the symbol written, the direction
An empty cellnot allowed in a DFAallowed, and means halt

What it does NOT mean

The four representations are not four machines. They are one machine written four ways, and a question may ask for any of them.

A quintuple is not a quadruple. Both models exist and are equivalent; this syllabus uses the quintuple, which does the writing and the moving in one instruction.

munotes.in319

Representations of a Turing Machine

The columns are tape symbols, not input symbols. The blank has a column, and so does every working symbol the machine invents.

An empty cell is not a mistake. It is a halt, and the pattern of empty cells is where the machine's rejections live.

The state in an instantaneous description is not part of the tape. It is written inside the tape string as a notation for where the head is, and it is not in the tape alphabet.

Quick revision

  • Four representations: quintuples, transition table, transition diagram, and a sequence of instantaneous

descriptions for one run.

  • A quintuple (q, X, p, Y, D) is delta(q, X) equal to (p, Y, D).
  • The table has one row per state and one column per TAPE symbol including the blank; a cell holds the next state,

the symbol written and the direction; an empty cell is a halt.

  • The diagram labels each arrow X slash Y comma D.
  • An instantaneous description is the tape with the state written immediately before the cell the head is on, so

the whole configuration is one string.

  • Converting is mechanical, and the check is that the number of filled cells equals the number of quintuples.
  • The table cannot hide a missing move and the diagram cannot hide a loop.

Test yourself

1. Write the quintuple for: in q1, reading a 1, write Y, move left, go to q2. (q1, 1, q2, Y, L), which is delta(q1, 1) equal to (q2, Y, L).

2. What do the columns of a Turing machine's transition table hold, and what does an empty cell mean? The tape symbols, including the blank and every working symbol. An empty cell means delta is undefined there, so the machine halts.

3. How is an instantaneous description written, and why is the state put inside the tape? The tape contents with the state written immediately before the cell the head is on. Putting it inside fixes the head position without a separate number, so the whole configuration is a single string, which is what the universal machine of chapter 73 needs.

4. A machine has 4 states and a tape alphabet of 3 symbols. How many cells does its table have, and how many quintuples can it have? Twelve cells, and at most twelve quintuples, usually fewer because delta is partial.

5. Which representation shows a scanning loop best, and which shows a missing move best? The diagram shows the loop, as an arrow from a state to itself. The table shows the missing move, as an empty cell; a diagram simply lacks an arrow and looks finished.

munotes.in320

Representations of a Turing Machine

6. Convert this row to quintuples: state q2, on 0 write 0 move left stay in q2, on X write X move right go to q0. (q2, 0, q2, 0, L) and (q2, X, q0, X, R).

Contents This chapter on its own page

munotes.in321

Chapter Sixty-Four

Acceptability by a Turing Machine: Accept, Reject and Loop

Syllabus topic Module 2, "Turing Machines: Acceptability by Turing Machines"

In one line

A Turing machine has three possible outcomes on an input, not two: accept, reject, and run for ever.

In the wording a student can write in an examination: a Turing machine M accepts a string w if, started in q0 with w on the tape, it reaches a configuration whose state is in F. The language accepted by M, written T(M) or L(M), is the set of all such w. On a string not in the language M may halt in a non accepting state, which is a rejection, or it may never halt, and the definition of acceptance does not distinguish the two.

The third outcome

Read the definition again and notice what it does not say. It says when a string is accepted. It says nothing about what must happen otherwise.

So on an input not in the language there are two possibilities, and the definition permits both.

OutcomeWhat the machine doesCounted as
acceptreaches an accepting statein the language
rejecthalts, with no move available, in a non accepting statenot in the language
loopnever haltsnot in the language

The third row is new to this book. A finite automaton makes exactly one move per input symbol and then stops. A pushdown automaton may take epsilon moves for ever, but the standard definitions arrange that it cannot usefully do so. A Turing machine has a head that can move back and forth over a tape it is rewriting, and nothing at all compels it to finish.

Why this cannot be defined away

The obvious reaction is to outlaw it: require every machine to halt on every input. That would make the third row disappear and the subject much simpler.

It cannot be done, and the reason is chapter 74. There is no way to tell which machines halt on every input, so "the machines that always halt" is not a set anybody can recognise, and a definition that quantified over it would not be usable.

What can be done, and is, is to give the class of always halting machines a name and study it separately. That is chapter 70's recursive languages, and the gap between them and the rest is the subject of the whole block.

Halting: two conventions

The literature has two ways of saying when a Turing machine stops, and both are in use. A student should recognise both and use one.

Convention one: halt on an undefined move. delta is partial, and the machine halts when it reaches a state and symbol for which delta gives nothing. It accepts if that state is in F, and rejects otherwise. This is chapter 62's definition and the one MU's syllabus follows.

munotes.in322

Acceptability by a Turing Machine: Accept, Reject and Loop

Convention two: two named halting states. delta is total, and the machine has a distinguished accepting state and a distinguished rejecting state, both of which have no outgoing moves. It halts exactly on reaching one of them.

The two are equivalent: add a rejecting state and send every undefined move to it, and convention one becomes convention two; delete the rejecting state and its incoming moves, and the reverse.

This book uses a mixture, which is worth stating plainly: the machines are written with a named accepting state and with undefined moves as rejections, because that is what makes the transition tables short.

The definition of the language

T(M) = { w : q0 w reaches a configuration with a state in F }

Two remarks that earn marks.

Acceptance does not require halting at the accepting state. In most treatments an accepting state has no outgoing moves so the machine halts there anyway, but the definition asks only that such a state be reached.

Rejection and looping are not distinguished. Both mean "not accepted", and the language is the same either way. This is exactly the asymmetry chapter 27 described: a machine that has not answered yet has not said no.

Worked: the three outcomes on one machine

start q0
blank B
accept qa
q0, 0 -> q1, X, R
q0, Y -> q3, Y, R
q0, B -> qa, B, S
q1, 0 -> q1, 0, R
q1, Y -> q1, Y, R
q1, 1 -> q2, Y, L
q2, 0 -> q2, 0, L
q2, Y -> q2, Y, L
q2, X -> q0, X, R
q3, Y -> q3, Y, R
q3, B -> qa, B, S

Accepts: ε, 01, 0011, 000111

Rejects: 0, 1, 10, 001, 011, 0101

Accepting. On 01 the machine crosses off the 0 and the 1 and reaches qa. That is the first row of the table.

Rejecting by an undefined move. On 001 the machine crosses off one 0 and one 1, returns, crosses off the second 0, scans right for another 1 and finds the blank. delta(q1, B) is undefined, so the machine halts in q1, which is not accepting. That is the second row, and the halt is the machine noticing there were more 0 than 1.

Rejecting by a wrong symbol. On 10 the machine is in q0 on a 1, and delta(q0, 1) is undefined, so it halts at once.

This machine never loops, and that is a property of this machine and not of the model. Every one of its moves either crosses off a symbol or advances towards one end, so no run can go on for ever. Machines that do loop are easy to write, and the next section writes one.

munotes.in323

Acceptability by a Turing Machine: Accept, Reject and Loop

A machine that loops

The shortest interesting one.

start p0
blank B
accept pa
p0, 0 -> p0, 0, R
p0, B -> p0, B, R
p0, 1 -> pa, 1, S

Accepts: 1, 01, 001, 101

Rejects: 0

That machine scans right looking for a 1. If it finds one it accepts. If the input has no 1 at all, it walks right over blanks for ever and never halts, which is the third outcome.

So 0 is rejected by looping, and the rejection list above is a claim about the language and not about the machine halting. The checker runs each rejected string for a fixed number of steps and reports it as not accepted, which is the honest reading of the definition: not accepted covers both halting without acceptance and not halting.

And that is the whole difficulty of the subject in one machine. Watching it run tells you nothing: after a million steps it may be about to find a 1, or it may be walking over blanks for ever. Only an argument about the machine settles it, and chapter 74 proves no general argument exists.

The three classes of language this creates

Chapter 27 named them without a machine. Here they are with one.

ClassThe machineOn a string in the languageOn a string not in it
recursive, or decidablealways haltsacceptsrejects, halting
recursively enumerablemay loopacceptsrejects or loops
neitherno machine at all

Every recursive language is recursively enumerable, because a machine that always halts satisfies the weaker condition. The converse fails, and chapter 74 provides the witness.

And the third row is not empty, by chapter 7's counting argument, which showed there are uncountably many languages and countably many machines.

Distinctions

RejectingLooping
The machinehalts in a non accepting statenever halts
The string isnot in the languagenot in the language
Distinguished by the definitionnono
Observable in finite timeyesNO
A finite automatonA Turing machine
Number of moves on input of length nexactly nany number, or unbounded
Always answersyesnot necessarily
Outcomestwothree

What it does NOT mean

Looping does not mean the machine is wrong. It is permitted by the definition, and chapter 70's recursively enumerable languages are defined using machines that do it.

Not accepted does not mean rejected. It covers both halting without acceptance and never halting, and the distinction is exactly what chapter 70 makes.

Running a machine for a long time proves nothing. A machine that has not stopped may stop on the next move. This is the practical face of chapter 74's theorem.

munotes.in324

Acceptability by a Turing Machine: Accept, Reject and Loop

An undefined move is not a crash. It is a halt, and under this book's convention it is a rejection.

Acceptance does not require the input to be consumed. There is no input pointer; the machine may accept with half the tape unvisited, and chapter 66's machines sometimes do.

Quick revision

  • Three outcomes: accept, reject by halting in a non accepting state, and loop by never halting.
  • T(M) is the set of strings on which an accepting state is reached. The definition says nothing about the other

strings, so rejection and looping are not distinguished.

  • Two halting conventions: an undefined move halts, or two named halting states with a total delta. They are

equivalent, and this book uses the first.

  • A machine that always halts gives a recursive language; one that may loop gives a recursively enumerable one;

and by chapter 7 there are languages with no machine at all.

  • Looping cannot be outlawed, because chapter 74 proves there is no way to recognise which machines always halt.
  • Watching a machine run can never establish that it will not stop.

Test yourself

1. Name the three outcomes and say which two the definition treats alike. Accept, reject by halting without acceptance, and loop by never halting. The definition treats rejecting and looping alike: both mean the string is not accepted.

2. Give the two halting conventions and say how each becomes the other. Either delta is partial and an undefined move halts, or delta is total with named accepting and rejecting states that have no outgoing moves. Add a rejecting state and route every undefined move to it to get from the first to the second; delete it to go back.

3. Write a Turing machine that loops on some input. One that scans right looking for a 1, moving right on a 0 and on a blank and accepting on a 1. On an input with no 1 it walks over blanks for ever.

4. A machine has run for a million steps without halting. What can you conclude? Nothing. It may halt on the next move or never. Chapter 74 proves there is no general method that settles the question.

5. Why can looping not simply be forbidden? Because there is no way to recognise which machines always halt, so "the always halting machines" is not a class anybody can identify, and a definition restricted to them would be unusable.

6. Which class of languages does an always halting machine give, and how does it relate to the other? The recursive, or decidable, languages. Every recursive language is recursively enumerable, since an always halting machine satisfies the weaker condition, and the converse fails.

Contents This chapter on its own page

munotes.in325

Chapter Sixty-Five

Designing and Describing a Turing Machine

Syllabus topic Module 2, "Turing Machines: Designing and Description of Turing Machines"

In one line

Design a Turing machine by deciding what it does to the tape and in what order, and describe it at the level of detail the question wants: the table, the method, or the idea. Designing a Turing machine is a matter of choosing the phases first and the states afterwards, which is the opposite of the order most students try.

In the wording a student can write in an examination: the design of a Turing machine proceeds by choosing a tape alphabet including the working symbols the machine will write, dividing the computation into phases each of which is a single sweep of the head, assigning a state to each phase and to each item of finitely many information the machine must carry between cells, and writing a transition for every state and tape symbol that can arise. A machine may be described formally by its transition function, by an implementation description stating the moves of the head and the symbols written, or by a high level description stating the algorithm.

The three levels of description

This is the part of the chapter that decides marks, because a question asking for one level and answered at another has answered a different question.

LevelWhat it givesLengthWhen MU asks for it
formalthe transition table or the quintuplesa table"construct a Turing machine for ..."
implementationthe sweeps, the symbols written, the head's patha numbered list"design and describe ..."
high levelthe algorithm, in the language of the taska paragraph"write a note on ..."

The formal description is the machine. Nothing is left to the reader and the machine can be executed from it.

The implementation description says what the head does without naming the states. "Scan right to the first 1, write Y over it, and scan back left to the leftmost X" is an implementation description of three states' worth of machine, and it is what an examiner usually wants when the question says design.

The high level description says what is computed without mentioning the tape. "Repeatedly cross off one symbol from each end and check that they match" is a high level description of the palindrome machine.

All three are legitimate, and a good answer often gives the high level first, the implementation next and the table last, because the table is unreadable without them.

The method for designing one

Step 1. Decide the tape alphabet. What working symbols does the machine need to write? A machine that crosses things off needs one crossed off symbol per symbol it crosses, or one that serves for several, and the choice affects everything after.

Step 2. Write the high level description first. In words, in the language of the task. If you cannot, the design is not ready.

munotes.in326

Designing and Describing a Turing Machine

Step 3. Break it into sweeps. A sweep is one pass of the head in one direction, doing one job. Almost every machine in this block is a small number of sweeps repeated.

Step 4. Give each sweep a state, and one more state for each thing that must be carried. A machine that erases a symbol at the left end and needs to know which symbol it was when it reaches the right end must carry that in its state, and that is one state per symbol of the alphabet.

Step 5. Fill in the table, and decide every dash on purpose. An empty cell is a rejection, and the pattern of empty cells is where the machine's correctness lives.

Step 6. Run it, on the shortest accepted string, the shortest rejected one, and the empty string.

The two costs to keep in mind

A state per thing carried. The machine has no variables. Anything it must remember while the head travels is recorded in the state, and there are finitely many states, so only finitely many things can be carried. A machine that needs to carry a count cannot carry it in its state and must write it on the tape.

A sweep per pass. Each sweep is one state and costs one pass over the tape. A machine that makes n sweeps over n cells does n squared moves, which is why the machine of chapter 62 is so much slower than the pushdown automaton for the same language.

Worked: one machine at all three levels

The task. Accept the palindromes over {a, b}.

High level description

Repeatedly remove the first symbol, check that the last symbol is the same, and remove it too. If at some point one symbol remains, or none, accept. If the two ends ever differ, reject.

Three sentences, no tape, no states.

Implementation description

  1. If the tape is blank, accept.
  2. Read the leftmost symbol, erase it, and remember which it was.
  3. Scan right to the first blank, then step back one cell, which is the rightmost remaining symbol.
  4. If that cell is blank, the string had one symbol, which has just been erased; accept.
  5. If it holds the remembered symbol, erase it and scan back left to the blank at the left end, step right, and

repeat from step 1.

  1. Otherwise reject.

Six numbered steps, the head's path is explicit, and no state is named. This is the level MU's "design and describe" question wants.

Note step 2's "remember which it was". That is the thing carried in the state, and because the alphabet has two symbols it costs two families of states, one for having erased an a and one for a b. Step 4 also needs its own state per symbol, which is why the table below has six states rather than three.

munotes.in327

Designing and Describing a Turing Machine

Formal description

start q0
blank B
accept qa
q0, B -> qa, B, S
q0, a -> r1, B, R
q0, b -> s1, B, R
r1, a -> r1, a, R
r1, b -> r1, b, R
r1, B -> r2, B, L
r2, a -> r3, B, L
r2, B -> qa, B, S
s1, a -> s1, a, R
s1, b -> s1, b, R
s1, B -> s2, B, L
s2, b -> s3, B, L
s2, B -> qa, B, S
r3, a -> r3, a, L
r3, b -> r3, b, L
r3, B -> q0, B, R
s3, a -> s3, a, L
s3, b -> s3, b, L
s3, B -> q0, B, R

Accepts: ε, a, b, aa, bb, aba, abba, ababa

Rejects: ab, ba, aab, abb, aabb, abab

Twenty moves and seven states. Read them against the implementation description.

StateIts jobWhich step
q0read and erase the leftmost symbol1 and 2
r1having erased an a, scan right to the blank3
r2check the rightmost symbol is an a4 and 5
r3having matched, scan back left5
s1, s2, s3the same three, for a b3, 4, 5
qaaccept

The two states doing the carrying are r1 and s1, and their existence is step 2's "remember which it was" made concrete. Doubling them is the price of a two symbol alphabet, and a three symbol alphabet would triple them.

The rejections are the missing cells. r2 on a b has no move, which is the case where the right end does not match a remembered a. That single empty cell is the whole of step 6.

A run on abba. The erased cells are shown as blanks, which is what the tape really holds, so the written portion grows blanks at the left end as the machine works inward:

q0 abba
B r1 bba
Bb r1 ba
Bbb r1 a
Bbba r1 B
Bbb r2 aB
Bb r3 bBB
B r3 bbBB
r3 BbbBB
B q0 bbBB
BB s1 bBB
BBb s1 BB
BB s2 bBB
B s3 BBBB
BB q0 BBB
BB qa BBB

Sixteen configurations, fifteen moves, on an input of four symbols. The checker re-executes every one against the twenty rules above.

Read the fifth and sixth. At Bbba r1 B the head has reached the blank past the right end, and the next move steps back to Bbb r2 aB, where the head is on the rightmost surviving symbol, an a, matching the a that was erased at the start. It is then erased and r3 scans back.

munotes.in328

Designing and Describing a Turing Machine

Read the ninth and tenth. At r3 BbbBB the head is on the blank left by the first erasure, and stepping right gives B q0 bbBB, with the machine back in q0 on the leftmost surviving symbol. That is the loop closing, and the two blanks at the ends are the two symbols already matched.

Read the last two. At BB q0 BBB every symbol has been erased, so q0 sees a blank and accepts. The tape is empty, which is what the high level description's "or none" case meant.

The trap: a machine that needs to count

Worth a section because it is the commonest design error.

Suppose a student designs a machine for a to the n b to the n by "count the a, then count the b, then compare". The counting has nowhere to go. The state set is finite and fixed before the input is seen, so it cannot hold a number that grows with the input.

The remedy is always the same: put what grows on the tape. The machine of chapter 62 does not count at all; it crosses off in pairs, so nothing has to be remembered between sweeps except which sweep it is in.

The general rule. In a state: which phase, and finitely many carried items. On the tape: everything that grows.

That is the same division chapter 55 drew between a pushdown automaton's states and its stack, and it holds for every machine in this book.

Distinctions

FormalImplementationHigh level
Mentions statesyesnono
Mentions the tapeyesyesno
Can be executed fromyeswith workno
Asked for by"construct""design and describe""write a note"
In a stateOn the tape
Which sweep the machine is inyes
Which of finitely many symbols was just erasedyes
How many symbols have been seenyes
Anything that grows with the inputyes

What it does NOT mean

An implementation description is not vague. It is precise about the head's path and the symbols written, and only leaves out the state names.

A high level description is not an excuse. It must say what the algorithm is, not what the language is.

A state cannot hold a count. Finitely many states were fixed before the input arrived, so anything growing goes on the tape.

Doubling the states for two symbols is not waste. It is what carrying one item of information across a sweep costs, and there is no cheaper way.

munotes.in329

Designing and Describing a Turing Machine

Running the machine is not optional. Every machine in this book is executed on its claimed strings, and the design method's last step says so for a reason.

Quick revision

  • Three levels of description: formal, which is the table; implementation, which is the head's path and the

symbols written; and high level, which is the algorithm in the language of the task.

  • Design: choose the tape alphabet, write the high level description, break it into sweeps, give each sweep a

state plus one per item carried, fill the table deciding every dash, then run it.

  • A state holds which phase the machine is in and finitely many carried items. Anything that grows goes on the

tape.

  • Carrying one symbol across a sweep costs one family of states per symbol of the alphabet.
  • The empty cells are the rejections, and each should be deliberate.
  • A machine making n sweeps over n cells does n squared moves, which is why the tape is slow.

Test yourself

1. Name the three levels of description and say what each leaves out. Formal leaves out nothing and is the table. Implementation leaves out the state names but gives the head's path and what is written. High level leaves out the tape entirely and gives the algorithm.

2. Give a high level description of the machine for 0 to the n 1 to the n. Repeatedly cross off the leftmost 0 and the leftmost 1; accept when none of either remains, and reject if one runs out before the other.

3. Why does the palindrome machine need six states rather than three? Because it must carry, across a whole sweep to the right, which symbol it erased at the left end. That is one item of information over a two symbol alphabet, so each of the three sweeps is doubled.

4. A student proposes counting the a, then the b, then comparing. What is wrong? The count has nowhere to live. States are finite and fixed before the input is seen, so a number that grows with the input cannot be held in one. The count must be written on the tape, or the design must avoid counting, as crossing off in pairs does.

5. What does an empty cell in the table mean, and why does it matter to the design? The machine halts there, which under this book's convention is a rejection. The pattern of empty cells is therefore where the machine's rejections are decided, so each one should be deliberate.

6. Roughly how many moves does a machine that makes n sweeps over n cells take? About n squared, which is why a Turing machine solving a problem a pushdown automaton can solve is much slower than the pushdown automaton.

Contents This chapter on its own page

munotes.in330

Chapter Sixty-Six

Turing Machine Construction: Machines That Recognise

Syllabus topic Module 2, "Turing Machines: Turing Machine Construction"

In one line

Three machines that answer yes or no, for three languages no earlier machine in this book can accept.

In the wording a student can write in an examination: the construction of a Turing machine for a language proceeds by choosing working symbols to mark the tape, dividing the computation into sweeps, and writing the transition for every state and tape symbol, with an undefined transition serving as a rejection.

Machine 1: a to the n b to the n c to the n

Why it matters. Chapter 50 proved this language is not context free, so no pushdown automaton accepts it. It is context sensitive, so chapter 61's linear bounded automaton does, and a Turing machine does because the linear bounded automaton is one.

High level description. Repeatedly cross off one a, one b and one c. Accept when none of the three remains and nothing else is left over.

Implementation description.

  1. If the tape is blank, accept: the empty string qualifies with n equal to 0.
  2. Scan right past any crossed off a. On the leftmost uncrossed a, cross it off as X and move right.
  3. Scan right past any a and any crossed off b. On the leftmost uncrossed b, cross it off as Y and move right.
  4. Scan right past any b and any crossed off c. On the leftmost uncrossed c, cross it off as Z and turn left.
  5. Scan left to the leftmost X, step right, and repeat from step 2.
  6. When step 2 finds a Y instead of an a, every a has been crossed off. Check that only Y and then only Z remain

before the blank, and accept.

Formal description.

start p0
blank B
accept qa
p0, B -> qa, B, S
p0, X -> p0, X, R
p0, Y -> p4, Y, R
p0, a -> p1, X, R
p1, a -> p1, a, R
p1, Y -> p1, Y, R
p1, b -> p2, Y, R
p2, b -> p2, b, R
p2, Z -> p2, Z, R
p2, c -> p3, Z, L
p3, a -> p3, a, L
p3, b -> p3, b, L
p3, c -> p3, c, L
p3, Y -> p3, Y, L
p3, Z -> p3, Z, L
p3, X -> p0, X, R
p4, Y -> p4, Y, R
p4, Z -> p5, Z, R
p5, Z -> p5, Z, R
p5, B -> qa, B, S

Accepts: ε, abc, aabbcc, aaabbbccc

Rejects: a, b, c, ab, abcc, aabc, aabbc, cba, acb

Twenty one moves, seven states. Read the state table against the implementation description:

StateIts sweep
p0find the leftmost uncrossed a, or discover there are none
p1scan right for the leftmost uncrossed b
p2scan right for the leftmost uncrossed c
p3scan all the way back left to the X block
p4, p5verify only Y then only Z remain
munotes.in331

Turing Machine Construction: Machines That Recognise

Why aabc is rejected. The first round crosses off one a, one b and one c. The second round reaches p1 looking for a b, scans right past the crossed off Y and finds the blank; delta(p1, B) is undefined, so the machine halts in p1, which is not accepting. That is the machine noticing there were more a than b.

Why cba is rejected. In p0 the machine is on a c and delta(p0, c) is undefined, so it halts at once. The block order is enforced by the missing cells, exactly as chapter 56's pushdown automaton enforced it.

Why abcc is rejected. After the one round, p0 sees a Y, moves to p4, scans Y and Z, and reaches the second c while in p5; delta(p5, c) is undefined and the machine halts.

A run on abc:

p0 abc
X p1 bc
XY p2 c
X p3 YZ
p3 XYZ
X p0 YZ
XY p4 Z
XYZ p5 B
XYZ qa B

Nine configurations, eight moves, and the checker re-executes each.

Machine 2: the palindromes over {a, b}

Why it matters. Chapter 37 proved this is not regular, and chapter 56 built a nondeterministic pushdown automaton for it that chapter 57 showed no deterministic one can match. The Turing machine below is deterministic, which makes a point worth stating: the Turing machine needs no guessing here, because it can go back.

High level description. Repeatedly remove the first symbol, check the last symbol matches, and remove it too.

The machine is chapter 65's, built there at all three levels of description, and it is printed again here so this chapter's three machines stand together.

start q0
blank B
accept qa
q0, B -> qa, B, S
q0, a -> r1, B, R
q0, b -> s1, B, R
r1, a -> r1, a, R
r1, b -> r1, b, R
r1, B -> r2, B, L
r2, a -> r3, B, L
r2, B -> qa, B, S
s1, a -> s1, a, R
s1, b -> s1, b, R
s1, B -> s2, B, L
s2, b -> s3, B, L
s2, B -> qa, B, S
r3, a -> r3, a, L
r3, b -> r3, b, L
r3, B -> q0, B, R
s3, a -> s3, a, L
s3, b -> s3, b, L
s3, B -> q0, B, R

Accepts: ε, a, b, aa, bb, aba, abba, ababa, aabbaa

munotes.in332

Turing Machine Construction: Machines That Recognise

Rejects: ab, ba, aab, abb, aabb, abab, abaa

The comparison worth drawing.

Chapter 56's pushdown automatonThis Turing machine
Deterministicno, and no deterministic one existsyes
Whyit must guess where the middle isit works from both ends and never needs the middle
Moves on an input of length nnabout n squared

So determinism was bought with time. That trade appears again in chapter 69, where a nondeterministic Turing machine is simulated by a deterministic one at an exponential cost.

Machine 3: binary strings whose value is divisible by 3

Why it matters. This language IS regular, and chapter 13 built a three state finite automaton for it. The Turing machine below is longer and slower, and it is here to make a point MU's classification question depends on: a Turing machine accepts every regular language, because a finite automaton is a Turing machine that never writes and never moves left.

High level description. Read left to right, keeping the remainder so far in the state. Accept if the remainder is 0 when the blank is reached.

Implementation description. There is nothing to write and nowhere to go but right. The remainder is one of three values, so three states carry it. That is the whole machine.

start m0
blank B
accept qa
m0, 0 -> m0, 0, R
m0, 1 -> m1, 1, R
m1, 0 -> m2, 0, R
m1, 1 -> m0, 1, R
m2, 0 -> m1, 0, R
m2, 1 -> m2, 1, R
m0, B -> qa, B, S

Accepts: ε, 0, 11, 110, 1001, 1100

Rejects: 1, 10, 100, 101, 111, 1011

Seven moves, four states, and every move writes back what it read and goes right. That is exactly a finite automaton, and the transition table is chapter 13's with a write and a direction added to every cell.

The general statement. For any finite automaton, build a Turing machine with the same states, the same transitions, a move that writes back what it reads and moves right, and a transition from each final state on the blank into an accepting state. So every regular language is accepted by a Turing machine, and the same argument with a stack simulated on the tape covers the context free languages. The Turing machine sits above everything.

What the three machines have in common

Each one marks the tape and revisits it, except the third, which is deliberately the case that does not need to.

Each rejection is a missing cell. There is no rejecting state in any of them; a run dies for want of a move, and the pattern of missing cells is the design.

munotes.in333

Turing Machine Construction: Machines That Recognise

Each is run. The claim lists are executed and every printed trace is re-executed, so none of these machines is offered on the strength of looking plausible.

And each is checked on the empty string, which is the case a designer forgets and which all three accept or reject on purpose.

Distinctions

MachineLanguageOut of reach ofBecause
1a to the n b to the n c to the npushdown automatatwo independent matchings, chapter 50
2palindromesdeterministic pushdown automatathe middle must be guessed, chapter 57
3divisible by 3nothing; it is regularit is here to show the containment
A finite automatonThe same thing as a Turing machine
Writesneverwrites back what it read
Movesforward one per symbolright one per symbol
Acceptsby the state when the input endsby a move on the blank into an accepting state

What it does NOT mean

Marking is not the only technique. It is the one that fits these three. Chapter 67's machines move and copy symbols instead.

A Turing machine being deterministic here is not a general fact. Chapter 69 defines the nondeterministic version and shows it accepts the same languages, at a cost.

A regular language does not need a Turing machine. Machine 3 is a demonstration of a containment, not a recommendation.

Slower is not worse for this subject. These chapters ask what can be computed; chapters 78 onward ask what it costs, and that is where machine 2's n squared becomes a fact worth quoting.

Quick revision

  • a to the n b to the n c to the n: cross off one of each per round, then verify only the crossed off symbols

remain. Seven states, twenty one moves. Not context free, so this is genuinely new ground.

  • Palindromes: erase from both ends, carrying the erased symbol in the state. Deterministic, unlike the pushdown

automaton for the same language, and slower by a factor of n.

  • Divisible by 3: a finite automaton written as a Turing machine, every move writing back and going right. It

shows every regular language is accepted by some Turing machine.

  • Rejections are missing cells in all three; no machine here has a rejecting state.
  • The empty string is decided on purpose in all three.

Test yourself

1. Give the high level description of the machine for a to the n b to the n c to the n. Repeatedly cross off one a, one b and one c, then accept when none of the three remains and nothing uncrossed is left over.

munotes.in334

Turing Machine Construction: Machines That Recognise

2. Why is aabc rejected, and in which state does the machine halt? After one round it looks for a second b, scans right past the crossed off Y and reaches the blank, where the transition is undefined. It halts in the state that was scanning for a b, which is not accepting.

3. Why is the palindrome Turing machine deterministic when the pushdown automaton cannot be? Because it can return along the tape and work from both ends, so it never has to know where the middle is. A pushdown automaton reads once and must guess.

4. How do you turn any finite automaton into a Turing machine? Keep the states and transitions, make every move write back the symbol it read and move right, and add a transition from each final state on the blank into an accepting state.

5. What does machine 3 demonstrate, given that the language is regular? That the containment of the regular languages in the recursive ones is real: every finite automaton is a Turing machine that never writes and never moves left.

6. Where are the rejections in these three machines? In the missing cells of the transition tables. None of the three has a rejecting state, so a run dies for want of a move, and each missing cell is a deliberate part of the design.

Contents This chapter on its own page

munotes.in335

Chapter Sixty-Seven

Turing Machine Construction: Machines That Compute

Syllabus topic Module 2, "Turing Machines: Turing Machine Construction"

In one line

A Turing machine computes a function when the tape it leaves when it halts is the answer.

In the wording a student can write in an examination: a Turing machine M computes a partial function f from Sigma star to Sigma star if, started with w on the tape and the head on its leftmost symbol, M halts with f(w) on the tape whenever f(w) is defined. A function so computed is called Turing computable, and the class of Turing computable functions is the formal counterpart of the informal notion of an effectively calculable function.

Accepting against computing

The same machine, read two ways.

An acceptorA transducer
The question isis this string in the language?what is f of this string?
The answer isthe state it halts inthe TAPE it halts with
WrittenT(M), a languagef, a function
Needs final statesyesnot really; halting is enough

Nothing in the definition changes. A Turing machine writes on its tape whatever it is doing, and calling the final tape an answer rather than ignoring it is a change of reading, not of machine.

Why it matters. Chapter 72's thesis is a claim about functions, not about languages: that everything effectively calculable is Turing computable. So the machines in this chapter are the evidence for it, and the chapter is where a student sees a Turing machine do arithmetic.

How a number is written on a tape

Two conventions, both used in this book and both examined.

Unary. The number n is written as n copies of 1. So 3 is 111 and 0 is the empty string. Arithmetic is easy and the tape is long.

Binary. The number n is written in base 2 with the most significant bit on the left. So 3 is 11. The tape is short and arithmetic takes more states.

Several numbers are separated by a symbol that is not 1, usually a 0 or a special marker. So 3 plus 2 in unary is 111011.

Machine 1: addition in unary

The function. Given 1 to the m, 0, 1 to the n, leave 1 to the (m plus n).

The idea, and it is a nice one. The tape already holds m plus n copies of 1 and one 0 in the middle. Turn the 0 into a 1 and there are m plus n plus 1 copies; then erase one from the end.

Implementation description.

  1. Scan right over the 1 until the 0 is found, and write a 1 over it.
  2. Scan right over the remaining 1 to the blank, then step back one cell.
  3. Erase that cell and halt.

Formal description.

munotes.in336

Turing Machine Construction: Machines That Compute

start a0
blank B
accept qa
a0, 1 -> a0, 1, R
a0, 0 -> a1, 1, R
a1, 1 -> a1, 1, R
a1, B -> a2, B, L
a2, 1 -> qa, B, S

Five moves, four states. The whole computation is three sweeps and no counting, which is the design lesson of chapter 65 applied: the tape holds what grows and the states hold only which sweep is running.

The run on 11011, which is 2 plus 2:

a0 11011
1 a0 1011
11 a0 011
111 a1 11
1111 a1 1
11111 a1 B
1111 a2 1B
1111 qa BB

Eight configurations, seven moves. The final tape is 1111, which is 4, and the trailing blanks are the erased cell and the tape beyond. The checker re-executes every move.

Three more results, each a run of the same machine:

InputMeansFinal tapeValue
110112 plus 211114
11012 plus 11113
101111 plus 311114
010 plus 111

The last row is the boundary case: m equal to 0 is the empty block before the 0, and the machine handles it because a0 meets the 0 immediately.

Machine 2: the successor in binary

The function. Given a binary number, leave the number one larger.

The idea. School arithmetic. Start at the least significant end and carry: every 1 becomes 0 and the carry continues; the first 0 becomes 1 and the carry stops. If every digit was a 1, a new 1 appears on the left.

Implementation description.

  1. Scan right to the blank past the right end, then step back one cell, which is the least significant digit.
  2. While the cell holds a 1, write 0 and move left.
  3. On the first 0, write 1 and halt.
  4. If the left end is passed, write 1 in the cell before the number and halt.

Formal description.

start s0
blank B
accept qa
s0, 0 -> s0, 0, R
s0, 1 -> s0, 1, R
s0, B -> s1, B, L
s1, 1 -> s1, 0, L
s1, 0 -> qa, 1, S
s1, B -> qa, 1, S

Six moves, three states.

Two runs, showing both cases. On 1011, which is 11:

s0 1011
1 s0 011
10 s0 11
101 s0 1
1011 s0 B
101 s1 1B
10 s1 10B
1 s1 000B
1 qa 100B

Nine configurations, eight moves, and the final tape is 1100, which is 12. The carry turned the last two 1 into 0 and then the 0 before them into a 1.

And on 11, which is 3, where every digit is a 1:

munotes.in337

Turing Machine Construction: Machines That Compute

s0 11
1 s0 1
11 s0 B
1 s1 1B
s1 10B
s1 B00B
qa 100B

Seven configurations, six moves, and the final tape is 100, which is 4. The head ran off the left end onto a blank and wrote the new leading 1 there, which is step 4.

That last case is why the machine needs three states and not two: the blank at the left has to be told apart from a 0, and both end the carry, but only one of them adds a digit.

Machine 3: doubling, in unary

The function. Given 1 to the n, leave 1 to the 2n.

Why it is harder. Addition and the successor each touched every cell once or twice. This one has to copy, and copying on a single tape means travelling the whole length once per symbol, which is where the n squared of chapter 65 comes from.

The idea. For each 1 of the original, mark it and append a marked 1 at the far right. When every original is marked, everything on the tape is a mark, and one final sweep turns them all back into 1.

Why the marks are needed. Without them the machine could not tell an original 1 from an appended one, and it would keep doubling for ever. The mark is what separates "already counted" from "not yet", and it is the same device chapter 66's machines used.

Implementation description.

  1. Find the leftmost unmarked 1, mark it as X, and move right.
  2. Scan right past everything to the blank, and write an X there.
  3. Scan back left to the left end, step right, and repeat from step 1.
  4. When no unmarked 1 remains, the head reaches the blank at the far right; scan back left turning every X into a

1, and halt.

Formal description.

start d0
blank B
accept qa
d0, X -> d0, X, R
d0, 1 -> d1, X, R
d0, B -> d3, B, L
d1, 1 -> d1, 1, R
d1, X -> d1, X, R
d1, B -> d2, X, L
d2, 1 -> d2, 1, L
d2, X -> d2, X, L
d2, B -> d0, B, R
d3, X -> d3, 1, L
d3, B -> qa, B, S

Eleven moves, five states.

The run on 1, the smallest case:

d0 1
X d1 B
d2 XX
d2 BXX
B d0 XX
BX d0 X
BXX d0 B
BX d3 XB
B d3 X1B
d3 B11B
qa B11B

Eleven configurations, ten moves, and one input symbol has become two. Read the third: the single 1 has become an X and a second X has been appended, so the tape is XX and the machine is on its way back left. Read the eighth and ninth: d3 is sweeping left turning each X into a 1.

munotes.in338

Turing Machine Construction: Machines That Compute

And on 11, where the round trips become visible, the machine takes twenty four moves. The first round marks the leftmost 1 and appends an X at the right, the second marks the other and appends again, and the final sweep converts all four marks, leaving 1111.

The cost. For an input of n the machine makes n round trips, each over a tape of length up to 2n, so it does about n squared moves. Measured: 10 moves for n equal to 1, 24 for n equal to 2, and 44 for n equal to 3. That growth is the subject of chapter 78, and it is the price of a single tape: chapter 68's two tape machine copies in a single pass.

Distinctions

UnaryBinary
n is written asn copies of 1the base 2 digits
Tape length for nnabout log n
Additionthree sweeps, five movesa carry, more states
Which this book usesfor addition and doublingfor the successor
AcceptingComputing
The answer isthe halting statethe halting TAPE
The machine differsnot at allnot at all
The object defineda languagea function

What it does NOT mean

A computing machine is not a different model. It is the same machine with the tape read as an answer.

Unary is not a toy. It makes the arithmetic short, which is why every textbook uses it here, and it makes the tape long, which is why no real computation does.

Copying is not free. On one tape it costs a pass per symbol, hence n squared, and chapter 68's second tape is what removes that.

A blank is not a zero. In machine 2 the blank at the left end and a 0 digit both stop the carry, and only the blank adds a digit, which is why they need separate moves.

Halting is the answer. There is no output instruction; the tape when the machine stops is the result.

Quick revision

  • A Turing machine computes f when it halts with f(w) on the tape. Same machine, different reading.
  • Unary: n is n ones, and several numbers are separated by a non 1 symbol. Binary: the digits, most significant

first.

  • Unary addition: turn the separating 0 into a 1 and erase one 1 from the right end. Five moves.
  • Binary successor: scan to the right end, then carry leftwards turning 1 into 0 until a 0 or the left blank is

reached, and write a 1 there.

munotes.in339

Turing Machine Construction: Machines That Compute

  • Unary doubling: mark each original and append a mark at the far right, then convert every mark back. It costs

about n squared moves, because copying on one tape needs a pass per symbol.

  • The marks are what separate counted from uncounted; without them the doubling machine would never stop.

Test yourself

1. What does it mean for a machine to compute a function? That started with w on the tape it halts with f(w) on the tape. The answer is what the tape holds when it stops, not the state it stops in.

2. Describe the unary addition machine in one sentence. Turn the 0 that separates the two blocks into a 1, then erase one 1 from the right hand end.

3. In the binary successor machine, why do the blank and the digit 0 need separate moves? Both stop the carry, but a 0 becomes a 1 in place while a blank means every digit was a 1, so a new leading 1 has to be written in a cell the number did not occupy.

4. Why does the doubling machine need to mark the original symbols? Because otherwise it could not tell an original 1 from one it had appended, and it would keep appending for ever.

5. Roughly how many moves does unary doubling take, and what would remove the cost? About n squared, since it makes one round trip per symbol over a tape of length up to 2n. A second tape removes it: chapter 68's two tape machine copies in a single pass.

6. Give the two ways of writing a number on a tape, with 5 in each. Unary, as 11111. Binary, as 101.

Contents This chapter on its own page

munotes.in340

Chapter Sixty-Eight

Variants of the Turing Machine: More Tapes, More Tracks, More Heads

Syllabus topic Module 2, "Turing Machines: Variants of Turing Machine"

In one line

Extra tapes, extra tracks and extra heads make a Turing machine more convenient and not more powerful, and each convenience is paid for in time.

In the wording a student can write in an examination: a multitape Turing machine has k tapes each with its own head, and its transition function reads k symbols and writes k symbols and moves k heads. A multitrack machine has one tape whose cells are divided into k tracks, read and written together by one head. A multihead machine has one tape and k heads. Each of these accepts exactly the languages a single tape single head machine accepts.

Why variants are worth studying at all

Two reasons, and both are examinable.

Because designing is easier with them. Chapter 67's doubling machine took n squared moves because copying on one tape needs a pass per symbol. With two tapes it is one pass. A great many constructions in this subject are stated on a multitape machine because the single tape version would be unreadable.

Because proving them no stronger is what makes the model robust. Chapter 72's Church Turing thesis rests on exactly this kind of result: every reasonable extension anybody has tried turns out to accept the same languages. Each simulation in this chapter is one piece of that evidence.

The multitrack machine

The definition. One tape, one head, but each cell is divided into k tracks. The head reads all k symbols of a cell at once, writes all k at once, and moves as usual.

The key observation, and it is the whole of the equivalence. A cell holding k symbols from an alphabet of size g is a cell holding one symbol from an alphabet of size g to the k. So a multitrack machine is an ordinary machine whose tape alphabet is the set of k tuples.

a cell of 3 tracks over {0, 1, B} is a cell over an alphabet of 27 symbols

Therefore it is not a new model at all, only a way of writing one. There is nothing to simulate: the translation is a renaming of the alphabet.

What it is for. Keeping a second piece of information beside each input symbol, without disturbing the input. Chapter 60's linear bounded automaton used exactly this to get a constant multiple of the input's space, and chapter 61's machine for the perfect squares computes on a second track.

Cost: none in time. One move of the multitrack machine is one move of the ordinary one. The cost is in the alphabet, which grows exponentially in the number of tracks, and an alphabet is finite however large, so nothing is lost.

munotes.in341

Variants of the Turing Machine: More Tapes, More Tracks, More Heads

The multitape machine

The definition. k tapes, each with its own head, each head moving independently. The transition function reads the k symbols under the k heads and writes k symbols and moves each head left, right or not at all.

delta(q, X1, ..., Xk) = (p, Y1, ..., Yk, D1, ..., Dk)

At the start the input is on tape 1 and the others are blank, and all heads are at the left.

What it is for. Anything needing two places at once. Copying, comparing two strings, keeping a counter beside a computation, or simulating a stack. Chapter 67's doubling in one pass: read tape 1 left to right, and write two copies of each symbol on tape 2.

The simulation, and why it costs a square

Theorem. Every k tape machine can be simulated by a single tape machine.

The construction. Use 2k tracks on the single tape, by the multitrack observation above. Tracks 1, 3, 5 and so on hold the contents of the k tapes; tracks 2, 4, 6 and so on hold a single mark each, showing where that tape's head is.

One move of the k tape machine becomes one sweep of the single tape machine:

  1. Sweep right from the left end to the rightmost mark, remembering in the state the k symbols found under the k

marks. That is one item of finitely many, so it fits in the state.

  1. Now the machine knows what the k tape machine would do, so sweep back left, and at each mark write the new

symbol and move the mark one cell left or right.

The cost. Each sweep crosses the whole written portion of the tape. After t moves of the k tape machine no tape is longer than t, so each simulated move costs about t moves, and t moves cost about t squared in total.

a k tape machine running in time t is simulated in time about t squared

That bound is the one to quote, and it is used in chapter 81: it is why the class P is the same whether it is defined on a one tape or a many tape machine, since the square of a polynomial is a polynomial.

And it is tight. There are languages a two tape machine decides in time proportional to n that a one tape machine needs time proportional to n squared for, so the square is not an artefact of a lazy simulation.

The multihead machine

The definition. One tape, k heads, all reading the same tape and moving independently.

The simulation. The same trick as for the multitape machine: one track for the tape and k tracks each holding one head mark. A move is a sweep to collect what the k heads see and a sweep back to act.

munotes.in342

Variants of the Turing Machine: More Tapes, More Tracks, More Heads

The cost. The same square, and for the same reason.

A subtlety worth a sentence. Two heads may be on the same cell and may try to write different symbols. The definition has to say what happens, and the usual rule is that the lowest numbered head wins, or that the machine is required never to do it. Either convention gives the same class of languages.

The two sided tape, and other small variations

Several more variants are standard and each is no stronger, and each has a one line reason.

VariantWhy it is no stronger
a tape infinite in both directionsfold it at the origin onto two tracks of a one sided tape
a head that may stay putreplace a stay by a move right then a move left
a machine with no writingthat is a two way finite automaton, and it is WEAKER: only the regular languages
two stacks instead of a tapeas strong as a tape; the two stacks hold the tape left and right of the head
several accepting stateskeep one, and move to it from all of them

The fourth row is worth noticing. A pushdown automaton with a second stack is as powerful as a Turing machine, because the tape to the left of the head is one stack and the tape to the right is the other, and moving the head is popping from one and pushing onto the other. So the jump in power from chapter 52 to chapter 62 is exactly one extra stack, which is a satisfying way to see it and is the reason chapter 56 said a second stack is not a variant of a pushdown automaton but a different model.

The third row is the only one that goes the other way. Taking writing away leaves a machine that can move both ways over a fixed tape, and such a machine accepts only the regular languages. So writing is where the power is, which chapter 62 claimed and this is the justification.

The pattern of all these proofs

Every simulation in this chapter has the same three parts, and an examination answer that gives them has answered the question.

One. Say how the richer machine's configuration is represented on the simpler one. Tracks, marks, a folded tape.

Two. Say how one move of the richer machine becomes several moves of the simpler one. Usually a sweep to gather and a sweep to act.

Three. Say what it costs, in moves, and note that the cost is polynomial so the class of languages is unchanged and the class P of chapter 81 is unchanged too.

munotes.in343

Variants of the Turing Machine: More Tapes, More Tracks, More Heads

Distinctions

VariantTapesHeadsReally a new model?Time cost
multitrackoneoneno, it is an alphabet renamingnone
multitapekk, one per tapenoabout a square
multiheadoneknoabout a square
two way infinite tapeoneonenonone, after folding
no writingoneoneit is WEAKER: regular only
two stacksnoneno, equal to one tape
More powerMore convenience
Any variant herenoyes
What is gainednothing in what can be computedshorter machines, fewer moves
What is paidnothingthe simulation's time, when reduced to one tape

What it does NOT mean

Equivalent does not mean equally fast. The languages are the same and the times are not, and chapter 78 onwards is entirely about the difference.

A multitrack machine is not a multitape machine. One head against k heads, and the difference is real: a multitrack machine cannot look at two distant places at once.

More tapes do not help with undecidability. Chapter 74's halting problem is unsolvable for a machine with any number of tapes, because they all accept the same languages.

The square is not an artefact. There are languages for which the one tape machine really does need the square, so the simulation is not merely unclever.

Removing writing is not a harmless variation. It drops the machine to the regular languages, which is the one change in this chapter that costs power.

Quick revision

  • Multitrack: one tape whose cells hold k symbols at once. It IS an ordinary machine over an alphabet of k tuples,

so there is nothing to simulate and no time cost.

  • Multitape: k tapes with k independent heads. Simulated on one tape with 2k tracks, k for contents and k for head

marks, one move becoming a sweep to gather and a sweep to act.

  • Cost of that simulation: a k tape machine running in time t becomes a one tape machine running in about t

squared, and the bound is tight.

  • That square is why the class P does not depend on the number of tapes.
  • Multihead: same trick, same square, plus a convention for two heads on one cell.
  • A two way infinite tape folds onto two tracks; a stay move is a right then a left; several accepting states

collapse to one.

  • Two stacks are as strong as a tape, so the step from chapter 52 to chapter 62 is exactly one extra stack.
  • Taking WRITING away leaves only the regular languages, so writing is where the power is.

Test yourself

1. Why is a multitrack machine not really a new model? Because a cell holding k symbols from an alphabet of size g is a cell holding one symbol from an alphabet of size g to the k. It is an ordinary machine with a renamed alphabet, so there is nothing to simulate.

munotes.in344

Variants of the Turing Machine: More Tapes, More Tracks, More Heads

2. Describe the simulation of a k tape machine on one tape. Use 2k tracks: k for the contents of the k tapes and k holding one mark each for the head positions. One move becomes a sweep right gathering the k symbols under the marks into the state, and a sweep back left writing the new symbols and moving the marks.

3. What does that simulation cost, and why does it matter? About the square of the running time, since each simulated move sweeps the whole written tape. It matters because the square of a polynomial is a polynomial, so the class P of chapter 81 does not depend on how many tapes the model has.

4. Which variation in this chapter makes the machine WEAKER, and what does it become? Removing the ability to write. The result is a two way finite automaton, which accepts only the regular languages.

5. How many stacks does it take to match a Turing machine, and why? Two. The tape to the left of the head is one stack and the tape to the right is the other, and moving the head is popping from one and pushing onto the other.

6. Give the three parts every one of these equivalence proofs has. How the richer machine's configuration is represented on the simpler one; how one of its moves becomes several moves of the simpler one; and what that costs in time.

Contents This chapter on its own page

munotes.in345

Chapter Sixty-Nine

Variants of the Turing Machine: Nondeterministic, Offline, and the Enumerator

Syllabus topic Module 2, "Turing Machines: Variants of Turing Machine"

In one line

A nondeterministic Turing machine, an offline one and an enumerator all accept exactly the languages an ordinary Turing machine accepts, and the first of the three costs an exponential to simulate.

In the wording a student can write in an examination: a nondeterministic Turing machine has a transition function mapping a state and a tape symbol to a set of triples, and accepts a string if some sequence of choices leads to an accepting state. Every nondeterministic Turing machine can be simulated by a deterministic one, so the two accept the same class of languages. An offline Turing machine has a separate read only input tape and one or more read write work tapes. An enumerator is a Turing machine with an output device that prints strings, and a language is recursively enumerable exactly when some enumerator prints it.

The nondeterministic Turing machine

The definition. One change to chapter 62's, and it is the change chapters 14 and 52 each made to their model:

delta(q, X) = a SET of triples (p, Y, D)

So at each step the machine may have several moves available, and the definition of acceptance is chapter 14's: the string is accepted if some sequence of choices reaches an accepting state.

The computation is a tree, not a line. The root is the starting configuration and a node's children are the configurations its available moves lead to. The machine accepts when the tree contains an accepting node anywhere.

A branch may be infinite, and the tree may be infinite even when an accepting node exists. That matters for the simulation below.

The simulation, and why depth first is wrong

Theorem. For every nondeterministic Turing machine there is a deterministic one accepting the same language.

The wrong construction. Explore the tree depth first: follow one branch to its end, then backtrack. This is what a programmer would write and it is wrong, because a branch may be infinite. The machine would go down an endless branch and never reach the accepting node sitting on the branch beside it.

The right construction: breadth first. Explore the tree level by level, so that every node at depth d is visited before any node at depth d plus 1.

How, on a three tape machine. Tape 1 holds the input, untouched. Tape 2 holds the working tape of one branch. Tape 3 holds a sequence of choices, a string over the digits 1 to b where b is the largest number of moves any situation offers.

  1. Write the next choice sequence on tape 3, in canonical order: first all sequences of length 0, then length 1,

and so on.

  1. Copy the input onto tape 2.
  2. Simulate the nondeterministic machine on tape 2, at each step taking the move named by the next digit of tape
  3. If a digit names a move that does not exist, abandon this sequence.
  4. If the simulation reaches an accepting state, accept. Otherwise go back to step 1 with the next sequence.
munotes.in346

Variants of the Turing Machine: Nondeterministic, Offline, and the Enumerator

Why it works. Every finite branch of the tree corresponds to some finite choice sequence, and every finite sequence is eventually written on tape 3, so every accepting node is eventually reached. Infinite branches are never followed to the end, because each simulation runs only as far as its choice sequence allows.

The cost. If the nondeterministic machine accepts within t steps, the deterministic one may have to try every choice sequence of length up to t, and there are about b to the t of those.

a nondeterministic machine running in time t becomes a deterministic one running in about b to the t

Exponential, and that is not known to be avoidable. Chapter 86's P against NP question is exactly the question of whether it can be improved to a polynomial for the machines that run in polynomial time, and nobody knows.

So this variant is the one where equivalence and efficiency come apart. The class of languages is the same; the class of languages decidable quickly may not be.

The contrast with the earlier models

ModelDeterministic equals nondeterministic?Cost of removing it
finite automatonyes, chapter 16exponentially many states, once
pushdown automatonNO, chapter 57impossible
linear bounded automatonnot known, chapter 60
Turing machineyes, this chapterexponentially much time, on every input

Three different answers in four rows, which is why each has to be proved rather than assumed.

The offline Turing machine

The definition. A machine with a separate read only input tape, bounded by end markers, and one or more read write work tapes.

Why it exists. Because it makes the space used by a computation measurable. On an ordinary machine the input occupies tape, so a machine cannot use less space than its input and asking about sublinear space is meaningless. On an offline machine the input tape is not counted, and the space is what the work tapes use.

That is chapter 79's definition, and it is why the offline machine is in the syllabus at all: without it the statement "this problem can be solved in logarithmic space" has nothing to mean.

It is no stronger. Copy the input onto a work tape and it becomes an ordinary multitape machine, which chapter 68 reduces to one tape.

And it is the machine chapter 60 restricted. A linear bounded automaton is an offline machine whose work space is bounded by the input's length, which is why the end markers appear in both definitions.

munotes.in347

Variants of the Turing Machine: Nondeterministic, Offline, and the Enumerator

The enumerator

The definition. A Turing machine with an extra output device, called a printer. The machine runs with a blank input tape, and whenever it likes it may print the current contents of a designated output tape and continue. It never halts unless it runs out of things to print.

The language it enumerates is the set of strings it eventually prints.

Three properties of the definition worth noticing.

Order does not matter. The strings may be printed in any order.

Repeats do not matter. A string printed twice is in the set once.

It need not halt. An infinite language needs a machine that runs for ever, and that is not a defect here: it is the point.

The theorem

Statement. A language is recursively enumerable if and only if some enumerator prints it.

The name is now explained. Chapter 27 called the class recursively enumerable and could only gesture at why. This is why: the class is exactly the languages a machine can list.

From an enumerator to an acceptor. Given an enumerator E and an input w, run E and compare each printed string with w. If one matches, accept. If w is never printed the machine runs for ever, which is permitted: the definition of accepting allows looping on a string not in the language.

From an acceptor to an enumerator. This is the direction with the technique, and the technique is the one chapter 27 called dovetailing.

The wrong construction: take the strings in canonical order, run the acceptor on each, and print the ones it accepts. It fails because the acceptor may loop on the first string, and nothing after it is ever reached.

The right construction: run in rounds. In round k, run the acceptor for k steps on each of the first k strings in canonical order. Print any string that is accepted within its allowance. Every string that is accepted at all is accepted within some number of steps, so it is printed in some round.

Dovetailing is the answer to every "but it might loop" objection in this block, and it is worth learning as a technique and not as a trick for this one theorem.

Enumerating in order

One more result, stated in chapter 27 and now provable.

A language is recursive if and only if some enumerator prints it in canonical order.

Why. If the strings come out in order, then to decide whether w is in the language, run the enumerator until it prints w, in which case accept, or prints something after w in the order, in which case reject. One of the two must happen, so the procedure halts.

munotes.in348

Variants of the Turing Machine: Nondeterministic, Offline, and the Enumerator

Conversely, a decider gives an in order enumerator: go through the strings in canonical order, run the decider on each, and print the ones it accepts.

So the three notions line up, and this completes chapter 27's table:

Can beMeans
accepted, possibly loopingrecursively enumerable
enumerated in some orderrecursively enumerable
enumerated in canonical orderrecursive
decided, always haltingrecursive

Distinctions

NondeterministicOfflineEnumerator
The changedelta returns a seta read only input tapean output printer, no input
Strongernonono
Cost to simulateexponential timenone worth statingnone
Why it existsit matches the way problems are statedto make sublinear SPACE measurableto explain the word enumerable
Depth firstBreadth first
On the computation treefollows one branch to the endvisits every node at depth d first
CorrectNO, an infinite branch traps ityes
What the choice sequence tape doesenumerates the branches by length

What it does NOT mean

Nondeterminism does not add power to a Turing machine. It adds convenience, and the simulation is exponential, which chapter 86 asks about.

Depth first search is not a valid simulation. An infinite branch traps it, and the accepting node beside it is never reached.

An enumerator is not required to print in order. Printing in order is a stronger property and it characterises the recursive languages.

An enumerator that never halts is not broken. An infinite language requires it.

The offline machine is not weaker. It is the ordinary machine with the input tape made read only, which costs nothing because the input can be copied.

Quick revision

  • Nondeterministic: delta returns a set, the computation is a tree, and acceptance means some node accepts.
  • Simulated BREADTH first, using a tape of choice sequences enumerated by length. Depth first fails because a

branch may be infinite.

  • The cost is about b to the t, which is exponential, and whether it can be improved is chapter 86's question.
  • Four models, three answers: finite automata yes, pushdown automata no, linear bounded automata unknown, Turing

machines yes.

  • Offline: a read only input tape plus work tapes, which is what makes sublinear space measurable and what chapter

60's machine restricts.

  • Enumerator: prints strings, in any order, with repeats allowed, never halting on an infinite language. A

language is recursively enumerable exactly when an enumerator prints it, which is where the name comes from.

  • Acceptor to enumerator needs DOVETAILING: in round k run the acceptor k steps on each of the first k strings.
  • Enumerating in CANONICAL order characterises the recursive languages.

Test yourself

1. Give the one change that makes a Turing machine nondeterministic, and the definition of acceptance. delta maps a state and a tape symbol to a SET of triples rather than one. The machine accepts when some sequence of choices reaches an accepting state.

munotes.in349

Variants of the Turing Machine: Nondeterministic, Offline, and the Enumerator

2. Why must the simulation search the computation tree breadth first? Because a branch may be infinite. A depth first search would follow such a branch for ever and never reach an accepting node on a different branch.

3. What does the simulation cost, and which famous question is about that cost? About b to the t for a machine running in time t with at most b choices per step, which is exponential. Whether the exponential can be reduced to a polynomial for polynomial time machines is the P against NP question of chapter 86.

4. Why does the offline machine exist? Because on an ordinary machine the input occupies tape, so no machine can use less space than its input, and sublinear space has nothing to mean. Making the input tape read only and not counting it is what allows space to be measured below the input's length.

5. Define an enumerator and state the theorem that explains the word enumerable. A Turing machine with a printer, run on a blank input, which may print strings as it goes and need never halt. A language is recursively enumerable exactly when some enumerator prints it.

6. Describe dovetailing and say what goes wrong without it. In round k, run the acceptor for k steps on each of the first k strings in canonical order, printing anything accepted within its allowance. Without it, running the acceptor to completion on the first string may never finish, and no later string is ever tried.

Contents This chapter on its own page

munotes.in350

Chapter Seventy

Recursive and Recursively Enumerable Languages

Syllabus topic Module 2, "Turing Machines: Decidability and Undecidability"

In one line

A language is recursively enumerable when a machine accepts exactly its members, and recursive when that machine also always halts.

In the wording a student can write in an examination: a language L is recursively enumerable if there is a Turing machine M with T(M) equal to L; on a string not in L, M may reject or may fail to halt. L is recursive, or decidable, if there is a Turing machine that accepts every string of L and rejects every string not in L, halting in both cases. Every recursive language is recursively enumerable, and the converse is false.

The two definitions, with the machine in hand

Chapter 27 gave these without a machine, because MU prints the label in Module 1. Nothing in them changes; what changes is that the words now refer to something defined.

RecursiveRecursively enumerable
Also calleddecidablesemi decidable, Turing recognisable
The machinea decider: always haltsa recogniser: may loop
On a memberaccepts, haltingaccepts, halting
On a non memberrejects, haltingrejects, or runs for ever

A decider is a recogniser, so every recursive language is recursively enumerable, and the proof is that sentence: a machine that always halts satisfies the weaker condition without alteration.

The converse fails, and chapter 74 provides the witness. Until then the two classes could in principle be equal, and this chapter's theorems are what make the difference usable once the witness exists.

The four closure properties

Both classes are closed under some operations and not others, and the differences are the useful part.

Both classes are closed under union

For recursive. Given deciders M1 and M2 for L1 and L2, build a machine that runs M1, then runs M2, and accepts if either accepted. Both always halt, so the combination does.

For recursively enumerable. The same construction does NOT work, because M1 may loop and M2 would never start. Run them side by side instead, one step of each in turn, and accept as soon as either accepts. That is chapter 27's dovetailing again, and this is its second use in the book.

Note the pattern. For a recursive class, sequencing is fine. For a recursively enumerable class, sequencing is wrong and dovetailing is right, every time. A student who remembers only that has most of this chapter.

Both are closed under intersection

Recursive. Run both deciders and accept if both accepted.

Recursively enumerable. Run both side by side and accept when both have accepted. Sequencing would be wrong for the same reason.

The recursive languages are closed under complement

The construction. Take the decider and swap its answers: accept where it rejected, reject where it accepted. It still always halts, so the result is a decider for the complement.

munotes.in351

Recursive and Recursively Enumerable Languages

This is chapter 38's construction, and the one condition it needs is exactly the one that holds here: the machine must always answer, so that swapping the answer is meaningful.

The recursively enumerable languages are NOT closed under complement

Swapping does not work, and the reason is precise. A recogniser has three outcomes and only two can be swapped. On a string where the machine loops there is nothing to swap, and the resulting machine loops too, so it accepts neither the string nor its complement's membership.

And the failure is real, not an artefact of that construction. Chapter 74 exhibits a recursively enumerable language whose complement is not recursively enumerable at all.

The theorem this chapter exists for

Statement. A language L is recursive if and only if both L and its complement are recursively enumerable.

The forward half. Suppose L is recursive. Then L is recursively enumerable, since a decider is a recogniser. And the complement is recursive by the closure above, hence recursively enumerable. Done.

The backward half, which is the useful one. Suppose M accepts L and N accepts the complement of L, each possibly looping on strings outside its own language. Build a machine D:

  1. On input w, run M and N side by side, one step of each in turn.
  2. If M accepts, accept.
  3. If N accepts, reject.

Why D always halts. Every string is either in L or in its complement. If w is in L then M accepts it, in some finite number of steps, and D reaches that point. If w is not in L then it is in the complement, so N accepts it in some finite number of steps. Either way one of the two accepts eventually, so D answers.

Why D is correct. It accepts exactly when M does, which is exactly when w is in L.

So L is recursive.

Why running M to completion first is wrong, which is the point an examiner checks: M may never finish on a string outside L, and then N is never started and D never answers. Dovetailing is the whole content of the proof.

The corollary that does the work

If L is recursively enumerable and its complement is not, then L is not recursive.

Proof. If L were recursive, its complement would be recursive by the closure property, hence recursively enumerable, which it is not.

This is how the second and third undecidable problems get proved. Chapter 74 proves one language undecidable the hard way, by diagonalisation, and then this corollary and the reductions of chapter 75 do the rest.

munotes.in352

Recursive and Recursively Enumerable Languages

And the picture it gives. For any language L, exactly one of four things holds:

L recursively enumerablenot
complement recursively enumerableL is recursiveimpossible
notL is r.e. and undecidableneither is r.e.

The top right cell is impossible by symmetry of the theorem, which is worth noticing: if the complement is recursively enumerable and L is not, then reading the theorem with L and its complement swapped puts us in the bottom left cell instead. So there are three genuine cases, and chapter 74 provides an example of the second and chapter 75 of the third.

Where the four language classes now stand

ClassMachineClosed under complementMembership
context sensitivelinear bounded automatonyes, since 1988decidable
recursivea Turing machine that always haltsyesdecidable
recursively enumerablea Turing machinenonot decidable in general
all languagesnone

And the containments are proper. Context sensitive inside recursive, by a diagonal argument over the context sensitive grammars, which can be listed. Recursive inside recursively enumerable, by chapter 74's halting language. Recursively enumerable inside all languages, by chapter 7's counting argument.

Three different techniques for three containments, and a student who can name which technique settles which gap understands the block.

Distinctions

SequencingDovetailing
Run M then Nyesno
Run one step of each in turnnoyes
Correct for decidersyesyes, needlessly
Correct for recognisersNOyes
WhyM may never finishneither machine is waited on
A deciderA recogniser
Halts on every inputyesnot necessarily
Its language isrecursiverecursively enumerable
Swapping its answers givesthe complementnothing useful

What it does NOT mean

Recursively enumerable does not mean you can list the non members. That is the complement, and the theorem says having both lists makes the language decidable.

A machine that has not halted has not rejected. It has not answered, and the whole chapter turns on that.

Not closed under complement does not mean no complement is recursively enumerable. Many are; the class is not closed because SOME complement is not.

Recursive here has nothing to do with a function calling itself, as chapter 27 said. It means decidable.

Dovetailing is not an optimisation. It is the correct construction, and sequencing is wrong rather than slow.

Quick revision

  • Recursively enumerable: a machine accepts exactly the members and may loop on a non member. Recursive: it also

always halts.

  • Every recursive language is recursively enumerable; the converse fails, and chapter 74 gives the witness.
  • Both classes are closed under union and intersection; for the recursively enumerable class the machines must be

run side by side, not one after the other.

  • The recursive languages are closed under complement, by swapping the decider's answers. The recursively
munotes.in353

Recursive and Recursively Enumerable Languages

enumerable ones are not, because a looping outcome cannot be swapped.

  • The theorem: L is recursive exactly when L and its complement are both recursively enumerable, proved by

dovetailing.

  • The corollary that does the work: if L is recursively enumerable and its complement is not, L is not recursive.
  • Three cases for any L: recursive; recursively enumerable and undecidable; or neither it nor its complement

recursively enumerable.

Test yourself

1. Give both definitions and the single difference. Recursively enumerable: some Turing machine accepts exactly the strings of L, and on the others may reject or loop. Recursive: some Turing machine accepts exactly the strings of L and rejects all the others, always halting. The difference is whether the machine must answer.

2. Prove that both classes are closed under union, and say where the two proofs differ. For recursive, run one decider then the other and accept if either accepted; both halt so the combination does. For recursively enumerable, run them side by side one step at a time and accept as soon as either accepts; sequencing would be wrong because the first machine may never finish.

3. Why are the recursively enumerable languages not closed under complement? Because swapping a recogniser's answers does not work: on an input where it loops there is nothing to swap, and the new machine loops too. And chapter 74 exhibits a recursively enumerable language whose complement is not recursively enumerable.

4. State and prove the complement theorem. L is recursive exactly when L and its complement are both recursively enumerable. Forwards, a decider gives recognisers for both. Backwards, run the two recognisers side by side; every input is in one language or the other, so one must accept eventually, and the combined machine answers accordingly and always halts.

5. Give the corollary and say what it is used for. If L is recursively enumerable and its complement is not, then L is not recursive. It is how the second and later undecidable problems are established once one has been proved undecidable directly.

6. Name the technique that settles each of the three containments. Context sensitive inside recursive by a diagonal argument over the grammars; recursive inside recursively enumerable by chapter 74's halting language; recursively enumerable inside all languages by chapter 7's counting argument.

Contents This chapter on its own page

munotes.in354

Chapter Seventy-One

Decidable and Undecidable, and What the Complement Tells You

Syllabus topic Module 2, "Turing Machines: Decidability and Undecidability"

In one line

A decision problem is decidable when some algorithm answers it for every instance and always stops, and undecidable otherwise.

In the wording a student can write in an examination: a decision problem is a question about an input admitting a yes or no answer. It is encoded as a language, the set of encodings of the instances whose answer is yes. The problem is decidable if that language is recursive, and undecidable otherwise. A problem may be undecidable and still have a recursively enumerable language, in which case the yes instances can be recognised but the no instances cannot be ruled out.

A problem is a language

This is the step that makes the whole block work and it is the one a note usually skips.

A decision problem is a question with a yes or no answer about an input: "is this number prime?", "does this grammar generate the empty string?", "does this machine halt on this input?".

To study it mathematically it must become an object, and the object is a language.

  1. Fix a way of writing each instance as a string. Chapter 73 does this for a Turing machine; for a number or a

graph it is routine.

  1. Collect the encodings of the instances whose answer is yes.
  2. That set of strings is a language, and it is the problem.

the problem is DECIDABLE exactly when that language is RECURSIVE

Three things follow, and each is used later.

Everything about languages applies to problems. The closure results of chapter 70 become statements about problems, and the corollary about complements becomes a way of proving problems undecidable.

The encoding does not matter, provided it is reasonable: any two sensible encodings can be converted into each other by a machine, so a problem decidable under one is decidable under the other.

An invalid string is a no. A string that encodes no instance at all is simply not in the language, and the machine rejects it, which is why the encoding chapter has to make invalid strings recognisable.

Decidable, undecidable, and the middle

The three cases of chapter 70, read as statements about problems.

The language of the problem isThe problem isWhat a machine can do
recursivedecidablealways answer, yes or no
recursively enumerable, not recursiveundecidable, but semi decidableconfirm every yes, never rule out a no
not recursively enumerableundecidable, and not even semi decidablenot even confirm the yes instances

The middle row is the interesting one and it is where the halting problem lives. A machine for it can watch and eventually say yes whenever the answer is yes; on a no it watches for ever.

munotes.in355

Decidable and Undecidable, and What the Complement Tells You

The third row is real too. Chapter 70's corollary says that if a problem is semi decidable and undecidable, then its complement is in the third row, so every example of the middle row supplies an example of the third.

Decidable problems you have met

Worth listing, because a block about undecidability can leave a student thinking nothing is decidable.

ProblemDecidableWhere
is w in L(M), for a finite automaton Myeschapter 39
is L(M) empty, for a finite automatonyeschapter 39
do two finite automata accept the same languageyeschapter 39
is w in L(G), for a context free grammaryes, by CYKchapter 51
is L(G) empty, for a context free grammaryeschapter 51
is w in L(M), for a linear bounded automatonyeschapter 61
is a given grammar ambiguousnochapter 77
do two context free grammars agreenochapter 77
does a Turing machine halt on a given inputnochapter 74

Read the pattern down the column. Everything about finite automata is decidable. Most things about context free grammars are. Almost nothing about Turing machines is. The stronger the model, the less can be decided about it, and that is the trade chapter 26 named for grammars and chapter 61 for the machines.

The complement, and what it tells you

Chapter 70's corollary in the form it is used.

If a problem is semi decidable and its complement is not, the problem is undecidable.

And the more useful contrapositive: if a problem and its complement are both semi decidable, the problem is decidable.

How this is used in practice, and it is the standard way of finishing a proof:

  1. Show the problem is semi decidable, usually by simulation: run the machine, or search for a derivation, and

say yes when found.

  1. Show the complement is not semi decidable, usually by a reduction from a problem already known to be in

that state.

  1. Conclude the problem is undecidable.

Chapter 75 is step 2 as a technique.

A worked classification

Three problems, each placed in one of the three rows, with the reason.

Does this Turing machine accept this string? Semi decidable: simulate the machine and accept if it accepts. Undecidable: chapter 74. So the middle row.

Does this Turing machine REJECT this string, or loop on it? That is the complement of the above. By the corollary it is not semi decidable, so the third row.

Does this Turing machine have at least three states? Decidable, trivially: count the states in the encoding. The top row.

The third of these is worth a sentence, because it shows what the undecidable results are and are not about. A question about the description of a machine is usually easy. A question about its behaviour is usually impossible, and chapter 76's Rice's theorem makes that precise.

munotes.in356

Decidable and Undecidable, and What the Complement Tells You

Two words that are used loosely and should not be

Solvable and unsolvable are used in this literature as synonyms for decidable and undecidable, and MU's syllabus uses "Unsolvable Problems" as a label. They mean the same thing.

Computable properly applies to functions and decidable to problems, and a problem's decidability is the computability of its yes or no answer as a function. The two are not different notions.

Recognisable means semi decidable, that is, recursively enumerable. A student meeting "Turing recognisable" in a modern textbook and "recursively enumerable" in MU's is meeting one class under two names.

Distinctions

DecidableSemi decidableNeither
The language isrecursiverecursively enumerablenot even that
A machine canalways answerconfirm a yesnot even that
Closed under complementyesno
Examplemembership for a finite automatonhaltingthe complement of halting
A question about a machine's descriptionabout its behaviour
Examplehow many states has itdoes it halt
Usuallydecidable, and easyundecidable
The general statementRice's theorem, chapter 76

What it does NOT mean

Undecidable does not mean unanswerable for a particular instance. Any one instance may be easy; the claim is that no single method works for all of them.

Undecidable does not mean hard. Chapter 86's NP complete problems are decidable and hard. Undecidable means no algorithm exists at all.

Semi decidable is not half decidable in any useful practical sense. A machine that says yes eventually and never says no is of limited use, and its value is theoretical: it is what puts the problem in the middle row rather than the bottom.

The encoding is not a detail to worry about. Any two reasonable encodings give the same answer, and a question about the encoding is not a question about the problem.

Decidable is not a property of a machine. It is a property of a problem, and it asks whether SOME machine decides it.

Quick revision

  • A decision problem is a yes or no question about an input; encode each instance as a string and collect the yes

instances, and the problem IS that language.

  • Decidable means the language is recursive; semi decidable means recursively enumerable; and a problem may be

neither.

  • Everything about finite automata is decidable, most things about context free grammars are, and almost nothing

about Turing machines is.

  • The corollary in use: semi decidable with a complement that is not semi decidable gives undecidable; both semi

decidable gives decidable.

  • The standard proof shape: show semi decidability by simulation, then show the complement is not semi decidable
munotes.in357

Decidable and Undecidable, and What the Complement Tells You

by a reduction.

  • A question about a machine's description is usually decidable; a question about its behaviour is usually not,

which chapter 76 makes precise.

  • Solvable and unsolvable mean decidable and undecidable; recognisable means recursively enumerable.

Test yourself

1. How does a decision problem become a language? Fix an encoding of instances as strings, and take the set of encodings of the instances whose answer is yes. The problem is decidable exactly when that language is recursive.

2. Give the three possible statuses of a problem, with an example of each. Decidable, such as membership for a finite automaton. Semi decidable but undecidable, such as whether a Turing machine halts on a given input. Not even semi decidable, such as the complement of that.

3. State the test using complements, in both directions. If a problem and its complement are both semi decidable then the problem is decidable. Contrapositively, if a problem is semi decidable and its complement is not, the problem is undecidable.

4. Is "does this Turing machine have more than five states" decidable? Why? Yes, and easily: count the states in the encoding. It is a question about the machine's description rather than its behaviour.

5. What is the difference between undecidable and intractable? Undecidable means no algorithm exists at all. Intractable, in the sense of chapter 86's NP complete problems, means an algorithm exists and no fast one is known. The two are unrelated.

6. Name three words that mean recursively enumerable. Recursively enumerable, semi decidable, and Turing recognisable. MU's syllabus uses the first, modern textbooks the third.

Contents This chapter on its own page

munotes.in358

Chapter Seventy-Two

The Church Turing Thesis

Syllabus topic Module 2, "Turing Machines: The Church-Turing thesis"

In one line

The thesis says that everything a human being could compute by following a definite method, a Turing machine can compute, and it is believed rather than proved.

In the wording a student can write in an examination: the Church Turing thesis asserts that the intuitive notion of an effectively calculable function coincides with the class of functions computable by a Turing machine. It is not a theorem, because the intuitive notion is not a mathematical object and so cannot be the subject of a proof. It is supported by the equivalence of every formal model of computation that has been proposed.

What it claims

Two notions are being identified.

The informal one: effectively calculable. A function is effectively calculable if there is a definite procedure a person could follow, with paper and pencil, needing no insight at any step, that produces the answer in finitely many steps. That is not a definition in the mathematical sense; it is a description of what people mean.

The formal one: Turing computable. A function is Turing computable if some Turing machine computes it, in the sense of chapter 67. That IS a definition, and everything about it can be proved.

The thesis is that they are the same class, and it goes one way in each direction:

every effectively calculable function is Turing computable

every Turing computable function is effectively calculable

The second half is uncontroversial. A Turing machine is a definite procedure with no insight in it; a person could follow one with paper and pencil, slowly. So the content of the thesis is the first half.

Why it is not a theorem

Because one side of it is not mathematics.

To prove "every effectively calculable function is Turing computable" one would have to have a definition of effectively calculable to argue from. The only definitions on offer are formal models, and proving the thesis against one of them proves the two models equivalent, which is a theorem about two formal objects and not the thesis.

So the thesis is a claim about the adequacy of a definition, of the same kind as the claim that the epsilon and delta definition captures what a limit is. Such claims are settled by evidence and by use, not by proof.

It could in principle be refuted. Exhibit a function that is effectively calculable, by a procedure everyone agrees requires no insight, and prove no Turing machine computes it. Nobody has, in ninety years.

The evidence

Four kinds, and an examination answer that gives two of them has answered the question.

One: the models all coincide

Every formal model of computation anybody has proposed turns out to compute exactly the Turing computable functions. Turing's own paper says so about the first of them, in its introduction:

munotes.in359

The Church Turing Thesis

In a recent paper Alonzo Church has introduced an idea of "effective calculability", which is equivalent to my

"computability", but is very differently defined.

The footnote on that sentence gives Church's paper as "An unsolvable problem of elementary number theory", American Journal of Mathematics volume 58 (1936), pages 345 to 363, and Turing adds that "the proof of equivalence between 'computability' and 'effective calculability' is outlined in an appendix to the present paper".

So the two definitions were given independently, in the same year, from quite different starting points, and turned out to define the same class. That is the first and strongest piece of evidence, and it is why the thesis carries both names.

And it has kept happening. Recursive functions, register machines, cellular automata, every programming language, and every variant of chapter 68 and 69 all give the same class. A definition that survives that many independent attempts to change it is not accidental.

Two: the robustness of the model itself

Chapters 68 and 69 proved that adding tapes, tracks, heads, nondeterminism or a read only input tape changes nothing. A model this insensitive to its own details is describing something other than its details.

Three: Turing's own argument

Turing did not merely define a machine and hope. Section 9 of his paper argues that the machine captures what a human computer does, and the argument is worth knowing because it is the only part of the subject that reasons about people rather than about mathematics.

His starting point is stated in section 1 of the paper: the justification for the definition, he writes, "lies in the fact that the human memory is necessarily limited". From there he compares a person computing to a machine with finitely many conditions, which he calls m-configurations, reading a tape divided into squares, aware only of the square being scanned, and able to remember earlier symbols only by changing its condition.

Read as an argument it has three steps, and they are the three restrictions of chapter 62's definition:

A person can hold only finitely many things in mind at once, so the machine has finitely many states.

A person writes on paper, which is divided into cells, and can attend to only one at a time, so the machine has a tape and a single scanned square.

A person can move attention only a bounded distance in one step, so the head moves one cell.

That is an argument about people, and it is the closest anyone has come to proving the thesis.

Four: nobody has beaten it

Ninety years of trying, including with models designed to be stronger, and the class has not moved.

munotes.in360

The Church Turing Thesis

What the thesis licenses

This is the practical reason the chapter matters, because the thesis is used constantly and usually silently.

It lets an algorithm be described in words. Chapter 65's high level descriptions, chapter 69's dovetailing, chapter 61's simulation of a grammar: none of those is a transition table, and each is accepted as defining a Turing machine. The licence for that is the thesis.

It turns "no Turing machine" into "no algorithm". When chapter 74 proves no Turing machine decides halting, the conclusion anybody wants is that no method decides halting, and the thesis is the step between them.

Without it, chapter 74 would be a fact about one machine model. With it, it is a fact about computation.

A student should say so in an answer. When a proof says "and by the Church Turing thesis, no algorithm exists", that sentence is doing real work and is not a flourish.

What the thesis does NOT claim

Five things, and each is a misreading that appears in examination answers.

It does not say a Turing machine is efficient. Nothing about time or space. Chapters 78 onward are a separate subject, and the thesis is silent about all of them.

It does not say the brain is a Turing machine. It says that what a person can compute by following a definite method is Turing computable. Anything a person does that is not following a definite method is outside the claim.

It does not say a real computer is a Turing machine. A real computer has finite memory, so strictly it is a finite automaton; the machine is the idealisation that removes the memory limit, which is the useful direction.

It is not proved and it is not provable as stated, because one side is informal.

It is not about what is true, only about what is calculable. Hilbert's question was about provability, which is a different notion, and chapter 74's answer to it goes through computability.

Distinctions

A theoremA thesis
Both sides aremathematicalone is informal
Settled byproofevidence and use
Can be refutedby finding an errorby exhibiting a counterexample
This onethe equivalence of two MODELSeffectively calculable equals Turing computable
Church's formulationTuring's
The formal notioneffective calculability, defined by his own calculuscomputability by a machine
Year19361936
Turing's own words"equivalent to my computability, but very differently defined"
Whose name is firstby convention, Church's

What it does NOT mean

Equivalence of two models is not the thesis. It is a theorem, and it is evidence for the thesis.

Believing the thesis is not an act of faith. It is an inference from ninety years of independent attempts converging.

munotes.in361

The Church Turing Thesis

The thesis does not make undecidability less real. It makes it more: it is what allows a result about one machine model to be read as a result about method in general.

Church and Turing did not collaborate on it. Turing's paper describes Church's work as a "recent paper" and outlines the equivalence in an appendix, which is what happens when two people reach the same place separately.

Quick revision

  • The thesis: the informal class of effectively calculable functions is exactly the formal class of Turing

computable functions.

  • Not a theorem, because one side is not a mathematical object. It is settled by evidence and could be refuted by

a counterexample.

  • The evidence: every model ever proposed coincides; the machine is insensitive to its own details; Turing's own

argument about the limits of a human computer; and ninety years without a counterexample.

  • Turing's paper itself records that Church's independent notion is "equivalent to my computability, but very

differently defined", and cites Church's 1936 paper in the American Journal of Mathematics.

  • Its use: it lets an algorithm be described in words, and it turns "no Turing machine can" into "no algorithm

can", which is what makes chapter 74 a result about computation rather than about one model.

  • It says nothing about efficiency, nothing about the brain, and nothing about real computers' finite memory.

Test yourself

1. State the thesis and say which half carries the content. That the effectively calculable functions are exactly the Turing computable ones. The half that carries the content is that every effectively calculable function is Turing computable; the converse is uncontroversial, since a Turing machine is itself a definite procedure.

2. Why can it not be proved? Because effective calculability is an informal notion rather than a mathematical object, so there is nothing to argue from. Proving it against a formal model only proves two models equivalent.

3. Give two pieces of evidence for it. That every formal model proposed, from Church's calculus onward, computes exactly the same class; and that the machine is insensitive to its own details, since tapes, tracks, heads and nondeterminism all change nothing.

4. What does Turing's own paper say about Church's work? That Church "has introduced an idea of effective calculability, which is equivalent to my computability, but is very differently defined", and that the proof of the equivalence is outlined in an appendix to Turing's paper.

5. What does the thesis license in a proof? Describing an algorithm in words rather than as a transition table, and reading a result of the form "no Turing machine can do this" as "no algorithm can do this".

munotes.in362

The Church Turing Thesis

6. Name two things the thesis does NOT claim. That a Turing machine is efficient, since it says nothing about time or space; and that the brain is a Turing machine, since it concerns only what can be done by following a definite method.

Contents This chapter on its own page

munotes.in363

Chapter Seventy-Three

The Universal Turing Machine

Syllabus topic Module 2, "Turing Machines: Universal Turing Machine"

In one line

There is one Turing machine that, given a description of any Turing machine and an input, does what that machine would do.

In the wording a student can write in an examination: a universal Turing machine U is a Turing machine which, given as input an encoding of a Turing machine M followed by a string w, simulates the computation of M on w, accepting if M accepts w and halting without accepting if M rejects, and failing to halt if M fails to halt. Its existence shows that a single fixed machine suffices for all computation.

Why this is the most consequential machine in the book

Because it says a machine need not be built for each job.

Every machine so far was made for one language. The universal machine is made for all of them: it takes the job as data. That is the idea of a stored program computer, and Turing published it in 1936, nine years before one was built.

And it is what makes chapter 74 possible. A proof that some problem about machines is undecidable needs machines to be objects that other machines can be handed. The encoding below is what makes them objects.

Encoding a machine as a string

The universal machine's input is a string, so a machine must be written as one. Any reasonable scheme works; this book uses the standard one because it is what MU's textbooks print and because it uses only two symbols, which keeps the point clear.

Write everything in unary, separated by single zeros.

ItemWritten as
state qii copies of 1
tape symbol Xjj copies of 1
direction Lone 1
direction Rtwo 1

Number the states from 1, with q1 the start state and q2 the single accepting state, which chapter 68 said costs nothing. Number the tape symbols from 1, with X1 the symbol 0, X2 the symbol 1 and X3 the blank, and further working symbols after those.

One move delta(qi, Xj) = (qk, Xl, Dm) becomes

1 to the i, 0, 1 to the j, 0, 1 to the k, 0, 1 to the l, 0, 1 to the m

that is, five unary numbers separated by single zeros.

The whole machine is its moves in any order, separated by double zeros, with a triple zero at the end to mark where the machine stops and the input begins.

Why the separators have different lengths. A single zero separates fields within a move, a double separates moves, and a triple separates the machine from its input. Because every field is a non empty block of 1, a run of zeros is never ambiguous.

munotes.in364

The Universal Turing Machine

A machine encoded, in full

Take the smallest useful machine: it accepts any string beginning with a 0, and rejects any beginning with a 1.

start q1
blank B
accept q2
q1, 0 -> q2, 0, R
q1, 1 -> q3, 1, R

Accepts: 0, 00, 01, 011

Rejects: ε, 1, 10, 11

Two moves. The second goes to q3, which has no moves at all, so the machine halts there without accepting, which is how it rejects a string beginning with 1. The empty string is rejected because q1 has no move on the blank.

Number everything. The states are q1 equal to 1, q2 equal to 11 and q3 equal to 111. The tape symbols are 0 equal to X1, 1 equal to X2 and the blank equal to X3, so they are 1, 11 and 111. The direction R is 11.

The first move, delta(q1, X1) equal to (q2, X1, R), becomes five unary fields separated by single zeros:

1 0 1 0 11 0 1 0 11

which written without spaces is

10101101011

The second move, delta(q1, X2) equal to (q3, X2, R), becomes

1 0 11 0 111 0 11 0 11

which written without spaces is

10110111011011

The whole machine is the two moves joined by a double zero:

101011010110010110111011011

Twenty seven symbols. Read it back to check: eleven symbols of the first move, then 00, then fourteen of the second, which is 11 plus 2 plus 14, and that is 27.

And with an input, after a triple zero, the machine and the string 01 together are

10101101011001011011101101100001

That string is a program and its data, written in one line over the alphabet {0, 1}, which is the whole point of the chapter.

One check the encoding must pass, and it is worth doing once. No triple zero may appear inside the code itself, or the boundary would be ambiguous. It cannot: every field is a non empty run of 1, so within a move the zeros are single, and between two moves the double zero has a 1 on each side. The code above contains no run of three zeros anywhere.

What the encoding buys

Three things, and the third is chapter 74's.

Machines can be counted. Every machine is a binary string, so the machines can be listed in canonical order, which is chapter 7's argument made concrete. That argument said there are countably many machines; here is the list.

Machines can be inputs. A machine can be handed to another machine, which is what the universal machine does.

A machine can be handed its own code. Nothing forbids running the machine whose code is s on the input s, and chapter 74's proof does exactly that. This is the self reference that makes the diagonal argument work, and it is available only because a machine is a string.

munotes.in365

The Universal Turing Machine

An invalid string is a machine too, by convention: a string that does not decode to a legal machine is taken to encode some fixed machine that accepts nothing. That convention matters, because it means every string encodes a machine, so the list of machines is the list of strings and no gaps have to be handled.

How the universal machine works

U is usually given three tapes, which chapter 68 says costs nothing but a square in time.

TapeHolds
1the code of M, unchanged throughout
2the tape of M as the simulation proceeds
3the current state of M, in unary

The cycle, repeated:

  1. Read the symbol under the simulated head on tape 2, and the state on tape 3.
  2. Scan tape 1 for the move whose first two fields match them.
  3. If there is none, M has halted; halt, accepting if the state on tape 3 is the accepting one.
  4. Otherwise write the third field onto tape 3, write the fourth field onto tape 2 at the head position, and move

the simulated head on tape 2 according to the fifth.

Step 2 is the only interesting one, and it is a search over a fixed string, which is a finite automaton's work. Everything else is copying.

So U is not clever. It is a lookup loop, and the surprise is not that it is difficult but that it is easy: a machine that does everything is no more complicated than the machines it simulates.

What U does on each of the three outcomes

This is the part chapter 74 needs, and it must be exact.

M on wU on the code of M and w
acceptsaccepts
rejects, haltinghalts without accepting
runs for everruns for ever

The third row is not a defect of U. U cannot detect that M is looping, because it is simulating faithfully and a faithful simulation of an endless computation is endless. Chapter 74 proves that no machine could do better.

So the language U accepts is the set of strings encoding a machine and an input where the machine accepts the input. That language has a name, the universal language, and chapter 74 shows it is recursively enumerable and not recursive, which is the middle row of chapter 71's table.

The consequences

The stored program computer. A processor is a universal machine; a program in memory is the code on tape 1. The idea that a computer is a machine that takes machines as data is this theorem.

munotes.in366

The Universal Turing Machine

An interpreter. A Python interpreter is a universal machine for Python, and the fact that one can be written in Python is the same self reference this chapter makes available.

A compiler. It reads a program, which is data, and writes another, which is also data.

And the negative results. Everything from chapter 74 to chapter 77 requires machines to be data, and this chapter is what makes them so.

Distinctions

An ordinary Turing machineThe universal machine
Built forone languageall of them
Its inputa string of the language's alphabeta machine and a string
Corresponds toa programa computer
Number neededone per jobone, ever
SeparatorLengthSeparates
a single zero1fields within one move
a double zero2one move from the next
a triple zero3the machine from its input

What it does NOT mean

U is not more powerful than the machines it simulates. It accepts a recursively enumerable language like any other Turing machine. Its power is in being one machine rather than many.

U cannot detect a loop. It simulates faithfully, so it loops when its subject loops, and chapter 74 proves that no machine can do otherwise.

The encoding is not the only one. Any scheme a machine can read works, and the results do not depend on which is chosen.

An invalid code is not an error case. By convention it encodes a fixed machine accepting nothing, so every string is a machine and no gaps arise.

Being able to hand a machine its own code is not a paradox. It is a fact about strings, and chapter 74 uses it to derive a contradiction from a supposed decider, which is a different thing.

Quick revision

  • A universal Turing machine takes the code of a machine M and a string w and simulates M on w, matching all

three outcomes including looping.

  • The encoding: everything in unary, separated by single zeros within a move, double zeros between moves, and a

triple zero before the input. States and symbols are numbered from 1, with q1 the start state and q2 the accepting one.

  • Every string encodes a machine, invalid ones by convention encoding a machine that accepts nothing.
  • U runs three tapes: the code, the simulated tape, and the simulated state; its cycle is find the matching move

and apply it.

  • The encoding buys three things: the machines can be listed, a machine can be given to a machine, and a machine

can be given ITS OWN code, which is what chapter 74 needs.

  • The idea is the stored program computer, published in 1936 and built nine years later.
munotes.in367

The Universal Turing Machine

Test yourself

1. What does a universal Turing machine do, and how does it behave in each of the three cases? Given the code of a machine M and a string w it simulates M on w: it accepts when M accepts, halts without accepting when M rejects, and runs for ever when M does.

2. Encode the move delta(q1, X1) = (q2, X1, R) in the scheme of this chapter. 1 0 1 0 11 0 1 0 11, that is, the string 10101101011: five unary fields separated by single zeros.

3. Why do the separators have three different lengths? So that the structure is unambiguous: fields within a move are separated by one zero, moves by two, and the machine from its input by three, and since every field is a non empty block of 1 a run of zeros is always readable.

4. Why does every string have to encode some machine? So that the machines can be listed as the strings are, with no gaps to handle. A string that does not decode to a legal machine is taken by convention to encode a fixed machine that accepts nothing.

5. What three things does the encoding make possible? Listing all the machines, handing one machine to another, and handing a machine its own code, which is the self reference chapter 74's proof uses.

6. Why can U not report that the machine it is simulating will never stop? Because it simulates faithfully, so an endless computation gives an endless simulation. Chapter 74 proves that no machine, however cleverly built, can do better.

Contents This chapter on its own page

munotes.in368

Chapter Seventy-Four

The Halting Problem

Syllabus topic Module 2, "Turing Machines: Halting Problem"

In one line

No program can read another program and say whether it will finish.

In the wording a student can write in an examination: the halting problem is the problem of determining, for an arbitrary Turing machine M and input w, whether M halts on w. It is undecidable: there is no Turing machine that, given the encoding of M and the string w, always halts and correctly reports whether M halts on w.

The problem, stated exactly

The input. An encoding of a Turing machine M, by chapter 73's scheme, followed by a string w.

The question. Does M halt when run on w?

As a language, by chapter 71's translation:

HALT = { the code of M, then w : M halts on w }

The claim. HALT is not recursive. No machine decides it.

What is being claimed and what is not. It is not claimed that halting is hard to decide for a particular machine. Any one machine may be easy: chapter 66's machines obviously halt and chapter 64's looping machine obviously does not on some inputs. The claim is that no single method works for all of them.

The proof

The proof is by contradiction and it is half a page. Every step is written out.

Step 1. Suppose, for contradiction, that HALT is decidable.

Then there is a Turing machine H which always halts and behaves as follows:

GivenH does
the code of M and w, where M halts on whalts and says YES
the code of M and w, where M does not halt on whalts and says NO

Note the two things the supposition buys: H always halts, and H is always right. Both are used.

Step 2. Build a new machine D from H.

D takes one input, a string s, and does this:

  1. Treat s as the code of a machine. Run H on the pair (s, s), that is, ask H whether the machine coded by s

halts when given its own code as input.

  1. If H says YES, go into a deliberate infinite loop.
  2. If H says NO, halt.

D is the machine a question may call the INVERTED halting machine, because it takes the decider's answer and does the opposite: it loops where H says halt, and halts where H says loop. MU's 2020-21 paper uses that name, so it is worth recognising; the machine is the D of step 2 and nothing else.

D exists, and that is the step to be sure of. H exists by the supposition; running H on (s, s) needs only copying the input, which chapter 67 showed a machine can do; and the two branches need one looping state and one halting state. So if H exists then D exists, and D is a perfectly ordinary Turing machine with a finite description.

munotes.in369

The Halting Problem

Step 3. D behaves like this, by construction:

D halts on s exactly when the machine coded by s does NOT halt on s

Read it again, because the whole proof is in that line. D was built to do the opposite of whatever H reports.

Step 4. Run D on its own code.

D has a code, because chapter 73 says every machine does. Call it d. Now ask: does D halt on d?

Put s equal to d in step 3's line. The machine coded by d is D itself. So:

D halts on d exactly when D does NOT halt on d

That is a contradiction. Whichever of the two is true, the line says the other is. There is no consistent answer.

Step 5. So the supposition was false. D was built from H by an ordinary construction, so if D cannot exist then H cannot. Therefore no machine decides HALT, and the halting problem is undecidable.

Where each ingredient was used

Worth listing, because an examination answer that names them shows the proof was understood rather than memorised.

IngredientWhere it came fromWhat it did
a machine is a stringchapter 73's encodinglet D be given its own code
every string is a machinechapter 73's conventionlet D treat any s as a machine
a machine can copy its inputchapter 67let D form the pair (s, s)
deliberate looping is allowedchapter 64's third outcomelet D do the opposite of halting
proof by contradictionchapter 6the shape of the whole argument

The one that is easiest to overlook is the third: D has to turn s into the pair (s, s), and that is a real though small piece of machine.

Why this is a diagonal argument

Because it is chapter 7's argument in another costume, and seeing that is worth a paragraph.

Chapter 7 listed the languages, built a table of which strings each contained, walked down the diagonal and flipped every entry. The result disagreed with every row.

Here, list the machines, and build a table whose row for machine number i and column for string number j says whether machine i halts on string j. D walks down the diagonal and flips. D is a machine, so it is some row of the table, and it disagrees with that row at the diagonal entry, which is the contradiction.

So the same technique settles both, and the difference is that chapter 7's argument named no language while this one names a problem anybody can state in a sentence. Naming one is the hard part, and it is what made Turing's paper famous.

munotes.in370

The Halting Problem

What the theorem does and does not say

Five statements, and the examination tests whether a student can tell them apart.

It does not say a particular program's halting cannot be determined. Most programs are obviously fine. The claim is about a method that works for every program.

It does not say the problem is hard. Hard is chapter 86. This problem has no algorithm at all, at any cost.

It does not depend on the machine model. Chapters 68 and 69 showed the variants are equivalent, so a proof about one is a proof about all, and by chapter 72's thesis it is a statement about method in general.

It does not say useful approximations are impossible. A tool can answer for many programs and give up on the rest, and real tools do. What cannot exist is one that always answers and is always right.

It does not make programming impossible. It says one particular question about programs has no general answer, and a great many other questions do.

What follows immediately

HALT is recursively enumerable. Simulate: run M on w with the universal machine of chapter 73, and say yes when it halts. If it never halts, the simulation never halts, which is permitted. So HALT sits in the middle row of chapter 71's table: semi decidable and undecidable.

The complement of HALT is NOT recursively enumerable, by chapter 70's corollary: if it were, HALT would be recursive, which it is not. So the complement is an example of the bottom row, which chapter 71 promised.

A whole family of problems follows. Chapter 75 turns this one result into many by reduction, and chapter 76 turns it into a general theorem about every non trivial property of a machine's language.

A related result MU may set: the halting problem for a fixed input

Sometimes the question is posed as: does M halt on the empty input?

Also undecidable, and the proof is a reduction, which chapter 75 does properly. In outline: given M and w, build a machine M prime that ignores its own input, writes w on the tape, and then runs M. Then M prime halts on the empty input exactly when M halts on w. So a decider for the fixed input version would give a decider for the general one, and there is none.

Note the direction, because it is the thing chapter 75 warns about. The general problem was reduced TO the fixed input problem, which shows the fixed input problem is at least as hard. Reducing the other way would prove nothing.

munotes.in371

The Halting Problem

Distinctions

The halting problemThe universal language
Asksdoes M halt on wdoes M accept w
Recursively enumerableyesyes
Recursivenono
Related bya reduction each way
UndecidableIntractable
Meansno algorithm existsno fast algorithm is known
Examplehaltingthe NP complete problems of chapter 84
Costnot applicable, there is nothing to runexponential, as far as anybody knows

What it does NOT mean

Halting is not undecidable for finite automata. They always halt. The theorem is about Turing machines, and the third outcome of chapter 64 is what makes the question meaningful at all.

The proof does not need the universal machine to be built. It needs only that machines are strings and that one machine can be run on another's code, which chapter 73 supplies.

D is not a strange or illegal machine. It is built by an ordinary construction from a machine assumed to exist. The impossibility falls on H, not on D.

A program that has run for a long time is not evidence. Chapter 64 said it and this is why.

Undecidable is not the same as unprovable. Hilbert's question was about provability and Turing answered it through computability, which is a route rather than an identity.

Quick revision

  • HALT is the set of pairs, a machine's code and an input, such that the machine halts on the input. It is

undecidable.

  • The proof: suppose a decider H exists; build D which, on input s, asks H whether the machine coded by s halts

on s, and then does the opposite; run D on its own code; D halts exactly when it does not.

  • Five ingredients: a machine is a string; every string is a machine; a machine can copy its input; deliberate

looping is allowed; and proof by contradiction.

  • It is chapter 7's diagonal argument, over a table of which machine halts on which string, with D as the flipped

diagonal.

  • HALT is recursively enumerable, by simulation; its complement is not, by chapter 70's corollary.
  • The theorem is about a general method, not about any particular program, and it says nothing about difficulty.
  • Halting on the empty input is also undecidable, by a reduction FROM the general problem.

Test yourself

1. State the halting problem as a language. The set of strings consisting of the encoding of a Turing machine M followed by a string w, such that M halts when run on w.

2. Give the proof in full. Suppose a machine H decides it, always halting and always right. Build D which, on input s, runs H on the pair (s, s) and then loops for ever if H said yes and halts if H said no. Then D halts on s exactly when the machine coded by s does not halt on s. Run D on its own code d: D halts on d exactly when D does not halt on d, which is a contradiction. So H does not exist.

munotes.in372

The Halting Problem

3. Which earlier chapter's argument is this, and what corresponds to the diagonal? Chapter 7's diagonalisation. The table has a row per machine and a column per string, saying whether that machine halts on that string, and D is built to disagree with every row at its own diagonal entry.

4. Is HALT recursively enumerable? Is its complement? HALT is, by simulating M on w and saying yes when it halts. Its complement is not: if it were, HALT would be recursive by chapter 70's theorem, and it is not.

5. A student says the theorem means we can never tell whether a program will finish. Correct them. For many particular programs it is obvious. The theorem says no single method decides it for every program, and useful tools that answer for many programs and give up on the rest are perfectly possible and exist.

6. Why is the machine D not a trick or an illegal object? Because it is built by ordinary means from H: copying an input, running H, and branching to a loop or a halt. If H exists then D exists, so the impossibility falls on H.

Contents This chapter on its own page

munotes.in373

Chapter Seventy-Five

Reduction: Proving a Second Problem Unsolvable

Syllabus topic Module 2, "Turing Machines: Introduction to Unsolvable Problems"

In one line

To prove a new problem unsolvable, show that solving it would solve one already known to be unsolvable.

In the wording a student can write in an examination: a problem A reduces to a problem B if there is an algorithm which converts every instance of A into an instance of B with the same answer. If A reduces to B then a decision procedure for B yields one for A; equivalently, if A is undecidable then so is B. A reduction therefore carries decidability forwards and undecidability backwards.

The definition

A reduces to B when there is a computable function f taking each instance x of A to an instance f(x) of B such that

the answer to x in A is YES exactly when the answer to f(x) in B is YES

Three conditions, all necessary.

f must be computable, by a machine that always halts. A conversion that cannot be carried out proves nothing.

f must preserve the answer in BOTH directions. Yes must go to yes and no must go to no. A function that only sends yes to yes is useless: it would allow a wrong answer on the no instances.

f must work for every instance, not for a convenient family.

The two uses, and the direction

This is the section the marks are in.

Use one, forwards: decidability. If A reduces to B and B is decidable, then A is decidable. Given an instance x of A, compute f(x) and run B's decider on it. The answer is the answer.

Use two, backwards: undecidability. If A reduces to B and A is undecidable, then B is undecidable. For if B had a decider, the procedure above would decide A, and it has none.

A reduces to B: decidability flows from B to A, undecidability flows from A to B

So to prove B undecidable, reduce a KNOWN undecidable problem TO it. Not the other way round.

The mnemonic that works. B is at least as hard as A. So put the hard thing on the left of "reduces to" and the thing you are attacking on the right.

The error, spelled out. A student proves B undecidable by reducing B to the halting problem. That shows only that B is no harder than halting, which every problem in this book already is. It proves nothing at all, and it is the commonest wrong answer on this topic.

The shape of every reduction proof

Four parts, and an answer with all four is complete.

  1. Name the known undecidable problem, A. Usually halting.
  2. Give the construction: from an instance x of A, describe the instance f(x) of B.
  3. Prove the equivalence, both ways: if x is a yes then f(x) is a yes, and if f(x) is a yes then x is a yes.
  4. Conclude: a decider for B would decide A, which is impossible, so B is undecidable.
munotes.in374

Reduction: Proving a Second Problem Unsolvable

Part 3 is where the work is and it is the part that gets skipped. Both directions must be argued.

Reduction 1: halting on the empty input

The problem. Given a machine M, does M halt when started on a blank tape?

Reduce from halting. Take an instance of halting: a machine M and a string w.

The construction. Build a machine M prime which, on any input:

  1. Ignores whatever is on its tape and erases it.
  2. Writes w on the tape.
  3. Runs M from the start.

M prime is computable from M and w. Writing w is a fixed sequence of moves, one per symbol of w, so M prime's description can be produced mechanically. That is part 2 done.

The equivalence, both ways.

If M halts on w, then M prime, started on a blank tape, writes w and then runs M on w, which halts. So M prime halts on the blank tape.

If M prime halts on the blank tape, then since everything it does after step 2 is M running on w, M must halt on w.

Conclusion. A decider for the blank tape problem would decide halting, by building M prime and asking. So the blank tape problem is undecidable.

Reduction 2: does a machine accept a particular string

The problem. Given M and w, does M accept w? This is the universal language of chapter 73.

Reduce from halting. Take M and w.

The construction. Build M prime that behaves exactly like M except that every halting configuration is redirected to the accepting state. So M prime accepts whenever M halts, for any reason.

The equivalence. M halts on w exactly when M prime accepts w: if M halts, M prime reaches the accepting state instead; if M prime accepts w, it did so at a point where M would have halted.

Conclusion. Undecidable.

Note the reverse also holds, so the two problems reduce to each other and are of the same difficulty. Either may be used as the known problem in a later reduction, and textbooks differ about which they start from.

Reduction 3: does a machine accept the empty language

The problem. Given M, is T(M) empty?

Reduce from the acceptance problem. Take M and w.

The construction. Build M prime which, on input x:

  1. Ignores x entirely.
  2. Runs M on w.
  3. If M accepts w, accept x.

What M prime's language is. If M accepts w, then M prime accepts every x, so its language is everything. If M does not accept w, then M prime accepts nothing, so its language is empty.

munotes.in375

Reduction: Proving a Second Problem Unsolvable

The equivalence. T(M prime) is empty exactly when M does not accept w.

Conclusion. A decider for emptiness would tell us whether M accepts w, which is undecidable. So emptiness is undecidable.

This construction is worth memorising. "Build a machine that ignores its input and runs M on w" is the single most useful gadget in this block, and chapter 76 generalises it into a theorem.

Reduction 4: do two machines accept the same language

The problem. Given M1 and M2, is T(M1) equal to T(M2)?

Reduce from emptiness, which reduction 3 just settled.

The construction. Given M, let M1 be M and let M2 be a fixed machine that accepts nothing, which is one state and no moves.

The equivalence. T(M1) equals T(M2) exactly when T(M) is empty.

Conclusion. Undecidable.

Notice how short that was. Once a stock of undecidable problems exists, new ones cost three lines. That is the compounding chapter 74 made possible and it is why this technique matters.

The stock of undecidable problems, so far

ProblemProved byIn
does M halt on wdiagonalisationchapter 74
does M halt on the blank tapereduction from haltinghere
does M accept wreduction from haltinghere
is T(M) emptyreduction from acceptancehere
do M1 and M2 agreereduction from emptinesshere
is T(M) finite, regular, context freeRice's theoremchapter 76
does a Post correspondence instance have a solutionPost's own proofchapter 77
is a context free grammar ambiguousreduction from Postchapter 77

One diagonal argument and everything else by reduction, which is the pattern of the whole subject.

Distinctions

A reduces to BB reduces to A
SaysB is at least as hard as AA is at least as hard as B
To prove B undecidable, you wantthis onenot this one
Gives a decider forA, if B has oneB, if A has one
What f must doWhy
be computable, always haltingor the conversion cannot be performed
send yes to yesotherwise a yes instance is missed
send no to nootherwise a no instance is answered wrongly
work on every instancea method must be general

What it does NOT mean

A reduction is not a proof that two problems are the same. It is one directional, and A reducing to B says nothing about B reducing to A unless that is shown separately.

The converted instance need not resemble the original. In reduction 3 an instance of acceptance becomes an instance of emptiness, which is a question about a different machine.

munotes.in376

Reduction: Proving a Second Problem Unsolvable

f need not be efficient. For undecidability any always halting f will do. Chapter 83 introduces a stricter kind of reduction, where f must run in polynomial time, and the distinction matters there and not here.

A reduction does not solve anything. It transfers a question, and the transfer is what carries the impossibility.

Reducing to the halting problem proves nothing. Every problem in this book reduces to it.

Quick revision

  • A reduces to B when a computable, always halting f turns every instance of A into an instance of B with the

same answer, in both directions.

  • Decidability flows from B to A; undecidability flows from A to B.
  • To prove B undecidable, reduce a KNOWN undecidable problem TO B. B is then at least as hard.
  • Four parts to the proof: name the known problem, give the construction, prove the equivalence both ways, and

conclude.

  • The stock gadget: build a machine that ignores its own input and runs M on w, so that its language is either

everything or nothing according to whether M accepts w.

  • One diagonal argument, in chapter 74, and everything after it by reduction.

Test yourself

1. Define a reduction and give the three conditions on f. A reduces to B when there is a function f from instances of A to instances of B with the answer preserved. f must be computable by an always halting machine, must preserve the answer in both directions, and must be defined on every instance.

2. Which way does a reduction carry undecidability, and what is the standard error? From A to B: if A is undecidable and A reduces to B then B is. The standard error is reducing the new problem to the halting problem, which shows only that it is no harder than halting and proves nothing.

3. Prove the blank tape problem undecidable. From M and w build M prime that erases its tape, writes w and runs M. M prime halts on a blank tape exactly when M halts on w, and M prime is computable from M and w, so a decider for the blank tape problem would decide halting.

4. Describe the gadget used to prove emptiness undecidable. A machine that ignores its own input entirely, runs M on w, and accepts its input if M accepts w. Its language is everything if M accepts w and empty otherwise, so emptiness of its language answers a question about M and w.

5. Prove that equivalence of two machines is undecidable, in three lines. Given M, take M1 as M and M2 as a machine with no moves, whose language is empty. The two agree exactly when T(M) is empty, which is undecidable, so equivalence is.

munotes.in377

Reduction: Proving a Second Problem Unsolvable

6. Why must the equivalence in a reduction be proved in both directions? Because a function sending only yes instances to yes instances would let a no instance be converted to a yes one, and the decider for B would then give the wrong answer for A.

Contents This chapter on its own page

munotes.in378

Chapter Seventy-Six

Rice's Theorem

Syllabus topic Module 2, "Turing Machines: Introduction to Unsolvable Problems"

In one line

Every non trivial question about what a machine's language IS, rather than about how the machine is written, is undecidable.

In the wording a student can write in an examination: let P be a property of the recursively enumerable languages, that is, a set of such languages. P is trivial if it holds of all of them or of none. Rice's theorem states that for every non trivial P, the problem of deciding whether T(M) has the property P is undecidable.

What the theorem is about, and what it is not

The distinction chapter 71 drew, made into a theorem.

A property of the LANGUAGE. Is T(M) empty? Is it finite? Is it regular? Does it contain the string ab? Each of these is a question about the set of strings the machine accepts, and two machines with the same language must get the same answer.

A property of the MACHINE. How many states has M? Does M ever write a blank? Does M halt within a hundred steps on the empty input? Each of these can differ between two machines that accept the same language.

Rice's theorem is about the first kind only. The second kind is often decidable and sometimes trivially so, and the theorem says nothing about it.

The test to apply. Ask: if I replaced M by a completely different machine accepting exactly the same language, would the answer have to be the same? If yes, it is a property of the language and Rice's theorem applies. If no, it is not, and Rice's theorem is silent.

Trivial and non trivial

A property is trivial when every recursively enumerable language has it, or none does.

PropertyTrivial?
T(M) is recursively enumerabletrivial: every one is
T(M) is a set of stringstrivial: every one is
T(M) is uncountabletrivial: none is
T(M) is emptynot trivial: some are, some are not
T(M) is finitenot trivial
T(M) is regularnot trivial
T(M) contains abnot trivial

The trivial ones are decidable, and obviously: the answer is always yes, or always no, and a machine that prints it ignores its input. That is why they are excluded, and it is the only exclusion the theorem needs.

The statement

Rice's theorem. Let P be any non trivial property of the recursively enumerable languages. Then

{ the code of M : T(M) has the property P } is undecidable

Rice's own paper puts it in the vocabulary of its time. Its introduction says the most interesting result of the paper, and the one that gives it its name, "is the fact that no nontrivial class is completely recursive", where a class is a set of recursively enumerable sets and completely recursive is his term for decidable, which he glosses in the same sentence.

munotes.in379

Rice's Theorem

The paper is "Classes of Recursively Enumerable Sets and Their Decision Problems", Transactions of the American Mathematical Society volume 74 (1953), pages 358 to 366.

The proof

By reduction from the acceptance problem of chapter 75, using that chapter's stock gadget.

Setting up. Let P be non trivial. Write empty for the empty language.

Assume the empty language does NOT have the property P. If it does, run the whole proof on the negation of P instead, which is also non trivial, and a decider for one gives a decider for the other by swapping the answers. So this assumption costs nothing.

Since P is non trivial, some recursively enumerable language does have it. Call it L, and let ML be a machine accepting L. L is not empty by the assumption above.

The reduction. Given an instance of acceptance, a machine M and a string w, build a machine M prime which, on input x:

  1. Save x.
  2. Run M on w. If M does not accept w, this never finishes.
  3. If M accepts w, run ML on the saved x, and accept x if ML does.

What T(M prime) is, and this is the whole proof.

If M accepts w, step 2 finishes for every x, so M prime accepts exactly the x that ML accepts, and T(M prime) is L, which has the property P.

If M does not accept w, step 2 never finishes, so M prime accepts nothing, and T(M prime) is the empty language, which does not have the property P.

So T(M prime) has the property P exactly when M accepts w.

The conclusion. M prime is computable from M and w. A decider for P applied to M prime would tell us whether M accepts w, which chapter 75 proved undecidable. So no decider for P exists.

Note where the non triviality was used, twice: to get a language L that has the property, and, through the opening manoeuvre, to make the empty language one that does not.

What follows at once

Every one of these is undecidable, and each is one line of Rice's theorem rather than a reduction of its own.

Question about T(M)Non trivial?Hence
is it emptyyesundecidable
is it finiteyesundecidable
is it infiniteyesundecidable
is it regularyesundecidable
is it context freeyesundecidable
is it recursiveyesundecidable
does it contain a particular stringyesundecidable
does it equal a fixed languageyesundecidable
do two machines have the same languagenot directlysee below

The last row needs a word. Equality of two machines is a property of a pair, not of one language, so Rice's theorem does not apply as stated. Chapter 75 proved it undecidable by a three line reduction instead.

munotes.in380

Rice's Theorem

And chapter 75's first four results are now redundant, except as practice: emptiness follows from Rice, and so would finiteness and regularity had they been asked. That is the point of a general theorem.

What Rice's theorem does NOT cover

Five kinds of question, and each is worth recognising because an examination may set one to see whether the theorem is being applied blindly.

Properties of the machine's description. How many states, how many tape symbols, whether a particular state is reachable in the transition graph. All decidable by inspection.

Properties involving a step bound. Does M halt on w within 100 steps? Decidable: simulate 100 steps.

Trivial properties. Is T(M) recursively enumerable? Always yes.

Properties of pairs. Do two machines agree? Not a property of one language.

Questions about a machine of a restricted kind. Is a given FINITE AUTOMATON's language empty? Perfectly decidable, by chapter 39. Rice's theorem is about Turing machines, and the restricted models have their own answers.

That last exclusion is the one most often missed. The theorem is not "all questions about languages are undecidable"; it is about the languages of Turing machines, presented as Turing machines.

The moral

Chapter 71 observed that questions about a machine's description tend to be decidable and questions about its behaviour tend not to be. Rice's theorem is that observation proved, for the sharpest version of behaviour: what the machine accepts.

So there is no point looking for a clever decider for a question about a machine's language. The theorem closes the whole family at once, and a student who recognises a question as falling under it has answered it.

Distinctions

A property of the languageof the machine
Two machines with the same languagemust get the same answermay differ
Exampleis T(M) finitehow many states has M
Rice's theorem appliesyes, if non trivialno
Usuallyundecidabledecidable
TrivialNon trivial
Holds ofall languages, or nonesome but not all
Decidableyes, print the constant answerno, by Rice
Exampleis T(M) recursively enumerableis T(M) empty

What it does NOT mean

It does not say all questions about machines are undecidable. Only non trivial questions about their languages.

It does not apply to finite automata or pushdown automata. Chapter 39 and chapter 51 decide plenty about those. The theorem concerns Turing machines.

It does not say the property itself is undecidable. "Is this language finite" may be perfectly easy when the language is given some other way. It is deciding it FROM A MACHINE that is impossible.

munotes.in381

Rice's Theorem

Trivial does not mean uninteresting. It means holding of all or of none, which is a technical condition and the only one excluded.

It gives no information about difficulty. Everything it covers is equally impossible; there are no degrees in its statement.

Quick revision

  • A property of the recursively enumerable languages is trivial when every one has it or none does.
  • Rice's theorem: for every NON trivial such property, deciding whether T(M) has it is undecidable.
  • The test: would a different machine with the same language have to get the same answer? If yes it is a

property of the language and the theorem applies.

  • Proof: by reduction from acceptance, using a machine that saves its input, runs M on w, and only then runs a

machine for a language that has the property. Its language is that one if M accepts w and empty otherwise.

  • Consequences in one line each: emptiness, finiteness, infinity, regularity, context freeness, recursiveness,

containing a given string.

  • Not covered: the machine's description, step bounded questions, trivial properties, properties of pairs, and

questions about restricted models.

  • Rice 1953, Transactions of the American Mathematical Society volume 74, pages 358 to 366; his own phrase is

that no nontrivial class is completely recursive.

Test yourself

1. State Rice's theorem. For every non trivial property of the recursively enumerable languages, the problem of deciding whether a given Turing machine's language has that property is undecidable.

2. What makes a property trivial, and why are the trivial ones excluded? It holds of every recursively enumerable language or of none. Those are decidable outright, by a machine that ignores its input and prints the constant answer, so they are the one exclusion the theorem needs.

3. Give the test for whether the theorem applies to a question. Ask whether a completely different machine accepting the same language would have to get the same answer. If yes it is a property of the language; if no it is a property of the machine and the theorem is silent.

4. Which of these does Rice's theorem settle: is T(M) infinite; has M seven states; does M halt on the blank tape within fifty steps; is T(M) equal to T(M2)? Only the first. The second is a property of the description, the third has a step bound and is decidable by simulation, and the fourth is a property of a pair rather than of one language.

5. Describe the machine the proof builds. One that saves its own input, runs M on w and gets stuck for ever if M does not accept, and otherwise runs a machine for a language known to have the property, accepting whatever that machine accepts. So its language is that one if M accepts w and empty otherwise.

munotes.in382

Rice's Theorem

6. Does Rice's theorem mean emptiness is undecidable for a finite automaton? No. It concerns Turing machines, and emptiness for a finite automaton is decided by a reachability search, as chapter 39 showed.

Contents This chapter on its own page

munotes.in383

Chapter Seventy-Seven

The Post Correspondence Problem, and the Undecidable Problems of Grammars

Syllabus topic Module 2, "Turing Machines: Introduction to Unsolvable Problems"

In one line

Given a list of pairs of strings, can some sequence of them be chosen so that the tops and the bottoms spell the same thing? No method decides it.

In the wording a student can write in an examination: an instance of the Post correspondence problem is a finite list of pairs of non empty strings over an alphabet. A solution is a non empty sequence of indices, with repetition allowed, such that concatenating the first components in that order gives the same string as concatenating the second components. The problem of determining whether an arbitrary instance has a solution is undecidable.

The problem, as Post posed it

Post's 1946 paper is "A Variant of a Recursively Unsolvable Problem", Bulletin of the American Mathematical Society volume 52 (1946), pages 264 to 268. It opens by defining strings over two letters and then poses what it calls the correspondence decision problem: whether, for an arbitrary finite set of pairs of corresponding non null strings, there is a solution matching the two concatenations. The paper then proves that in its full generality the problem is recursively unsolvable.

Why it looks like a puzzle and is not. It has the shape of a game with dominoes and it is undecidable, which is why it is worth including: undecidability is not confined to questions about machines. And once a problem is undecidable it can be used to prove others undecidable, and Post's is the one that reaches the grammars.

An instance, and a solution

An instance is a list of pairs. Write them as a table, tops in one row and bottoms in the other.

Post's own example, from his paper.

Pair123
topbbabb
bottombbabb

Is there a solution? Take the sequence of indices 1, 2, 2, 3.

Pieces chosenJoined
topsbb, ab, ab, bbbababb
bottomsb, ba, ba, bbbbababb

The two agree, both being bbababb, seven symbols. So this instance has a solution, and Post's paper gives exactly this one.

Notice the asymmetry of the evidence. Producing the sequence 1, 2, 2, 3 settles the yes case completely, in one line, and anybody can check it. Settling a no case needs an argument about every possible sequence, of which there are infinitely many. That asymmetry is the shape of undecidability, and it is the same shape chapter 44 noted for ambiguity.

Two cases where there is no solution, and both are Post's

His paper gives two sufficient conditions, and each is worth knowing because an examination may set an instance satisfying one.

If every top is longer than its bottom, there is no solution: the joined tops would be longer than the joined bottoms, so they cannot be equal.

munotes.in384

The Post Correspondence Problem, and the Undecidable Problems of Grammars

If every top begins with a different letter from its bottom, there is no solution: the first pair chosen would make the two strings differ at their very first symbol.

Pair12
topabb
bottomba

Here the first pair has a longer top than bottom, and the second begins differently from its bottom, so a sequence starting with either pair is already lost.

These conditions are sufficient and not necessary. An instance failing both tests may still have no solution, and deciding that in general is exactly what cannot be done.

The modified problem

For the reductions it is convenient to require the solution to begin with pair 1. That is the modified Post correspondence problem.

It is also undecidable, and the two are interreducible: the modified problem is a special case, and the general one reduces to it by padding the strings with a marker symbol so that only one pair can begin a solution.

The reductions below use the modified form, which is standard.

Why it is undecidable

The full proof is a reduction from the halting problem and it is long. What it does is worth knowing even though MU does not set the details.

The idea. Given a machine M and an input w, build an instance whose pairs are made from M's moves, arranged so that the only way to make the tops and the bottoms agree is to spell out an accepting computation of M on w. The bottom row runs one configuration ahead of the top row, and each pair either copies a symbol across unchanged or applies one move of M, so a matching sequence is exactly a valid run.

Then a solution exists exactly when M accepts w, and a decider for the correspondence problem would decide acceptance, which chapter 75 proved impossible.

What to take from it. It is the same trick as chapter 59's triple: encode a computation as a combinatorial object, so that a question about the object becomes a question about the computation.

What it proves about grammars

This is why the chapter is in the syllabus, and it closes chapter 51's table.

The construction. From an instance with pairs (x1, y1) to (xn, yn), and fresh symbols a1 to an not in the alphabet, build two grammars:

Gx: S -> x(i) S a(i) | x(i) a(i), one pair of productions for each i

Gy: S -> y(i) S a(i) | y(i) a(i), one pair of productions for each i

What they generate. Gx generates every string made of some tops joined, followed by the indices used, in reverse order. Gy does the same with the bottoms. So a string lies in both languages exactly when one index sequence gives the same joined string on both rows, which is exactly a solution.

munotes.in385

The Post Correspondence Problem, and the Undecidable Problems of Grammars

L(Gx) intersect L(Gy) is non empty exactly when the instance has a solution

Four results fall out, each one line from that.

Is the intersection of two context free languages empty? Undecidable, immediately.

Is a context free grammar ambiguous? Undecidable. Combine Gx and Gy into one grammar with a new start symbol S and the two productions S to Sx and S to Sy, renaming the variables apart. A string with two derivation trees is one derivable both ways, which is one in the intersection. So the combined grammar is ambiguous exactly when the instance has a solution.

Are two context free grammars equivalent? Undecidable. The route is through complements: the complement of L(Gx) is context free, and so is the complement of L(Gy), and so is their union, and that union is everything exactly when the intersection of the two original languages is empty.

Is a context free language equal to Sigma star? Undecidable, by the same union.

And the ones that remain decidable, from chapter 51, for contrast: membership, emptiness and finiteness for a single grammar. The line is worth stating: a question about ONE grammar's strings tends to be decidable, and a question comparing TWO tends not to be.

A worked construction, small enough to read

Take the two pair instance with tops a and ab, bottoms ab and b, and fresh symbols 1 and 2.

S->a S 1 | a 1

Accepts: a1, aa11, aaa111

Rejects: ε, a, 1, 1a, aa1

That is Gx for the first pair alone, written out so the shape is visible: the top string, then the rest, then the index. A full Gx has two productions per pair and generates all the sequences.

Read the shape. Every string it generates is some tops joined, then the indices of the pairs used in reverse order. The reversal is what makes the grammar context free rather than needing two counts at once: the indices come back off in the order a stack would give them, which is chapter 52's observation again.

The whole undecidability picture

ProblemStatusProved by
haltingundecidablediagonalisation, chapter 74
acceptance, blank tape, emptiness of T(M)undecidablereduction, chapter 75
every non trivial property of T(M)undecidableRice, chapter 76
Post correspondenceundecidablereduction from halting
emptiness of an intersection of two CFLsundecidablereduction from Post
ambiguity of a CFGundecidablereduction from Post
equivalence of two CFGsundecidablereduction from Post
membership, emptiness, finiteness for one CFGDECIDABLEchapter 51
everything in chapter 39 about finite automataDECIDABLEchapter 39
the Entscheidungsproblem, whether a statement of first order logic is provableundecidableTuring, 1936
munotes.in386

The Post Correspondence Problem, and the Undecidable Problems of Grammars

The last row is where the subject began. Hilbert asked for a definite method deciding any statement of first order logic, Turing's paper of chapter 62 invented the machine in order to answer him, and the answer was no. So a first order theory can be undecidable, and the one Hilbert asked about is.

One diagonal argument at the top and everything else by reduction, which is the sentence that summarises the whole block.

Distinctions

The general problemThe modified problem
The solutionany non empty index sequencemust begin with pair 1
Undecidableyesyes
Used in reductionsrarelyusually
A yes instanceA no instance
Evidencethe index sequence, checkable in a linean argument about infinitely many sequences
Semi decidableyes, by searching sequences in order of lengthno

What it does NOT mean

It is not a puzzle with a trick. Particular instances are often easy either way; no method handles all of them.

A no answer is not always hard to see. Post's two conditions settle many instances at a glance, and they are sufficient rather than necessary.

The problem is not about machines. That is what makes it useful: it shows undecidability is not confined to self referential questions about programs, and it gives a combinatorial starting point for reductions into grammars.

The grammar results are not about Turing machines. They are about context free grammars, which chapter 51 showed are otherwise well behaved. The undecidability arrives through Post's problem rather than directly.

Undecidable does not mean useless. Parser generators decide ambiguity for restricted grammar classes and reject the rest, which is how the practical world lives with this result.

Quick revision

  • An instance is a finite list of pairs of non empty strings; a solution is a non empty index sequence, repeats

allowed, making the joined tops equal the joined bottoms.

  • Post 1946, Bulletin of the American Mathematical Society volume 52, pages 264 to 268.
  • His own example: tops bb, ab, b and bottoms b, ba, bb, solved by 1, 2, 2, 3, both sides giving bbababb.
  • Two sufficient conditions for no solution, both his: every top longer than its bottom, or every top beginning

with a different letter from its bottom.

  • The modified problem requires the solution to begin with pair 1, is also undecidable, and is the one used in

reductions.

  • Undecidable by a reduction from halting, in which a matching sequence is forced to spell out an accepting

computation.

  • From it: emptiness of an intersection of two context free languages, ambiguity of a grammar, equivalence of two
munotes.in387

The Post Correspondence Problem, and the Undecidable Problems of Grammars

grammars, and whether a context free language is everything.

  • A question about one grammar's strings tends to be decidable; one comparing two tends not to be.

Test yourself

1. Define an instance and a solution. An instance is a finite list of pairs of non empty strings. A solution is a non empty sequence of indices, with repetition allowed, such that joining the first components in that order gives the same string as joining the second components.

2. Solve the instance with tops bb, ab, b and bottoms b, ba, bb. The sequence 1, 2, 2, 3 works: the tops give bb, ab, ab, b which is bbababb, and the bottoms give b, ba, ba, bb which is also bbababb.

3. Give two conditions under which an instance certainly has no solution. If every top is longer than its corresponding bottom, since the joined tops would then be longer; or if every top begins with a different letter from its bottom, since the very first pair chosen would make the two strings differ.

4. What are the two grammars built from an instance, and what does their intersection mean? One generating the joined tops followed by the reversed index sequence, and one doing the same with the bottoms. A string lies in both exactly when one index sequence gives the same result on both rows, so the intersection is non empty exactly when the instance has a solution.

5. Name three questions about context free grammars that this makes undecidable. Whether the intersection of two context free languages is empty; whether a grammar is ambiguous; and whether two grammars generate the same language.

6. Why is a yes instance easy to confirm and a no instance not? A yes instance has a finite witness, the index sequence, which anybody can check in a line. A no instance is a claim about infinitely many sequences and has no finite witness, so the problem is semi decidable and not decidable.

Contents This chapter on its own page

munotes.in388

Chapter Seventy-Eight

Time Complexity

Syllabus topic Module 2, "Computability and Complexity: Time Complexity and Space Complexity"

In one line

The time a machine takes is the number of moves it makes, measured against the length of its input, in the worst case.

In the wording a student can write in an examination: let M be a Turing machine that halts on every input. The time complexity of M is the function T(n) giving the maximum number of moves M makes on any input of length n. A language has time complexity T(n) if some machine deciding it has that time complexity, and the class of languages decidable within a time bound is written with that bound.

Why steps and not seconds

Because seconds depend on the machine it is run on and steps do not.

A count of moves is a property of the algorithm, and it is the same whoever runs it. That is the point Hartmanis and Stearns make in the opening of their paper: they choose the multitape Turing machine as the model, they say, because "all digital computers in a slightly idealized form belong to the class of multitape Turing machines", so a result about moves is a result about computers.

And the count is the right thing even for a real program, because hardware gets faster by constant factors and algorithms differ by factors that grow. A machine twice as fast halves the seconds and does not change the function.

The three choices in the definition

Each has to be made and each is made the same way by convention.

Measured against what? The length of the input, written n. Not the input itself: there are infinitely many inputs and the function must be one function.

Which input of that length? The worst one. So T(n) is a maximum, and a machine is as slow as its worst case.

Counting what? Moves, each costing one, whatever it does. Reading, writing and moving the head are one move together.

Why the worst case and not the average. Because an average needs a distribution over the inputs, and there is no natural one. Average case analysis is a real subject and it needs an assumption the worst case does not.

The origin of the term

Hartmanis and Stearns introduced it. Their paper is "On the Computational Complexity of Algorithms", Transactions of the American Mathematical Society volume 117 (1965), beginning at page 285, and it records that it was received by the editors on 2 April 1963 and in revised form on 30 August 1963.

Their opening sets out the problem the block solves, and it is worth quoting because it is the whole motivation in one sentence:

some computable sequences are very easy to compute whereas other computable sequences seem to have an inherent

complexity that makes them difficult to compute

munotes.in389

Time Complexity

That is the observation Module 2 has not yet addressed. Chapters 62 to 77 divided the computable from the uncomputable. This block divides the computable into the easy and the hard.

The speed up theorem, and why constants are dropped

Their paper proves a result that decides how complexity is stated for ever after. In their words, there is a speed up theorem "which states that ST = SkT for positive numbers k": the class of sequences computable within T(n) steps is the same as the class computable within k times T(n) steps, for any positive constant k.

So a constant factor is not a difference. A machine running in 100n steps and one running in n steps decide the same languages within the same class, and no statement in this subject may depend on the difference.

How the speed up is achieved. By enlarging the tape alphabet so that one new symbol packs several old ones, and then processing several old cells in one move. That is chapter 68's multitrack observation used for speed rather than for space.

The consequence for notation. Since constants do not matter, the notation used must ignore them, which is exactly what chapter 80's big O does. The two chapters fit together: the theorem says constants are meaningless and the notation is built so that they cannot be expressed.

Measured examples from this book

Every one of these was counted by running the machine, not estimated.

MachineChapterInputMoves
0 to the n 1 to the n62001113
palindromes65abba15
a to the n b to the n c to the n66abc8
unary addition67110117
binary successor6710118
unary doubling67110
unary doubling671124
unary doubling6711144

The doubling machine's three rows are the interesting ones, because they are the same machine on inputs of 1, 2 and 3 symbols. The moves go 10, 24, 44, and the differences are 14 and 20, so the second difference is 6: a quadratic. That is the n squared chapter 67 predicted, and here it is measured rather than asserted.

Compare the first row with a pushdown automaton. Chapter 52's machine reads 0011 in four moves. The Turing machine takes thirteen, because it crosses the tape repeatedly. Same language, same answer, three times the work.

The standard growth rates

The functions that actually turn up, in increasing order.

T(n)CalledA machine doing this
constantconstant timelooking at one cell
log nlogarithmicbinary search, on a suitable model
nlinearone pass over the input
n log nlinearithmicthe best comparison sorts
n squaredquadraticone pass per symbol, as chapter 67's copier
n cubedcubicCYK, chapter 51
2 to the nexponentialtrying every subset
n factorialfactorialtrying every ordering
munotes.in390

Time Complexity

The line that matters is drawn between n to some power and 2 to the n, and chapter 81 draws it: everything above the line is called tractable and everything below is not.

Why that line and not another is chapter 81's argument, and it is not obvious; the short version is that polynomials compose and close under the operations an algorithm does, and exponentials swamp any hardware improvement.

What the model costs

A complexity statement is only meaningful once the model is fixed, because the model changes the answer.

Change of modelEffect on time
a constant factor faster hardwarenone, by the speed up theorem
one tape instead of kup to a square, chapter 68
deterministic instead of nondeterministicup to an exponential, chapter 69
a random access machine instead of a tapeup to a polynomial

The first and last rows are why the class P of chapter 81 is robust, and the third row is why chapter 86's question is open. A student who knows which changes cost a polynomial and which cost an exponential has the shape of the whole block.

Distinctions

ComputabilityComplexity
Askscan it be done at allwhat does it cost
Answered inchapters 62 to 77chapters 78 to 86
A negative result meansno algorithm existsno fast algorithm is known, or exists
Depends on the modelno, by chapter 72yes, up to a polynomial
Worst caseAverage case
T(n) isthe maximum over inputs of length nthe mean
Needsnothing extraa distribution over inputs
Used herealwaysnot at all

What it does NOT mean

Time is not seconds. It is moves, and a faster computer changes the seconds and not the count.

A constant factor is not a difference. The speed up theorem says so, and every statement in this block is written to be insensitive to it.

Time complexity is not defined for a machine that may loop. The definition requires the machine to halt on every input, which is why this block is about decidable problems only.

A measured count on one input is not T(n). T(n) is the maximum over all inputs of that length, and the tables above give data points rather than the function.

Faster on the model is not always faster in practice. The model fixes what is counted, and chapter 80's notation drops exactly the factors a practical comparison may care about.

Quick revision

  • T(n) is the maximum number of moves on any input of length n, for a machine that halts on every input.
  • Moves rather than seconds, because seconds depend on the hardware and moves do not.
  • Worst case, because an average needs a distribution and there is no natural one.
  • Hartmanis and Stearns introduced the term in 1965, and their speed up theorem says a constant factor changes
munotes.in391

Time Complexity

nothing, which is why the notation of chapter 80 ignores constants.

  • The speed up is achieved by packing several old cells into one new symbol, which is chapter 68's multitrack

idea used for time.

  • The doubling machine of chapter 67, measured at 10, 24 and 44 moves for inputs of 1, 2 and 3 symbols, is a

quadratic, which is what a copy on one tape costs.

  • The model matters up to a polynomial, except for nondeterminism, which costs an exponential and is chapter

86's open question.

Test yourself

1. Define the time complexity of a machine, naming the three choices in the definition. The maximum number of moves the machine makes on any input of length n. The three choices are: measured against the length of the input, taken over the worst input of that length, and counting moves rather than time.

2. Why is time counted in moves? Because a count of moves is a property of the algorithm and is the same on any hardware, while seconds depend on the machine it is run on. Hardware improves by constant factors, which the speed up theorem says do not matter.

3. State the speed up theorem and say what follows for notation. The class of sequences computable within T(n) steps equals the class computable within k times T(n) steps for any positive constant k. It follows that no statement may depend on a constant factor, so the notation used must be unable to express one, which is what big O does.

4. The doubling machine takes 10, 24 and 44 moves on inputs of 1, 2 and 3 symbols. What growth is that? Quadratic. The first differences are 14 and 20 and the second difference is 6, which is constant, so the moves grow with the square of the input.

5. Which changes of model cost a polynomial, and which an exponential? Reducing many tapes to one costs a square, and moving between a tape machine and a random access machine costs a polynomial. Removing nondeterminism costs an exponential, and whether that can be reduced is chapter 86's question.

6. Why is time complexity only defined for a machine that halts on every input? Because otherwise there is no number of moves to take a maximum of. The whole block is therefore about decidable problems, and undecidable ones have no complexity at all.

Contents This chapter on its own page

munotes.in392

Chapter Seventy-Nine

Space Complexity

Syllabus topic Module 2, "Computability and Complexity: Time Complexity and Space Complexity"

In one line

The space a machine uses is the number of tape cells it visits, and unlike time it can be used again.

In the wording a student can write in an examination: the space complexity of a Turing machine that halts on every input is the function S(n) giving the maximum number of tape cells visited on any input of length n. For sublinear bounds the offline machine is used, whose read only input tape is not counted, so that S(n) measures only the work tapes.

The one asymmetry, and everything follows from it

A cell can be written again. A moment cannot be lived again.

So a machine may use the same ten cells a million times, and its space is ten while its time is a million. The reverse cannot happen: a machine cannot use a thousand cells in ten moves, because it takes a move to reach each new cell.

space at most time, always, up to a constant

That one inequality is the first of this chapter's results and the easiest, and the rest of the chapter is what happens because the reverse fails so badly.

Measuring it: the offline machine

On an ordinary Turing machine the input sits on the tape, so the machine occupies n cells before it does anything, and no machine can use less space than its input. Asking about a machine that uses less than n cells would then be asking about nothing.

So sublinear space is measured on the offline machine of chapter 69: the input sits on a separate read only tape, which is not counted, and S(n) counts the work tape cells only.

That is why the offline machine is in the syllabus at all, and it is why chapter 60's linear bounded automaton has end markers: both definitions exist to make the space used by a computation a well defined quantity.

For bounds of n or more the distinction does not matter, and most of this chapter is at that level.

The space classes

BoundNameWhat lives there
constantthe regular languagesa finite automaton has no tape at all
log nL, for logarithmic spaceenough to hold a counter or a pointer into the input
nlinear spacechapter 60's linear bounded automaton, the context sensitive languages
n to a powerPSPACEa great deal, including all of P
unboundedthe recursive languages, and beyond

The second row deserves a sentence. Logarithmic space is enough to write down a position in the input, since a number up to n takes about log n digits, and not enough to write down a copy of the input. So a logarithmic space machine can point but not remember, and that turns out to be a natural and much studied class.

munotes.in393

Space Complexity

And the third row is chapter 61's, where the machine's space is the input's length and the languages are exactly the context sensitive ones.

The two results that differ from time's

Space is more powerful than time, cell for cell

Any machine running in time T uses at most T cells, as above. But a machine using S cells may run for far longer than S steps, because it may revisit them.

How long can it run? Chapter 61 counted it. With q states, a tape alphabet of size g and S cells there are q times S times g to the S configurations, and a halting machine cannot repeat one, so:

a machine using space S runs for at most about g to the S steps

So space S buys time exponential in S, and that is why PSPACE contains P and is believed to be much larger.

Nondeterminism costs much less for space than for time

This is the striking one, and the contrast with chapter 69 is the point.

For time, removing nondeterminism costs an exponential, as chapter 69 showed, and whether that can be improved is open.

For space, it costs only a square. That is Savitch's theorem: a nondeterministic machine using space S can be simulated by a deterministic one using space about S squared. So

nondeterministic space S is inside deterministic space S squared

Why the difference. A deterministic simulation of a nondeterministic machine has to explore a tree. For time that means visiting every branch, and there are exponentially many. For space it does not: the branches can be explored one at a time and the same cells reused for each, so only the depth of the recursion is paid for. The asymmetry at the top of this chapter is doing all the work.

And the consequence, from chapter 61. Nondeterministic linear space equals the context sensitive languages, and by Savitch it sits inside deterministic space n squared. Whether it equals deterministic LINEAR space is the open question chapter 60 recorded from Immerman's own paper.

And nondeterministic space is closed under complement

Chapter 61 gave this as Immerman's 1988 result. It is worth seeing here in its general form, because it is the other place where space behaves better than time: nondeterministic space is closed under complementation for any bound at least logarithmic, while whether nondeterministic TIME is closed under complement is not known and is one of the standard open questions.

Three results, all in space's favour, and all traceable to reuse.

Measured examples

Both columns below were measured by running the machines, not estimated.

munotes.in394

Space Complexity

MachineChapterInputCells visitedMoves
0 to the n 1 to the n620011513
palindromes65abba515
unary addition671101167
unary doubling671410
unary doubling6711624
unary doubling67111844

Read the last three rows. The doubling machine's space grows linearly, 4 then 6 then 8, two cells per extra input symbol, while its time grows quadratically, 10 then 24 then 44. The same machine is cheap in space and expensive in time, which is the asymmetry made visible, and it is expensive in time precisely because it keeps going back over the cells it already has.

The picture, with both measures

ClassDefined byContains
Ldeterministic log space
NLnondeterministic log spaceL, and inside L squared by Savitch
Pdeterministic polynomial TIMEL and NL
NPnondeterministic polynomial timeP
PSPACEpolynomial SPACENP, since a polynomial time machine visits polynomially many cells
EXPTIMEexponential timePSPACE, by the configuration count above

Every containment in that table is known, and almost none of them is known to be strict. The one that is, and it is the only one this syllabus needs, is that P is properly inside EXPTIME, by the hierarchy theorem of chapter 86.

Distinctions

TimeSpace
Reusablenoyes
Cost of removing nondeterminismexponential, chapter 69a square, by Savitch
Closed under complement, nondeterministicallynot knownyes, Immerman 1988
Measured onany machinethe offline machine, for sublinear bounds
The other bounds itspace is at most timetime is at most exponential in space
An ordinary machineAn offline machine
The input occupiescountable tapea read only tape, not counted
Minimum spacenas little as constant
Used forbounds of n or moresublinear bounds

What it does NOT mean

Space is not memory in bytes. It is tape cells, and the cells hold symbols from a fixed finite alphabet.

Small space does not mean fast. The doubling machine uses 8 cells and 44 moves, and a machine using S cells may run for exponentially many steps in S.

Savitch's theorem does not say nondeterminism is free for space. It says it costs a square, which for a polynomial bound is still a polynomial and for a linear bound is not linear, which is why chapter 60's question is open.

Logarithmic space is not nothing. It is enough to hold a position in the input, which is enough for a great many algorithms.

The offline machine is not a different model. Chapter 69 showed it is the ordinary machine with the input tape made read only, which costs nothing.

munotes.in395

Space Complexity

Quick revision

  • S(n) is the maximum number of tape cells visited on any input of length n, for a machine that halts on every

input.

  • Space can be reused and time cannot, and every difference between the two measures follows from that.
  • Space is at most time, always. Time is at most about g to the S for space S, by counting configurations.
  • Sublinear space needs the offline machine, whose read only input tape is not counted; that is why it exists.
  • Logarithmic space holds a position in the input but not a copy of it.
  • Removing nondeterminism costs an exponential in time and only a SQUARE in space, by Savitch's theorem, because

the branches of the tree can reuse the same cells.

  • Nondeterministic space is closed under complement, by Immerman 1988; the corresponding question for time is

open.

  • L inside NL inside P inside NP inside PSPACE inside EXPTIME, with only the outer pair known to be strict.

Test yourself

1. Define space complexity, and say why the offline machine is needed. The maximum number of tape cells visited on any input of length n. The offline machine is needed for sublinear bounds, because on an ordinary machine the input itself occupies n cells and no machine could use fewer.

2. Give the inequality between space and time, in both directions. Space is at most time, since reaching a new cell costs a move. Time is at most about g to the S for space S, since a halting machine cannot repeat a configuration and there are that many.

3. State Savitch's theorem and explain why space is cheaper than time here. A nondeterministic machine using space S can be simulated deterministically in space about S squared. It is cheaper because the branches of the computation tree can be explored one at a time reusing the same cells, while in time each branch must be paid for separately.

4. The doubling machine uses 6 cells and 24 moves on one input and 8 cells and 44 on the next. What does that show? That space and time can grow at different rates for the same machine: the space is linear and the time quadratic, because the machine keeps revisiting cells it already has.

5. What does logarithmic space suffice for, and what not? It suffices to hold a position or a counter in the input, since a number up to n needs about log n digits. It does not suffice to hold a copy of the input.

6. Name two ways in which nondeterminism behaves better for space than for time. Removing it costs only a square rather than an exponential, by Savitch; and nondeterministic space is known to be closed under complement, while the corresponding question for nondeterministic time is open.

Contents This chapter on its own page

munotes.in396

Chapter Eighty

Big O Notation, and the Family Around It

Syllabus topic Module 2, "Computability and Complexity: Big-O Notation"

In one line

Big O says one function grows no faster than another, ignoring constant factors and small inputs.

In the wording a student can write in an examination: f(n) is O(g(n)) if there exist positive constants c and n0 such that f(n) is at most c times g(n) for every n at least n0. The notation describes an upper bound on the rate of growth, with constant factors and finitely many small cases disregarded.

Why the notation has to ignore constants

Because chapter 78's speed up theorem says a constant factor is not a difference, so a notation that could express one would be saying something the subject holds to be meaningless.

And because the constants are not knowable anyway. They depend on the machine model, the tape alphabet and the encoding, none of which is fixed by the problem.

So the notation is built to be blind to exactly what the subject is blind to, and that fit is the reason it is used rather than something more precise.

The definition, with both constants

f(n) is O(g(n)) when there are constants c > 0 and n0 such that

f(n) at most c times g(n) for every n at least n0

Both constants are needed and each does a different job.

c absorbs the constant factor. Without it, 3n would not be O(n), which would make the notation useless.

n0 absorbs the small cases. Without it, n squared plus 100 would not be O(n squared), because at n equal to 1 the left side is 101 and the right is 1. The bound is about eventual behaviour, and n0 is where "eventually" begins.

The order of the quantifiers. There EXIST c and n0 such that FOR ALL n beyond n0. Choosing c after seeing n would allow anything.

Proving a bound

To show f(n) is O(g(n)), produce the two constants and verify the inequality. That is the whole method.

Worked: 3n squared plus 5n plus 2 is O(n squared).

For every n at least 1, both n and 1 are at most n squared. So

3 n squared + 5 n + 2 at most 3 n squared + 5 n squared + 2 n squared = 10 n squared

Take c equal to 10 and n0 equal to 1. Done.

The general recipe for a polynomial. Replace every lower power by the highest, add the coefficients, and that sum is c with n0 equal to 1. A polynomial of degree d is always O(n to the d), and that one line answers most examination questions on this topic.

Worked: 2 to the (n plus 1) is O(2 to the n).

2 to the (n plus 1) is 2 times 2 to the n, so c equal to 2 and n0 equal to 0 work. A constant in the exponent is a constant factor, which is worth noticing because it is the reason 2 to the n and 2 to the (n plus 1) are the same growth while 2 to the n and 3 to the n are not.

munotes.in397

Big O Notation, and the Family Around It

Disproving a bound

To show f(n) is NOT O(g(n)), the negation has to be argued: for EVERY c and n0 there is an n beyond n0 with f(n) greater than c times g(n).

Worked: n squared is not O(n).

Suppose it were, with constants c and n0. Take n to be the larger of n0 and c plus 1. Then n is at least n0, so the bound should hold, and

n squared = n times n at least (c + 1) times n > c times n

which contradicts it. So no such c exists.

The pattern. Given the supposed c, produce an n depending on c that breaks it. Both worked disproofs in this book do that, and it is the only method.

The family

Big O is one of five, and each says something different.

NotationSaysIn words
f is O(g)f at most c times g, eventuallyg is an upper bound on f's growth
f is Omega(g)f at least c times g, eventuallyg is a lower bound
f is Theta(g)both of the above, with different constantsf and g grow at the same rate
f is little o of gf over g tends to 0f grows strictly slower
f is little omega of gf over g tends to infinityf grows strictly faster

Theta is what a student usually means. Saying an algorithm is O(n squared) leaves open that it is also O(n), which would be a better statement. Saying it is Theta(n squared) says the bound is tight.

Omega is what a lower bound proof gives. "Every comparison sort takes Omega(n log n) comparisons" is a statement about every possible algorithm, which is a different kind of claim from a bound on one.

Little o is strict. n is little o of n squared; n squared is not little o of n squared, while it is O of it.

The hierarchy, with real numbers

The order of growth, and a table of what the functions actually are, so the words have sizes attached.

nlog n, to the nearest wholen log nn squared2 to the n
103331001024
204864001048576
50628225001125899906842624

Read the last column down. At n equal to 10 an exponential algorithm does about a thousand steps, at 20 about a million, and at 50 over a thousand million million. At a thousand million steps a second, the last of those is 1,125,900 seconds, which is about thirteen days.

munotes.in398

Big O Notation, and the Family Around It

Read the n squared column beside it. At n equal to 50 it is 2500, which is instant.

That gap is the whole of chapter 81's argument. A polynomial algorithm scales and an exponential one does not, and the difference is not a matter of degree.

The three things students get wrong

One. Writing equals. The convention is to write f(n) equals O(g(n)), and it is a historical abuse of the symbol. O(g(n)) is a set of functions, and the correct reading is "f belongs to O(g(n))". The abuse is harmless until somebody writes O(n) equals O(n squared) and then reverses it, which is false in one direction. This book writes "is O(g(n))" throughout for that reason.

Two. Using O when Theta is meant. Every function that is O(n) is also O(n squared) and O(2 to the n). Saying an algorithm is O(2 to the n) when it is really linear is true and useless.

Three. Applying it to one input. Big O is about a function of n, not about a running time on one input. "My program took O(n) seconds" is not a statement.

And a fourth that is worth naming. Big O hides the constant, and for two algorithms with the same growth the constant is exactly what decides which to use. It is a tool for comparing growth and not for comparing implementations.

Distinctions

OOmegaTheta
Boundupperlowerboth
Proved bygiving c and n0giving c and n0giving both pairs
Used foran algorithm's costa problem's difficultya tight statement
f is O(f)yesyesyes
Olittle o
n against n squaredn is O(n squared), and so is n squaredn is little o of n squared; n squared is not
Allows equality of growthyesno
Defined byconstantsa limit

What it does NOT mean

O(g) is not "the running time is g". It is an upper bound, and a loose one is still true.

The equals sign is not equality. It is membership, and reading it as equality leads to false statements.

Big O is not about small inputs. n0 exists so that any finite number of small cases may be ignored, which is why an algorithm that is O(n) may be slower than one that is O(n squared) on every input a person will ever run.

A constant factor is not hidden dishonestly. Chapter 78's speed up theorem says it is not a property of the algorithm, so there is nothing to hide.

munotes.in399

Big O Notation, and the Family Around It

2 to the n and 3 to the n are not the same. Their ratio grows without bound, so neither constant factors nor n0 can bridge them, unlike 2 to the n and 2 to the (n plus 1).

Quick revision

  • f is O(g) when there are positive constants c and n0 with f(n) at most c times g(n) for every n at least n0.
  • c absorbs the constant factor and n0 absorbs the small cases, and the existential quantifiers come first.
  • To prove a bound, give c and n0. For a polynomial, replace every lower power by the highest and add the

coefficients.

  • To disprove one, take the supposed c and produce an n depending on it that breaks the inequality.
  • Omega is a lower bound, Theta is both, little o is strictly slower and little omega strictly faster.
  • Theta is usually what is meant; O alone leaves a loose bound true.
  • The equals sign in f equals O(g) is an abuse for membership, and this book writes "is O(g)".
  • At n equal to 50, n squared is 2500 and 2 to the n is over a thousand million million, which is chapter 81's

argument in one row.

Test yourself

1. Give the definition of big O, with both constants and the order of the quantifiers. f is O(g) when there EXIST positive constants c and n0 such that FOR ALL n at least n0, f(n) is at most c times g(n). The constants are chosen first and must work for every large n.

2. Prove that 4n squared plus 7n plus 1 is O(n squared). For n at least 1 both n and 1 are at most n squared, so the expression is at most 4 plus 7 plus 1, that is 12, times n squared. Take c equal to 12 and n0 equal to 1.

3. Prove that n squared is not O(n). Suppose constants c and n0 exist. Take n to be the larger of n0 and c plus 1. Then n squared is n times n, which is at least (c plus 1) times n, which exceeds c times n, contradicting the bound.

4. What do each of c and n0 do, and what breaks without them? c absorbs the constant factor; without it 3n would not be O(n). n0 absorbs the small cases; without it n squared plus 100 would not be O(n squared), since the inequality fails at n equal to 1.

5. Why is Theta usually what a student means? Because O alone gives an upper bound that may be loose: an algorithm that is O(n) is also O(n squared), so stating O(n squared) is true and says less than it appears to. Theta asserts the bound is tight.

munotes.in400

Big O Notation, and the Family Around It

6. Is 2 to the (n plus 1) the same growth as 2 to the n? Is 3 to the n? 2 to the (n plus 1) is 2 times 2 to the n, so it is a constant factor away and the growth is the same. 3 to the n is not: the ratio to 2 to the n grows without bound, so no constant can bridge them.

Contents This chapter on its own page

munotes.in401

Chapter Eighty-One

Class P

Syllabus topic Module 2, "Computability and Complexity: Class P and Class NP"

In one line

P is the class of problems a deterministic machine decides in a number of steps bounded by some polynomial in the length of the input.

In the wording a student can write in an examination: P is the class of languages L for which there is a deterministic Turing machine deciding L and a polynomial p such that the machine halts within p(n) steps on every input of length n. Problems in P are called tractable, and P is taken as the formal counterpart of the informal notion of a problem that can be solved efficiently.

The definition

L is in P when some deterministic machine decides L in time O(n to the k), for some fixed k

Three words are doing work.

Deterministic, which distinguishes it from chapter 82's class.

Decides, so the machine halts on every input. A machine that may loop has no time complexity at all, by chapter 78, so undecidable problems are not in P and are not out of it either: the question does not arise.

Some fixed k, chosen before the input. The machine may not take n to the n steps, because that is not a polynomial: the exponent has to be a constant.

Why polynomial, and not some other line

The definition looks arbitrary and it is not. Four reasons, and an answer giving two of them is a good answer.

One: polynomials are closed under the things algorithms do

A polynomial of a polynomial is a polynomial. So:

Doing thisCostsAnd the result is
running one polynomial algorithm after anotherthe sumpolynomial
running one inside a loop of polynomially many stepsthe productpolynomial
feeding one polynomial algorithm's output to anotherthe compositionpolynomial

So a program built out of efficient parts is efficient. No other natural class has that property: if "efficient" meant "at most n squared steps", then calling an n squared routine n squared times would not be efficient, and the notion would be useless for building things.

Two: it does not depend on the model

Chapter 78's table said it. Changing between a one tape machine and a many tape one costs a square. Changing between a tape machine and a random access machine costs a polynomial. So the class P is the same on all of them, and a statement that a problem is in P is a statement about the problem rather than about the machine.

That robustness is the strongest argument, and it is why P and not some sharper class is the object of study. Anything defined more finely would be a property of the model.

munotes.in402

Class P

Three: the gap is real and it is enormous

Chapter 80's table showed it. At n equal to 50, n squared is 2500 and 2 to the n is over a thousand million million.

And faster hardware does not help an exponential. A machine a thousand times faster lets a polynomial algorithm handle an input about 32 times larger, if the polynomial is n squared. It lets an exponential algorithm handle an input ten symbols larger, because a thousand is about 2 to the 10. So a hardware improvement moves a polynomial boundary by a factor and an exponential boundary by an addition, and the addition is always small.

Four: it matches practice, roughly

Most problems found to be in P turn out to have algorithms with small exponents, usually 1, 2 or 3, and most problems not known to be in P have no practical algorithm at all. The correspondence is not exact, and the next section is about where it fails.

Where the identification breaks down

An honest chapter says this, and MU's students should be able to.

A polynomial of high degree is not practical. An algorithm running in n to the 100 steps is in P and is useless: at n equal to 10 it does 10 to the 100 steps, which is more than the atoms in the observable universe.

A large constant is not practical either. An algorithm running in a million times n steps is linear and is in P, and the constant may sink it.

An exponential with a small base may be fine in practice. An algorithm running in 1.01 to the n steps is not in P and is perfectly usable up to a large n.

So P is a useful idealisation and not a definition of practical. It is the right class to reason about because of reasons one and two, and the fit with practice is reason four, which is the weakest of the four.

What is in P

Every one of these has a known polynomial algorithm, and several appeared earlier in this book.

ProblemIn P byWhere in this book
does this finite automaton accept this stringrunning it, linear in the stringchapter 39
is this finite automaton's language emptyreachability, linear in the machinechapter 39
do two finite automata agreethe product and emptiness, quadraticchapter 39
does this context free grammar generate this stringCYK, cubicchapter 51
is a number even, or divisible by 3reading digits, linearchapter 13
sorting, searching, shortest paths, matchingstandard algorithms
is a number primea polynomial algorithm exists

The last row is worth a sentence. Primality was for a long time known to be in NP and not known to be in P, and a polynomial algorithm was found. It is the standard example of a problem moving into P, and it is a reminder that "not known to be in P" is a statement about knowledge.

munotes.in403

Class P

What is not known to be in P

Chapter 82's class NP is full of them, and chapter 84 names the hardest. The important point for this chapter is the shape of the claim.

No problem in NP has been PROVED to be outside P. Not one. Whether any is, is chapter 86's question, and it is open.

So "not in P" is almost always shorthand for "not known to be in P", and a careful answer says which is meant. The only problems proved outside P are ones proved outside by the hierarchy theorem of chapter 86, and they are artificial.

Closure properties

P is closed under the operations, which follows from reason one above and is worth stating because chapter 84 uses it.

OperationIn P?Why
complementyesswap the decider's answers; it still halts in time
unionyesrun both and take either
intersectionyesrun both and take both
concatenationyestry every split point, which is n of them
Kleene closureyesdynamic programming over the prefixes

The first row is the one that matters, because chapter 82's class NP is NOT known to be closed under complement, and that asymmetry is one of the two things that make NP interesting.

Distinctions

PThe recursive languages
Machinedeterministic, polynomial timedeterministic, halts eventually
Containsthe tractable problemseverything decidable
RelationshipP is inside the recursive languages, properly
A problem outsidemay be hard or undecidableis undecidable
PolynomialExponential
A thousand times faster hardwarehandles about 32 times more input, for n squaredhandles about 10 more symbols
Composes with itselfyes, still polynomialno
Calledtractableintractable

What it does NOT mean

In P does not mean fast. An n to the 100 algorithm is in P and unusable.

Not in P does not mean slow. A 1.01 to the n algorithm is outside P and usable for a long way.

Not known to be in P is not the same as not in P. No problem in NP has been proved outside P, and saying otherwise is the most common error on this topic.

P is not about the problem's inputs being small. It is about the growth of the cost with the input's size.

An undecidable problem is not outside P in the interesting sense. It has no complexity at all, because complexity is defined only for machines that halt.

Quick revision

  • P is the class of languages decided by a deterministic Turing machine within p(n) steps for some fixed
munotes.in404

Class P

polynomial p.

  • Four reasons for the line: polynomials compose, so programs built from efficient parts are efficient; the

class does not depend on the machine model; the gap between polynomial and exponential is enormous and hardware does not close it; and it roughly matches practice.

  • Faster hardware multiplies the input a polynomial algorithm can handle and merely ADDS to what an exponential

one can.

  • It is an idealisation: n to the 100 is in P and useless, and 1.01 to the n is outside P and usable.
  • Membership and emptiness for finite automata, CYK for context free grammars, sorting, shortest paths and

primality are all in P.

  • P is closed under complement, union, intersection, concatenation and closure. The complement row is the one

chapter 82 will contrast.

  • No problem in NP has been proved outside P, so "not in P" almost always means "not known to be in P".

Test yourself

1. Define P. The class of languages L for which some deterministic Turing machine decides L, halting within p(n) steps on every input of length n for some fixed polynomial p.

2. Give two reasons why the line is drawn at polynomials. Because polynomials are closed under sum, product and composition, so a program built from efficient parts is efficient; and because the class is the same on every reasonable machine model, since the conversions between models cost only a polynomial.

3. What does a thousandfold hardware improvement do for a polynomial algorithm, and for an exponential one? For an n squared algorithm it multiplies the manageable input size by about 32. For a 2 to the n algorithm it adds about 10 to it, since a thousand is roughly 2 to the 10.

4. Is an algorithm running in n to the 100 steps in P? Is it practical? It is in P, since 100 is a fixed exponent. It is not practical: at n equal to 10 it does 10 to the 200 steps.

5. Name three problems in P from earlier in this book. Membership for a finite automaton, emptiness for a finite automaton, and membership for a context free grammar by the CYK algorithm, which is cubic.

6. A student writes that the travelling salesman problem is not in P. What is wrong? It is not KNOWN to be in P. No problem in NP has been proved to lie outside P, and whether any does is the open question of chapter 86.

Contents This chapter on its own page

munotes.in405

Chapter Eighty-Two

Class NP, and the Certificate

Syllabus topic Module 2, "Computability and Complexity: Class P and Class NP"

In one line

NP is the class of problems whose YES answers can be checked quickly, even if nobody knows how to find them quickly.

Two definitions, and they are the same class

The first. L is in NP when some NONDETERMINISTIC Turing machine decides L in polynomial time. That is chapter 81's definition with determinism dropped, and it is where the name comes from: nondeterministic polynomial time.

The second. L is in NP when there is a polynomial time deterministic machine V, the verifier, and a polynomial q, such that:

w is in L if and only if some string c with length at most q(|w|) makes V accept the pair (w, c)

The string c is the certificate, also called the witness or the proof. The verifier does not have to find it. It has to check it.

The second definition is the useful one and the one to reach for in an examination, because showing a problem is in NP then means exhibiting a certificate and a checker, which is a concrete thing to write down.

Why the two agree

Both directions are short, and MU can ask for either.

A verifier gives a nondeterministic machine. Guess the certificate one symbol at a time, at most q(n) symbols, then run V. Guessing costs q(n) steps and V costs a polynomial, so the whole machine is polynomial, and it has an accepting run exactly when some certificate works.

A nondeterministic machine gives a verifier. Let the certificate be the LIST OF CHOICES the machine makes, one entry per step. A polynomial time machine takes at most p(n) steps, so the list is at most p(n) entries long, which is the polynomial bound. The verifier reads the list and follows the machine deterministically, making the choice the list names at each step, which costs p(n) steps. It accepts exactly when that run accepts, so the machine has an accepting run exactly when some list makes the verifier accept.

So the certificate is the accepting run. That single sentence is the whole equivalence, and it is worth remembering in that form.

What NP does not stand for

NP does not stand for "not polynomial". This is the commonest error on the topic and it is not a small one, because the true statement is nearly the opposite: every problem in P is in NP.

WrongRight
NP means not polynomialNP means nondeterministic polynomial
NP problems are the hard onesNP contains sorting, and every other problem in P
NP means no polynomial algorithm existsfor no problem in NP has that been proved
NP problems cannot be solvedall of them can; the question is how fast
munotes.in406

Class NP, and the Certificate

P is inside NP, and the proof is one line: a deterministic machine is a nondeterministic machine that happens never to have a choice. In the verifier form, take the certificate to be the empty string and let V be the decider, which ignores it.

Four certificates, worked and checked

Each of these is a problem MU or any textbook names, with the certificate written out and the check done.

Satisfiability

The problem. Given a formula in conjunctive normal form, is there an assignment making it true?

p | q | -r
-p | -q
q | r

The certificate is the assignment. The verifier substitutes and evaluates: one pass over each clause, so its cost is linear in the formula's length.

F1 is satisfied by p=0, q=1, r=1

F1 is satisfiable

F1 has 3 satisfying assignments

Why it is in NP and not obviously in P. The certificate is n bits long, so it is short, and checking it is linear. Finding it appears to need a search of 2 to the n assignments. Nobody has proved it does.

Clique

The problem. Given a graph and a number k, are there k vertices all joined to one another?

a - b
a - c
a - d
b - c
b - d
c - d
a - e
b - e
c - f
e - f

The certificate is the list of k vertices. The verifier checks each of the k(k-1)/2 pairs, which is quadratic in k and therefore polynomial.

G1 has a clique of size 4: a, b, c, d

G1 has no clique of size 5

max clique(G1) = 4

Read the second claim carefully. Proving there is NO clique of size 5 is not what the certificate does. A certificate establishes a YES answer only, and that asymmetry is the next section.

Hamiltonian path

The problem. Given a graph, is there a path visiting every vertex exactly once?

a - b
b - c
c - d
d - e
e - f
a - c
b - e

The certificate is the ordering. The verifier checks that every vertex appears exactly once and that each consecutive pair is joined, which is linear in the graph.

G2 has a Hamiltonian path: a, b, c, d, e, f

And a graph where no certificate exists, for contrast:

c - x
c - y
c - z

G3 has no Hamiltonian path

Why G3 has none, in words a student can write. A path has two ends, so at most two of its vertices have one neighbour on the path. Here x, y and z each have exactly one neighbour in the whole graph, so all three would have to be ends, and a path has only two.

munotes.in407

Class NP, and the Certificate

Subset sum

The problem. Given a set of whole numbers and a target, is there a subset adding up to the target?

set: 3 34 4 12 5 2
target: 9

S1 has a subset summing to the target: 4, 5

set: 3 34 4 12 5 2
target: 30

S2 has no subset summing to the target

The certificate is the subset. The verifier adds it up. The addition of numbers written in n bits costs a polynomial in n, so the verifier is polynomial.

The second instance is the one to notice. S2 has no subset summing to 30, and the only way shown here to know that is to try all 64 subsets. There is no short certificate for a NO answer, and nobody knows of one.

The asymmetry, and co-NP

A certificate proves YES and says nothing about NO. So NP is not obviously closed under complement, and this is one of the two facts that make the class interesting.

The complements have their own class. co-NP is the class of problems whose complement is in NP, that is, the problems whose NO answers have short certificates.

ProblemIn NP becauseIn co-NP because
is this formula satisfiablean assignmentnot known
is this formula unsatisfiablenot knownan assignment, for the complement
is this formula a tautologynot knownan assignment making it false
is this number primeit is in P, so in bothit is in P, so in both

Whether NP equals co-NP is open, and it is a different open question from chapter 86's, though a negative answer to this one would settle that one too: if P equalled NP then NP would equal co-NP, because P is closed under complement by chapter 81.

How big NP is

The bounds that are known are these, and they are worth memorising as a chain.

P is inside NP, and NP is inside EXPTIME

Why NP is inside EXPTIME. Try every certificate. There are at most 2 to the q(n) of them and each check costs a polynomial, so the total is exponential. So every problem in NP is solvable, and by an algorithm a student can describe in one sentence. That sentence is the answer to "can NP problems be solved at all", which is yes, and slowly.

Neither containment is known to be strict. Chapter 86 is about the first of them.

Distinctions

PNP
Machinedeterministic, polynomialnondeterministic, polynomial
Other definitionthe decider itselfa verifier and a certificate
Certificate needednoneat most polynomially long
Closed under complementyes, provednot known
Containsthe tractable problemsP, and the search problems
munotes.in408

Class NP, and the Certificate

VerifierDecider
Inputthe instance AND a certificatethe instance alone
Jobcheck a proposed answerfind the answer
Cost herepolynomialunknown for the NP complete problems
AnswersYES only, via a certificateboth

What it does NOT mean

NP does not mean not polynomial. P is inside NP.

A problem in NP is not undecidable. Every one of them is decidable by trying every certificate, which is exponential and finite.

A certificate is not an algorithm. It is the answer's evidence. The verifier is the algorithm, and it needs to be handed the certificate.

A certificate for a NO answer is not part of the definition. That is co-NP, and it is a different class.

A nondeterministic machine is not a machine that guesses correctly. It is a machine that accepts when SOME run accepts, which is the same thing formally and is the honest way to say it.

Quick revision

  • NP is nondeterministic polynomial time: the class decided by a nondeterministic Turing machine in polynomial

time.

  • Equivalently, the class with a polynomial time verifier and a polynomially long certificate. The certificate

IS the accepting run, and that is the whole proof that the two definitions agree.

  • P is inside NP, because a deterministic machine is a nondeterministic one with no choices.
  • Certificates: an assignment for satisfiability, a vertex list for clique, an ordering for a Hamiltonian

path, a subset for subset sum.

  • A certificate establishes YES only. The class for NO certificates is co-NP, and whether it equals NP is open.
  • NP is inside EXPTIME, because trying every certificate takes exponential time, so every NP problem is

decidable.

Test yourself

1. Give both definitions of NP. The class of languages decided by a nondeterministic Turing machine in polynomial time; equivalently the class of languages L for which some deterministic polynomial time verifier V and some polynomial q satisfy: w is in L exactly when V accepts (w, c) for some c no longer than q of the length of w.

2. Prove that a nondeterministic polynomial time machine gives a verifier. Take the certificate to be the list of the machine's choices, one per step. The machine runs at most p(n) steps, so the list is at most p(n) long. The verifier follows the machine deterministically, taking the choice the list names, for p(n) steps. Some list makes it accept exactly when some run of the machine accepts.

3. What is the certificate for clique, and what does the verifier cost? The k vertices. The verifier checks all k(k-1)/2 pairs for an edge, which is quadratic in k.

4. Show that P is inside NP, in one sentence. A deterministic polynomial time decider is a verifier that ignores its certificate, so taking the certificate to be the empty string puts every language in P into NP.

munotes.in409

Class NP, and the Certificate

5. Why is NP not obviously closed under complement? Because a certificate establishes a YES answer, and nothing in the definition supplies a short proof of a NO answer. The class of problems with short NO certificates is co-NP, and whether it equals NP is open.

6. A student writes that NP problems cannot be solved. Correct it. Every problem in NP can be solved: try every certificate, of which there are at most 2 to the q(n), checking each in polynomial time. That is an exponential time algorithm, so NP is inside EXPTIME. What is unknown is whether any of them needs more than polynomial time.

7. Which of these graphs has no Hamiltonian path, and why? The three pointed star, because x, y and z each have one neighbour and would all have to be an end of the path, and a path has two ends.

Contents This chapter on its own page

munotes.in410

Chapter Eighty-Three

Polynomial Reductions

Syllabus topic Module 2, "Computability and Complexity: Polynomial Reductions"

In one line

A polynomial reduction converts every instance of one problem into an instance of another with the same answer, in polynomial time. A polynomial time reduction is the same thing said the other way round, and both names are used.

The definition

A reduces to B in polynomial time when there is a function f, computable by a deterministic machine in polynomial time, such that for every instance x of A:

x is a yes instance of A if and only if f(x) is a yes instance of B

Three things must be true, and an answer that establishes two of them has not finished.

What must be shownWhat goes wrong without it
1f turns a yes into a yesthe reduction loses solutions
2f turns a no into a nothe reduction invents solutions
3f runs in polynomial timethe reduction proves nothing at all

Condition 3 is the new one, and it is the one this chapter is about. Chapter 75 needed only that f be computable.

Why condition 3 is not a formality

Here is a transformation that satisfies conditions 1 and 2 for ANY pair of problems A and B, as long as B has at least one yes instance and one no instance.

Given x, solve A on x by brute force. If the answer is yes, output a fixed yes instance of B. If it is no,

output a fixed no instance of B.

Conditions 1 and 2 hold perfectly. Yes goes to yes and no goes to no, exactly. And it establishes nothing whatever, because the work was done inside f. If a reduction of this kind counted, every problem would reduce to every other, and the whole idea would be empty.

So the time bound is what carries the meaning. The reduction has to be cheaper than solving the problem, or it says nothing about how hard the problem is.

The two uses, and the direction

The direction is the same as chapter 75's and students get it backwards just as often, so it is worth setting down again with the clock in it.

A reduces to B in polynomial time: a fast algorithm flows from B to A, hardness flows from A to B

Use one, forwards. If A reduces to B and B is in P, then A is in P. Given x, compute f(x), which costs a polynomial, then decide it with B's polynomial algorithm, which costs a polynomial of a polynomial. The whole thing is polynomial, by chapter 81's first closure argument.

Use two, backwards. If A reduces to B and A is hard, then B is hard. Because if B were easy, use one would make A easy.

munotes.in411

Polynomial Reductions

The mnemonic. B is at least as hard as A. So the KNOWN hard problem goes on the left. Every proof in chapter 85 reduces a problem already known to be NP complete TO the new problem, never the other way round, and a proof written the wrong way round establishes nothing.

And the relation carries. If A reduces to B and B reduces to C, then A reduces to C, because a polynomial of a polynomial is a polynomial. That is why one theorem, Cook's, was enough to start a list that now runs into thousands.

Reduction one: independent set reduces to clique

The two problems. An independent set is a set of vertices no two of which are joined. A clique is a set of vertices every two of which are joined. The definitions are exact opposites, which is what makes the reduction one line.

The transformation. Given the pair (G, k), output the pair (complement of G, k). The complement has the same vertices and exactly the edges G lacks.

Here is a graph, the five cycle:

a - b
b - c
c - d
d - e
e - a

G1 has an independent set of size 2: a, c

G1 has no independent set of size 3

max independent(G1) = 2

And here is its complement, which f produces:

a - c
a - d
b - d
b - e
c - e

G2 = complement(G1)

G2 has a clique of size 2: a, c

G2 has no clique of size 3

max clique(G2) = 2

Notice what the complement of the five cycle is. Another five cycle, a, c, e, b, d, which is why both numbers come out at 2.

Both directions, in one sentence each. If S is an independent set of G, no two of its vertices are joined in G, so every two are joined in the complement, so S is a clique there. And the same sentence read backwards. k does not change, which is what makes this the easiest reduction in the subject.

The check that matters is not the one instance but the identity, and it holds for every k at once:

independent(G1, k) iff clique(complement(G1), k)

The running time. Building the complement means deciding, for each of the n(n-1)/2 pairs, whether G has that edge. That is quadratic in the number of vertices. Polynomial, so condition 3 holds.

Reduction two: independent set reduces to vertex cover

The second problem. A vertex cover is a set of vertices touching every edge.

The transformation. Given (G, k), output (G, n - k), where n is the number of vertices. The graph does not change at all. Only the number does.

munotes.in412

Polynomial Reductions

Take the same five cycle. n is 5, so an independent set of size 2 should correspond to a cover of size 3:

G1 has a vertex cover of size 3: a, b, d

G1 has no vertex cover of size 2

min cover(G1) = 3

Why it works. S is an independent set exactly when every edge has at most one end in S, which is exactly when every edge has at least one end OUTSIDE S, which is exactly when the vertices outside S form a cover. So the independent sets and the covers are complements of one another as sets of vertices, and their sizes add up to n.

independent(G1, k) iff cover(G1, n - k)

Read the table across to see it. These are the real numbers for the five cycle:

Size of the independent setDoes G1 have oneSize of the cover it leavesDoes G1 have one
0yes5yes
1yes4yes
2yes3yes
3no2no
4no1no
5no0no

The running time. Subtracting k from n. Constant, once the graph is copied, and copying is linear. This is about as cheap as a reduction gets, and it is worth seeing one like it so that the idea of a reduction does not become confused with the idea of hard work.

Reduction three: Hamiltonian cycle reduces to Hamiltonian path

The two problems. A Hamiltonian cycle visits every vertex exactly once and returns to its start. A Hamiltonian path visits every vertex exactly once and need not return. A graph may have a path and no cycle, so the reduction is not trivial and cannot simply hand the graph over unchanged.

The transformation. Given G, pick any vertex v. Build G' like this:

StepWhat to do
1replace v by two new vertices, v1 and v2
2join each of them to every neighbour v had
3add a new vertex s, joined to v1 only
4add a new vertex t, joined to v2 only

Why s and t are there. They have one neighbour each, so any Hamiltonian path in G' must have them as its two ends, by the argument chapter 82 used on the star. That forces the path to run from s through v1, across all the old vertices, into v2 and out to t, which is precisely a cycle through v cut open.

Worked on a yes instance

a - b
b - c
c - d
d - a

G3 has a Hamiltonian cycle: a, b, c, d

Splitting a gives:

munotes.in413

Polynomial Reductions

s - a1
a1 - b
a1 - d
a2 - b
a2 - d
b - c
c - d
a2 - t

G4 has a Hamiltonian path: s, a1, b, c, d, a2, t

Line the two up. The cycle is a, b, c, d, back to a. The path is s, then a1 standing for the first visit to a, then b, c, d in the same order, then a2 standing for the return to a, then t. The path is the cycle with its start split in two and a handle attached at each end.

Worked on a no instance

Two triangles sharing the vertex c. Every vertex has at least two neighbours, so nothing about the degrees settles it, and it still has no Hamiltonian cycle, because any cycle through both triangles would have to pass through c twice:

a - b
b - c
c - a
c - d
d - e
e - c

G5 has no Hamiltonian cycle

Splitting a, whose neighbours are b and c:

s - a1
a1 - b
a1 - c
a2 - b
a2 - c
b - c
c - d
d - e
e - c
a2 - t

G6 has no Hamiltonian path

So a no instance went to a no instance, which is condition 2 on this pair, and the two instances together show both conditions at work on real graphs.

The running time. One vertex is duplicated, its neighbour list is copied twice, and two vertices with one edge each are added. That is linear in the size of the graph. Polynomial, so condition 3 holds.

The three reductions side by side

Independent set to cliqueIndependent set to coverHamiltonian cycle to path
The graphreplaced by its complementunchangedone vertex split, two added
The numberunchangedn minus kthere is none
Cost of fquadraticlinearlinear
Hardest part of the proofneither, it is symmetricneitherforcing the ends to be s and t

Distinctions

Reduction (chapter 75)Polynomial reduction
f must becomputablecomputable in polynomial time
Used to proveundecidabilityhardness within the decidable
What flows backwardsundecidabilitythe absence of a fast algorithm
A brute force fis alloweddestroys the proof
A reduces to BB reduces to A
SaysB is at least as hard as AA is at least as hard as B
To prove B hardthis is the one to writethis proves nothing about B
If B is in Pthen A is in Pnothing follows about A

What it does NOT mean

It does not mean the two problems are the same problem. Independent set and clique are different questions; the reduction says only that an algorithm for one answers the other.

munotes.in414

Polynomial Reductions

It does not mean the reduction runs both ways. A reduction has a direction. Both of the first two here happen to be reversible, and the third is not obviously so.

It does not mean f may look at the answer. A transformation that solves A in order to decide what to output satisfies the first two conditions and is worthless, which is why condition 3 exists.

It does not make A and B equally hard. It makes B at least as hard as A. B may be much harder.

The number k need not stay the same. In the second reduction it becomes n minus k, and forgetting to change it is the commonest slip in writing this one out.

Quick revision

  • A reduces to B in polynomial time when a polynomial time computable f sends yes instances to yes instances

and no instances to no instances.

  • Three conditions: yes to yes, no to no, and f polynomial. The third is what separates this from chapter 75.
  • Without the time bound, every problem reduces to every other by solving it inside f, so the bound carries all

the meaning.

  • Direction: a fast algorithm flows from B to A, hardness flows from A to B, so the known hard problem goes on

the left.

  • The relation carries: A to B and B to C gives A to C, because a polynomial of a polynomial is a polynomial.
  • Independent set to clique: take the complement, keep k. Quadratic.
  • Independent set to vertex cover: keep the graph, use n minus k. Linear.
  • Hamiltonian cycle to Hamiltonian path: split one vertex in two and hang a new vertex off each half, which

forces those two to be the ends of the path. Linear.

Test yourself

1. State the three conditions a polynomial reduction must satisfy. That f sends every yes instance of A to a yes instance of B, that it sends every no instance to a no instance, and that f is computable in polynomial time.

2. Why is the time bound essential? Without it, f could solve A by brute force and output a fixed yes or no instance of B. That satisfies the first two conditions for every pair of problems and proves nothing, so every problem would reduce to every other.

3. To prove a new problem B is hard, which way round do you write the reduction? Reduce a problem already known to be hard to B. The known problem goes on the left, because the reduction shows B is at least as hard as it.

4. Reduce independent set to clique, and say what happens to k. Replace the graph by its complement and leave k alone. A set is independent in G exactly when it is a clique in the complement, since the complement has exactly the edges G lacks.

munotes.in415

Polynomial Reductions

5. Reduce independent set to vertex cover. Keep the graph and ask for a cover of size n minus k. A set is independent exactly when the vertices outside it form a cover, so the two sizes add up to n.

6. In the Hamiltonian reduction, what forces the path to begin at s and end at t? Each of s and t has exactly one neighbour. A path has two ends, and any vertex of degree one that lies on the path must be one of them, so s and t are the two ends.

7. If A reduces to B in polynomial time and B is in P, what follows? And if A is not in P? A is in P, because computing f and then running B's algorithm is a polynomial of a polynomial. If A is not in P then B is not in P either, by the same argument read backwards.

Contents This chapter on its own page

munotes.in416

Chapter Eighty-Four

NP Complete and NP Hard, and Cook's Theorem

Syllabus topic Module 2, "Computability and Complexity: NP Complete and NP Hard Problems"

In one line

NP hard means at least as hard as everything in NP. NP complete means that, and in NP as well.

The two definitions

A problem B is NP hard when every problem A in NP reduces to B in polynomial time.

A problem B is NP complete when B is NP hard AND B is itself in NP.

NP complete = NP hard AND in NP

So NP complete is the smaller notion, and every NP complete problem is NP hard. The reverse fails, and the example is in this chapter.

The plain reading. An NP hard problem is one a fast algorithm for which would give a fast algorithm for everything in NP. An NP complete problem is one of those which is itself an NP problem, so it sits inside NP, at the top.

The difference, which is what gets asked

NP hardNP complete
Every NP problem reduces to ityes, by definitionyes
Must itself be in NPnoyes
Must be decidablenoyes
Must be a decision problemnoyes
Examplethe halting problemsatisfiability
Another examplefinding the largest cliquedeciding whether a clique of size k exists
If it is in Pthen P equals NPthen P equals NP

Rows 3 and 4 are the ones students miss. NP hard says nothing about the problem itself. It is a statement about everything BELOW the problem, so a problem can be NP hard and undecidable, and a problem can be NP hard and not even be a yes or no question.

The consequence that makes the idea useful

If any NP complete problem is in P, then P equals NP.

The proof is three lines and MU can ask for it. Let B be NP complete and suppose B is in P. Take any A in NP. Since B is NP hard, A reduces to B in polynomial time. By chapter 83's first use, A is then in P. A was any problem in NP, so all of NP is in P, and since P is inside NP by chapter 82, the two are equal.

And the contrapositive. If P is not equal to NP, then no NP complete problem is in P. So the whole class stands or falls together: find a polynomial algorithm for ONE of them and you have one for all of them, and prove that ONE of them has none and you have proved it of all of them.

That is why the class is worth naming. A long list of separate problems became one problem.

Cook's theorem

The statement in the form textbooks now use. Satisfiability is NP complete.

The statement Cook printed. His Theorem 1, on page 152 of the scan, says that if a set of strings is accepted by some nondeterministic Turing machine within polynomial time, then that set is P reducible to the set of tautologies in disjunctive normal form.

munotes.in417

NP Complete and NP Hard, and Cook's Theorem

Three differences between the two statements are worth knowing, because a student who has only seen the modern form will not recognise the original.

Cook, 1971The modern textbook
The target problemtautologies in disjunctive normal formsatisfiability in conjunctive normal form
The reductionby oracle, his "query machine"by transformation, one instance to one instance
The phrase usedP reduciblepolynomial time reducible, or Karp reducible

And the modern form is inside his own proof. The proof on the same page builds, from a machine M and an input w, a formula A(w) in CONJUNCTIVE normal form which is satisfiable exactly when M accepts w. The step to tautologies is then De Morgan: the negation of A(w) is in disjunctive form and is a tautology exactly when M does NOT accept w.

The paper's own note about its venue is handwritten across the top left corner of the first page, which is why the citation beside this book records it as an annotation and not as a printed journal line.

Why the duality holds, checked

The bridge between the two statements is worth one worked example, because MU can set either form.

A formula with no satisfying assignment:

x | y
-x
-y

F1 is unsatisfiable

Negate it. Each clause becomes a term with every literal flipped, by De Morgan:

-x & -y
x
y

D1 = negation(F1)

D1 is a tautology

And the other way, with a formula that IS satisfiable:

x | y
-x
-x & -y
x

F2 is satisfiable

D2 = negation(F2)

D2 is not a tautology

So the two questions are the same question upside down. A formula is unsatisfiable exactly when its negation is a tautology, and the negation of a conjunctive normal form is a disjunctive one. Cook asked about tautologies; we ask about satisfiability; it is one theorem.

How the proof works, in the shape a student can reproduce

The full construction is long. The IDEA is short, and the idea is what an examination wants.

StepWhat is done
1the machine runs at most p(n) steps, so it uses at most p(n) tape squares
2introduce a variable for every fact about the computation: square s holds symbol a at step t, the machine is in state q at step t, the head is on square s at step t
3write clauses saying each square holds exactly one symbol, the machine is in exactly one state, the head is in one place
4write clauses saying step 0 is the starting configuration with w on the tape
5write clauses saying each step follows from the one before by a legal move
6write a clause saying some step is accepting
munotes.in418

NP Complete and NP Hard, and Cook's Theorem

The formula is satisfiable exactly when an accepting computation exists, because a satisfying assignment IS an accepting computation written out as truth values, and every accepting computation gives a satisfying assignment.

And the formula is polynomially long, because there are a polynomial number of squares, steps and states, so a polynomial number of variables and clauses. That is condition 3 of chapter 83, and without it the proof would establish nothing.

What Cook's theorem made possible

Before it there was no first problem. To show something NP hard you had to reduce EVERY problem in NP to it, and that means arguing about an arbitrary machine, as Cook did.

After it, one reduction is enough. To show a new problem B is NP complete, show B is in NP and reduce ONE known NP complete problem to B. The chain then carries, because reduction is transitive by chapter 83.

Every problem in NP reduces to satisfiability, and satisfiability reduces to B, so every problem in NP reduces to B

That is the whole technique of chapter 85, and it is why a list that began with one problem in 1971 grew as fast as it did.

The map

ClassWhat it holdsKnown relationships
Pthe tractable problemsinside NP
NPshort certificates for yesholds P, inside EXPTIME
NP completethe hardest problems IN NPinside NP; equals P exactly when P equals NP
NP hardeverything at least as hard as NPholds all of NP complete, and much more

Reading the map. NP complete is a band inside NP. NP hard is a region that overlaps NP in exactly that band and then extends outward past everything decidable. If P equals NP the band collapses into P and the three inner classes become one; if not, P and NP complete are disjoint, and there is a third region between them whose members are in NP, not in P and not NP complete.

The halting problem: NP hard, not NP complete

This is the standard example and MU has asked for it.

It is NP hard. Take any A in NP. A is decidable, by chapter 82's exponential search. So given an instance x, write down a machine that runs that decider on x and then halts if the answer is yes and loops forever if the answer is no. Writing the machine down is a matter of copying a fixed program and inserting x, which is polynomial. That machine halts exactly when x is a yes instance of A, so A reduces to the halting problem in polynomial time.

munotes.in419

NP Complete and NP Hard, and Cook's Theorem

It is not in NP. Everything in NP is decidable, and the halting problem is not, by chapter 74.

So it is NP hard and not NP complete, and it shows how far outside NP the NP hard region reaches.

A second example, of a different kind. "What is the size of the largest clique in this graph" is NP hard and is not NP complete either, for a duller reason: it is not a yes or no question at all, and NP is a class of LANGUAGES. Optimisation problems are routinely NP hard and never NP complete.

Distinctions

NP hardNP completeIn NP
Definitioneverything in NP reduces to itNP hard and in NPhas a short certificate
Can be undecidableyesnono
Can be an optimisation problemyesnono
Contains the othersholds NP completeholds neitherholds neither
Proof neededone reduction from a known onethat, and a certificatea certificate
Cook's theoremIts use
What it establishesone problem is NP completea starting point
How it is provedby encoding an arbitrary machinenever again, after 1971
Later proofsreduce a known complete problemone reduction each

What it does NOT mean

NP complete does not mean unsolvable. Every NP complete problem is decidable, and solvable in exponential time by trying every certificate.

NP hard does not mean in NP. The halting problem is NP hard and is not even decidable.

NP complete does not mean proved to need exponential time. Nothing of the kind has been proved for any of them; that is chapter 86's open question.

Proving one NP complete problem hard does not require touching the others. The consequence carries automatically, which is the point of the class.

A reduction FROM the new problem proves nothing. It must go from a known NP complete problem TO the new one. Chapter 83 gives the mnemonic.

Cook's theorem is not about clique or the travelling salesman. It is about satisfiability, in his paper about tautologies, and everything else came afterwards by reduction.

Quick revision

  • B is NP hard when every problem in NP reduces to B in polynomial time. B is NP complete when it is NP hard

and is itself in NP.

  • NP hard says nothing about B itself, so an NP hard problem may be undecidable or may not be a decision

problem at all.

  • If any NP complete problem is in P then P equals NP, because every NP problem reduces to it and a polynomial
munotes.in420

NP Complete and NP Hard, and Cook's Theorem

of a polynomial is a polynomial.

  • Cook's theorem gave the first one. His Theorem 1 is phrased in terms of tautologies in disjunctive normal

form; the conjunctive formula satisfiable exactly when the machine accepts is built inside his proof.

  • The construction introduces a variable for every fact about the computation and clauses forcing those facts

to describe a legal accepting run, and the formula is polynomially long.

  • After Cook, proving a new problem NP complete needs only membership in NP and ONE reduction from a known

complete problem, because reduction is transitive.

  • The halting problem is NP hard and not NP complete, because it is not in NP: it is undecidable.

Test yourself

1. Define NP hard and NP complete, and say which is stronger. B is NP hard when every problem in NP reduces to it in polynomial time. B is NP complete when it is NP hard and also lies in NP. NP complete is the stronger condition, so every NP complete problem is NP hard.

2. Prove that if one NP complete problem is in P then P equals NP. Let B be NP complete and in P, and let A be any problem in NP. A reduces to B in polynomial time, so computing the reduction and then running B's polynomial algorithm decides A in polynomial time. So NP is inside P, and P is inside NP, so they are equal.

3. State Cook's theorem in both forms. Modern: satisfiability is NP complete. Cook's own Theorem 1: any set accepted by a nondeterministic Turing machine in polynomial time is P reducible to the set of tautologies in disjunctive normal form.

4. Sketch the proof. Given a machine running in p(n) steps, introduce a propositional variable for each statement of the form "square s holds symbol a at step t", "the state at step t is q" and "the head is at square s at step t". Write clauses forcing each of those to be single valued, clauses fixing step 0 to the input configuration, clauses making each step follow the last by a legal move, and one clause requiring an accepting step. The formula is satisfiable exactly when an accepting computation exists, and it has polynomially many variables and clauses.

5. Why is the halting problem NP hard but not NP complete? NP hard because every problem in NP is decidable, so an instance can be turned in polynomial time into a machine that halts exactly when the answer is yes. Not NP complete because NP contains only decidable problems and the halting problem is undecidable.

6. Give an NP hard problem that is not a decision problem. Finding the size of the largest clique in a graph. NP is a class of languages, that is, of yes or no questions, so an optimisation problem cannot be in NP and cannot be NP complete.

munotes.in421

NP Complete and NP Hard, and Cook's Theorem

7. A student proves a new problem B is NP complete by reducing B to satisfiability. What is wrong? The direction. Reducing B to satisfiability shows B is no harder than satisfiability, which is already known for everything in NP. To prove B is NP hard, satisfiability, or another known NP complete problem, must reduce TO B.

Contents This chapter on its own page

munotes.in422

Chapter Eighty-Five

Proving a Problem NP Complete: Three Reductions Worked

Syllabus topic Module 2, "Computability and Complexity: NP Complete and NP Hard Problems"

In one line

To prove B is NP complete: show B is in NP, then reduce a problem already known to be NP complete to B.

The recipe

Step 1. Show B is in NP. Name the certificate and say what the verifier does and what it costs. This step is usually two sentences and is usually the one left out, and leaving it out downgrades the proof from NP complete to NP hard.

Step 2. Pick a known NP complete problem A and reduce A to B. Three things are needed, which are chapter 83's three conditions:

What to show
2athe transformation f, described exactly
2byes goes to yes, and no goes to no
2cf runs in polynomial time

Why this is enough. Every problem in NP reduces to A, because A is NP complete. A reduces to B. Reduction is transitive, so every problem in NP reduces to B, so B is NP hard. With step 1, B is NP complete.

The direction, once more. A reduces to B, where A is the OLD problem and B is the NEW one. A proof that reduces the new problem to a known one establishes nothing about the new problem.

Reduction one: satisfiability to 3-satisfiability

3-satisfiability is satisfiability restricted to formulas whose every clause holds exactly three literals. It looks easier and is not.

Step 1, in NP. The certificate is the assignment, the verifier substitutes and evaluates, and that costs one pass over the formula. Every 3CNF formula is a CNF formula, so this is chapter 82's verifier unchanged.

Step 2, the transformation. Take each clause of the original in turn, and depending on how many literals it holds, replace it as the table says. The variables written y and z are BRAND NEW, a fresh set for each clause.

Literals in the clauseWhat replaces it
one, say afour clauses: a with each of the four settings of two new variables y1 and y2
two, say a or btwo clauses: (a or b or y) and (a or b or not y)
threenothing changes
k, more than threea chain of k minus 2 clauses linked by k minus 3 new variables

The chain, written out. For a clause of the form a1 or a2 or a3 or a4, with one new variable z1, the replacement is (a1 or a2 or z1) and (not z1 or a3 or a4).

Worked

p
q | r
-p | q | -r
p | -q | r | s

F1 is satisfiable

One clause of each kind: one literal, two, three, four. The conversion produces this:

p | y1 | y2
p | y1 | -y2
p | -y1 | y2
p | -y1 | -y2
q | r | y3
q | r | -y3
-p | q | -r
p | -q | z1
-z1 | r | s
munotes.in423

Proving a Problem NP Complete: Three Reductions Worked

F2 is in 3CNF

F2 = 3CNF(F1)

Read the groups down. The first four came from the single literal p, and they force p to be true whatever y1 and y2 are, because the four of them between them try every setting. The next two came from q or r, and force q or r for the same reason. The seventh is copied unchanged. The last two are the chain.

Both directions, in words.

Yes to yes. Take an assignment satisfying F1. Each clause has a true literal. For the one and two literal cases the new clauses are true whatever the new variables are. For the chain, set the new variables so the chain "passes" the satisfied position: if the true literal is in the first pair set z1 false, and if it is in the last pair set z1 true.

No to no, which is the same thing said backwards: take an assignment satisfying F2 and read off the original variables. The four clauses from p force p true. The two from q or r force q or r. The chain forces one of the four original literals true, because z1 is either false, in which case (p or not q) must hold, or true, in which case (r or s) must hold.

The cost. A clause of k literals becomes at most 4 clauses when k is small and k minus 2 clauses when k is large, so the new formula is at most a constant times the size of the old one, and each clause is handled by itself. Linear, so condition 2c holds.

Reduction two: 3-satisfiability to clique

This is the reduction Cook's own paper carries out, in its Theorem 2, against the subgraph problem rather than clique. The joining rule below is his.

Step 1, in NP. The certificate is the list of k vertices, and the verifier checks each of the k(k-1)/2 pairs for an edge. Chapter 82 gave this.

Step 2, the transformation. Given a 3CNF formula with m clauses, build a graph:

What to do
verticesone for every literal OCCURRENCE, so 3m of them, labelled by clause and literal
edgesjoin two vertices when they are in DIFFERENT clauses and are not a literal and its negation
the numberask for a clique of size m, the number of clauses

Worked

x | y | -z
-x | -y | z
x | -y | -z

F3 is in 3CNF

Three clauses, so nine vertices and a clique of size 3 is what we ask for. The vertex named 2.-y is the occurrence of not y in the second clause:

munotes.in424

Proving a Problem NP Complete: Three Reductions Worked

1.x - 2.-y
1.x - 2.z
1.x - 3.x
1.x - 3.-y
1.x - 3.-z
1.y - 2.-x
1.y - 2.z
1.y - 3.x
1.y - 3.-z
1.-z - 2.-x
1.-z - 2.-y
1.-z - 3.x
1.-z - 3.-y
1.-z - 3.-z
2.-x - 3.-y
2.-x - 3.-z
2.-y - 3.x
2.-y - 3.-y
2.-y - 3.-z
2.z - 3.x
2.z - 3.-y

G1 = clique-graph(F3)

G1 has a clique of size 3: 1.x, 2.-y, 3.x

Read the clique back as an assignment. The three vertices say: the first clause is satisfied by x, the second by not y, the third by x. So set x true and y false. z is not mentioned, so it may be anything.

F3 is satisfied by x=1, y=0, z=0

F3 is satisfied by x=1, y=0, z=1

Both directions, in words.

A satisfying assignment gives a clique. Pick one true literal in each clause and take its vertex. There are m of them, one per clause, so no two are in the same clause. No two are complementary, because both are true under the same assignment and a literal and its negation cannot both be true. So every pair is joined, and they form a clique of size m.

A clique of size m gives a satisfying assignment. No two vertices of a clique are in the same clause, since same clause vertices are never joined, so the m vertices are one from each clause. No two are complementary, since complementary vertices are never joined. So setting every named literal true is consistent, and it makes one literal true in every clause.

The smallest no instance shows the second direction working. This formula has one literal per clause and no satisfying assignment:

x
-x

F4 is unsatisfiable

Its graph has two vertices and no edge at all, because the only cross clause pair is complementary:

1.x
2.-x

G3 = clique-graph(F4)

G3 has no clique of size 2

The cost. There are 3m vertices, so fewer than 9m squared over 2 pairs to consider, and each is settled by comparing two labels. Quadratic in the formula, so condition 2c holds.

Reduction three: clique to vertex cover

Step 1, in NP. The certificate is the set of vertices in the cover, and the verifier checks that every edge has an end in it, which is linear in the number of edges.

Step 2, the transformation. Given (G, k), output (complement of G, n minus k), where n is the number of vertices.

Two things change here, the graph AND the number, and forgetting either is the usual mistake. Chapter 83 had a reduction where only the number changed and one where only the graph changed; this one is both at once, and it is those two put together.

munotes.in425

Proving a Problem NP Complete: Three Reductions Worked

Worked

a - b
b - c
c - a
c - d
d - e

G4 has a clique of size 3: a, b, c

max clique(G4) = 3

Five vertices, so we ask the complement for a cover of size 5 minus 3, which is 2:

a - d
a - e
b - d
b - e
c - e

G5 = complement(G4)

G5 has a vertex cover of size 2: d, e

min cover(G5) = 2

Look at the two answers together. The clique in G4 is a, b, c. The cover in G5 is d, e. They are complementary sets of vertices, and that is the whole reduction.

Both directions, in one argument. A set S is a clique in G exactly when no two of its vertices are joined in the complement, which is exactly when every edge of the complement has at least one end OUTSIDE S, which is exactly when the vertices outside S form a cover of the complement. The set outside S has n minus k vertices.

And the identity holds at every size at once, not only at the one worked above:

clique(G4, k) iff cover(complement(G4), n - k)

kClique of size k in G4Cover of size 5 minus k in the complement
0yesyes, all five
1yesyes
2yesyes
3yesyes, d and e
4nono
5nono

The cost. Building the complement is quadratic, as chapter 83 counted, and subtracting is free. Polynomial, so condition 2c holds.

The chain these three build

Every problem in NP reduces to satisfiability, which reduces to 3-satisfiability, which reduces to clique, which reduces to vertex cover

So all four are NP complete, on the strength of Cook's theorem and three reductions, and none of the three had to mention a Turing machine.

ProblemKnown complete byThe reduction from
satisfiabilityCook's theoremevery NP problem, via the machine
3-satisfiabilityreduction onesatisfiability
cliquereduction two3-satisfiability
vertex coverreduction threeclique
independent setchapter 83's first reductionclique

The mistakes that cost marks

MistakeWhy it is fatal
reducing the new problem to the known oneproves the new problem is EASY, not hard
forgetting step 1proves NP hard only, which is a different claim
arguing only that yes goes to yesa reduction that sends every instance to a yes instance passes that test
not bounding f's running timechapter 83 showed such a reduction proves nothing
changing the graph and forgetting the numberthe third reduction above fails outright
starting from a problem not known to be NP completethe chain has no beginning
munotes.in426

Proving a Problem NP Complete: Three Reductions Worked

Distinctions

Proving B is in NPProving B is NP hard
What is supplieda certificate and a verifiera reduction from a known complete problem
Directionnot applicablefrom the OLD problem to the new one
Cost to checkthe verifier must be polynomialthe transformation must be polynomial
If only this is doneB is in NP, which is weakB is NP hard, which is not the same as complete

What it does NOT mean

A reduction is not an algorithm for either problem. It converts instances. Neither problem is solved by it.

3-satisfiability being NP complete does not make 2-satisfiability hard. Clauses of two literals give a problem that is in P, and the chain stops there.

The clique the reduction asks for is not any clique. It is one of size exactly m, the number of clauses, and the size is part of the instance.

The complement in the third reduction is not the complement of the answer. It is the complement of the GRAPH, and the answer's complement is a separate fact that falls out of it.

Quick revision

  • Two steps: B is in NP, by a certificate and a verifier; and a known NP complete problem reduces to B in

polynomial time.

  • The reduction runs from the OLD problem to the NEW one. The other direction proves nothing.
  • Satisfiability to 3-satisfiability: pad short clauses with new variables in every combination, and split

long clauses into a chain linked by new variables. Linear.

  • 3-satisfiability to clique: one vertex per literal occurrence, join vertices in different clauses that are

not complementary, ask for a clique of size m. Quadratic.

  • Clique to vertex cover: take the complement of the graph and ask for a cover of size n minus k. Quadratic.
  • In the second, a clique of size m must take one vertex per clause and cannot hold a literal and its negation,

which is exactly a consistent satisfying assignment.

  • In the third, the clique and the cover are complementary sets of vertices, which is why the sizes add to n.

Test yourself

1. Give the two steps of an NP completeness proof. Show the problem is in NP, by naming a certificate and a polynomial time verifier; then reduce a problem already known to be NP complete to it in polynomial time, showing yes goes to yes, no goes to no, and the transformation is polynomial.

2. Convert the clause p or q or r or s into 3CNF. Introduce one new variable z and write (p or q or z) and (not z or r or s). If z is false the first clause forces p or q, and if z is true the second forces r or s, so together they force the original clause.

munotes.in427

Proving a Problem NP Complete: Three Reductions Worked

3. Why does a single literal clause need four new clauses and not one? Because the new variables must not be able to satisfy the clause by themselves. Writing a with both settings of y1 and both of y2, four clauses in all, means every setting of the new variables leaves at least one clause needing a, so a is forced true.

4. In the 3-satisfiability to clique reduction, why are two vertices from the same clause never joined? So that a clique of size m is forced to take exactly one vertex from each clause. If they were joined, a clique could sit inside one clause and would say nothing about the others.

5. Why can a clique in that graph never contain both a literal and its negation? Because the construction refuses to join complementary vertices, and every two vertices of a clique are joined. That is what makes the assignment it names consistent.

6. Reduce clique to vertex cover, and state both changes. Replace the graph by its complement and replace k by n minus k. A set of k vertices is a clique in G exactly when the other n minus k vertices cover the complement.

7. A proof reduces the new problem B to satisfiability and concludes B is NP complete. What has been proved? Only that B is in NP, and even that only if the reduction is polynomial and satisfiability's verifier is used. Nothing about hardness. The reduction must run from satisfiability to B.

Contents This chapter on its own page

munotes.in428

Chapter Eighty-Six

Complexity Hierarchies, and the P Against NP Question

Syllabus topic Module 2, "Computability and Complexity: Complexity Hierarchies"

In one line

The classes form a chain from tiny space to enormous time, more room or more time always buys more, and yet almost none of the neighbouring links in the chain has been proved strict.

The chain

Each class holds the one before it. Every containment on this list is known; almost none is known to be strict.

ClassThe machineContained in
Ldeterministic, log spaceNL
NLnondeterministic, log spaceP
Pdeterministic, polynomial timeNP
NPnondeterministic, polynomial timePSPACE
PSPACEdeterministic, polynomial spaceEXPTIME
EXPTIMEdeterministic, exponential timeEXPSPACE
EXPSPACEdeterministic, exponential space

Why NP is inside PSPACE. Try every certificate one at a time, reusing the same tape for each. There are exponentially many, so it takes a long time, and it needs only room for one certificate and the verifier's own working space, which is polynomial. Chapter 79's asymmetry is doing the work: a tape square can be used again and a moment cannot.

Why PSPACE is inside EXPTIME. A machine using p(n) squares over an alphabet of size a, with s states, has at most s times p(n) times a to the p(n) distinct configurations. A halting machine never repeats one, so it halts within that many steps, which is exponential.

Savitch's theorem, and where nondeterminism is cheap

Chapter 79 stated it: a nondeterministic machine using space S can be simulated by a deterministic one using space S squared. So:

PSPACE = NPSPACE, because the square of a polynomial is a polynomial

There is no such theorem for time. Removing nondeterminism from a time bound costs an exponential as far as anyone knows, and whether it must is the question this chapter is named after.

That is worth pausing on. For SPACE, the nondeterministic and deterministic classes are the same and it is a theorem. For TIME, the corresponding question is open and has a million dollar prize on it. The difference between the two is that space can be reused.

The hierarchy theorems, which make the chain worth drawing

Without these the chain might collapse entirely, and every class in it be the same class.

The time hierarchy theorem. Given enough more time, a machine can decide something no faster machine can. Roughly: if g grows sufficiently faster than f, then there is a language decidable in time g and not in time f. The proof is a diagonal argument in the shape of chapter 74's.

The space hierarchy theorem. The same for space, and with a cleaner condition, because space does not pay the simulation overhead that time does: if f grows more slowly than g, there is a language decidable in space g and not in space f.

munotes.in429

Complexity Hierarchies, and the P Against NP Question

What they give us. These are the ONLY separations known along the chain, and they separate classes that are far apart.

SeparationKnown?By what
L is properly inside PSPACEyesthe space hierarchy theorem
PSPACE is properly inside EXPSPACEyesthe space hierarchy theorem
P is properly inside EXPTIMEyesthe time hierarchy theorem
NP is properly inside NEXPTIMEyesthe nondeterministic time hierarchy theorem
L is properly inside Pnot known
P is properly inside NPnot knownthe question below
NP is properly inside PSPACEnot known
PSPACE is properly inside EXPTIMEnot known

The strangest fact in the table

We know the two ends of a chain differ and cannot say where.

P is inside NP, NP is inside PSPACE, PSPACE is inside EXPTIME, and P is NOT equal to EXPTIME

So at least one of those three containments is strict. At least one of them is a real step up. And nobody knows which. It is entirely possible, so far as anything proved goes, that P equals NP equals PSPACE and the single strict step is the last one.

That sentence is the honest summary of this topic, and a student who can write it has understood the chapter better than one who recites the chain.

The P against NP question

The statement. Is P equal to NP? That is: if a solution can be CHECKED in polynomial time, can one always be FOUND in polynomial time?

Both forms are the same question, by chapter 82's verifier definition, and the second is the one to quote because it says what is at stake.

What each answer would mean

If P equals NPIf P is not equal to NP
every NP complete problem has a polynomial algorithmnone of them does
finding is as easy as checking, everywherechecking is genuinely easier than finding
public key cryptography built on hard problems loses its groundthe present state of affairs is confirmed
finding a proof becomes about as easy as verifying onemathematics stays hard in the way it appears to be
scheduling, routing and packing problems become tractablethe heuristics we use are not a failure of cleverness

The cryptography row deserves a caution. P equals NP would not by itself break every cipher overnight; the polynomial might be of high degree, as chapter 81 warned. It would remove the FOUNDATION those systems rest on, which is the belief that certain problems cannot be done quickly.

What a proof would have to look like

To proveIt suffices to
P equals NPgive ONE polynomial algorithm for ONE NP complete problem
P is not equal to NPprove that ONE problem in NP has no polynomial algorithm
munotes.in430

Complexity Hierarchies, and the P Against NP Question

The first is easier to state and nobody has done it. The second requires proving that no algorithm of a certain kind exists, which is a far harder shape of claim: it is a statement about every possible program.

Its status

Open. The Clay Mathematics Institute lists it among its Millennium Problems, and its own page for the problem still carries the heading "Unsolved". That page records that Stephen Cook and Leonid Levin formulated the question independently in 1971.

The prize. The Institute's index of the seven problems says its Board designated a seven million dollar fund, one million allocated to each problem, and that the prizes were announced at a meeting in Paris on 24 May 2000 at the College de France.

Cook said as much himself, in the paper. The discussion section of his 1971 paper, on page 154 of the scan, calls the set of tautologies a good candidate for a set outside polynomial time, says he feels it is worth spending considerable effort trying to prove that conjecture, and that such a proof would be a major breakthrough in complexity theory. It is still a conjecture.

Distinctions

Time hierarchySpace hierarchy
Saysmore time decides moremore space decides more
Condition on the gapwider, because simulation costs timenarrower, because it costs no extra space
GivesP properly inside EXPTIMEL properly inside PSPACE
Methoddiagonalisationdiagonalisation
Nondeterminism in spaceNondeterminism in time
Cost of removing ita square, by Savitchexponential, as far as is known
So the classes arePSPACE equals NPSPACE, provedP against NP, open
Why the differencespace can be reuseda step cannot be taken twice

What it does NOT mean

The chain being known does not mean the classes are known to differ. Almost every neighbouring pair is open.

P not equal to EXPTIME does not settle P against NP. It says one of three links is strict and does not say which.

Savitch's theorem does not say nondeterminism is free. It says it costs a square, which is free only because polynomials survive squaring.

The hierarchy theorems do not separate P from NP. They compare a class with a class of the same kind and a larger bound; P and NP have the SAME bound and different machines, which is exactly the case they do not cover.

P equals NP would not mean every hard problem becomes easy. Undecidable problems stay undecidable, and the polynomial might have a degree that makes it useless in practice.

An algorithm for one NP complete problem would settle it. Not an algorithm that works usually, and not one that works on the instances people happen to meet. A polynomial worst case bound, proved.

munotes.in431

Complexity Hierarchies, and the P Against NP Question

Quick revision

  • The chain is L, NL, P, NP, PSPACE, EXPTIME, EXPSPACE, each inside the next.
  • NP is inside PSPACE because certificates can be tried one at a time in the same space, and PSPACE is inside

EXPTIME by counting configurations.

  • Savitch: PSPACE equals NPSPACE, because squaring a polynomial gives a polynomial. There is no such theorem

for time.

  • The hierarchy theorems, proved by diagonalisation, give the only separations known: L properly inside

PSPACE, P properly inside EXPTIME, PSPACE properly inside EXPSPACE.

  • P is not EXPTIME, and P is inside NP inside PSPACE inside EXPTIME, so at least one of those three

containments is strict and nobody knows which.

  • P against NP asks whether finding is as easy as checking. To prove equality, give one polynomial algorithm

for one NP complete problem. To prove inequality, prove no algorithm exists for one NP problem.

  • The Clay Mathematics Institute's page still reads Unsolved, and records Cook and Levin as having formulated

the question independently in 1971.

Test yourself

1. Write the chain of classes and say which containments are known to be strict. L inside NL inside P inside NP inside PSPACE inside EXPTIME inside EXPSPACE. The known strict ones are L inside PSPACE, P inside EXPTIME and PSPACE inside EXPSPACE, all by hierarchy theorems. No neighbouring pair on the list is known to be strict.

2. Why is NP inside PSPACE? Because the certificates can be tried one after another, reusing the same tape. That takes exponential time and only polynomial space, since a certificate and the verifier's workspace are both polynomial.

3. State Savitch's theorem and its consequence for PSPACE. A nondeterministic machine using space S can be simulated deterministically in space S squared. Since the square of a polynomial is a polynomial, PSPACE and NPSPACE are the same class.

4. What do the hierarchy theorems say, and what do they give? That sufficiently more time, or more space, decides strictly more languages. They give P properly inside EXPTIME and L properly inside PSPACE. They do not separate P from NP, because those two have the same bound and differ in the machine.

5. Explain the fact that at least one of three containments is strict. P is inside NP inside PSPACE inside EXPTIME, and the time hierarchy theorem shows P is not equal to EXPTIME. If all three containments were equalities then P would equal EXPTIME, which is false, so at least one is strict. Which one is unknown.

6. What would suffice to prove P equals NP? And to prove it does not? A single polynomial time algorithm for a single NP complete problem. For the other direction, a proof that some problem in NP has no polynomial time algorithm, which is a claim about every possible program and is much harder to make.

munotes.in432

Complexity Hierarchies, and the P Against NP Question

7. Name one consequence of P equalling NP, and one caution about it. Every NP complete problem would have a polynomial algorithm, and the hardness assumptions under public key cryptography would lose their ground. The caution is that the polynomial could be of a degree high enough to be useless in practice, so "polynomial" and "fast" are not the same word.

Contents This chapter on its own page

munotes.in433

Chapter Eighty-Seven

The Machines and the Grammars, Side by Side

Syllabus topic Modules 1 and 2 together

In one line

Four kinds of grammar, four kinds of machine, and they match up one for one.

The table to redraw from memory

If a student can produce this from memory and explain any cell, Q.3 is answered.

TypeGrammar restrictionMachineLanguage classSeparated from the level below by
3one variable on the left; the right side a terminal, a terminal then a variable, or epsilonfinite automatonregular
2one variable on the left; anything on the rightpushdown automatoncontext freea to the n, b to the n
1the right side is never shorter than the leftlinear bounded automatoncontext sensitivea to the n, b to the n, c to the n
0no restrictionTuring machinerecursively enumerablethe halting problem's language

Read the machine column downward. No memory, then a stack, then a tape no longer than the input, then an unbounded tape. Each level adds exactly one thing, and each addition buys exactly one step up the language hierarchy.

The same idea, four times over

Here is one language at each level, with the grammar, and with a machine or an expression beside it, so the match can be seen rather than recited.

Type 3: an even number of a's

S->bS | aA | ε
A->bA | aS

The machine reads the variables as states, which is what makes type 3 and the finite automaton the same thing:

H0Stateab
start finale0e1e0
e1e0e1

Accepts: ε, b, aa, abba, bb

Rejects: a, ab, ba, aaa, bab

And the same language as an expression:

(b + a b* a)*

L(H0) = L(H0R)

L(H3) = L(H0R)

Three descriptions, one language. Grammar, machine, expression. That equivalence is what chapter 40 called Kleene's theorem, and it holds only at this level.

Type 2: a to the n followed by b to the n

S->aSb | ε

Accepts: ε, ab, aabb, aaabbb

Rejects: a, b, ba, aab, abab

No finite automaton can do this, by chapter 36's pumping lemma: the machine would have to remember how many a's it had seen, and it has only finitely many states. The stack remembers:

start q0
stack Z
final qf
q0, a, Z -> q0, A Z
q0, a, A -> q0, A A
q0, b, A -> q1, ε
q1, b, A -> q1, ε
q1, ε, Z -> qf, Z
q0, ε, Z -> qf, Z

L(P1) = L(H2)

One A pushed per a, one popped per b, and the bottom marker reached exactly when the counts agree.

munotes.in434

The Machines and the Grammars, Side by Side

Type 1: a to the n, b to the n, c to the n

S->aSBc | abc
cB->Bc
bB->bb

Accepts: abc, aabbcc, aaabbbccc

Rejects: ε, a, ab, abcc, aabc, aabbc

Look at the second production. cB->Bc has two symbols on each side and a terminal on the left, which no context free grammar may have. That is precisely the extra power, and it is used to carry a B leftward past the c's until it meets a b and becomes one. No stack can do this, by chapter 50's pumping lemma for context free languages: matching three counts needs two things remembered at once, and a stack gives one.

Type 0: the language nothing below can accept

At the top there is no grammar to print that a machine here could check, and the reason is the point. Chapter 74's language, the set of machine and input pairs for which the machine halts, has a type 0 grammar and no type 1 one. A linear bounded automaton always halts, because it has finitely many configurations and can be stopped when one repeats, so every context sensitive language is decidable. The halting problem's language is not.

What each machine can and cannot do

Finite automatonPushdown automatonLinear bounded automatonTuring machine
Memorynone, beyond the stateone stacktape, no longer than the inputunbounded tape
Can count one thingnoyesyesyes
Can count two things at oncenonoyesyes
Always haltsyesyes, if the epsilon moves are controlledyesno
Nondeterminism adds powernoYESnot knownno
Decides its own membership problemyes, in linear timeyes, in cubic time by CYKyes, in exponential timeno

Row 5 is the one worth memorising, because the answer is different in every column. For finite automata the subset construction of chapter 16 removes nondeterminism. For pushdown automata it cannot be removed, and the language of even length palindromes is the standard witness. For linear bounded automata it is an open question, which is where chapter 61 left it. For Turing machines the simulation of chapter 69 removes it at an exponential cost in time, and the class of languages is unchanged.

Closure, across all four

Yes means the class is closed under the operation: applying it to members gives a member.

OperationRegularContext freeContext sensitiveRecursiveRecursively enumerable
Unionyesyesyesyesyes
Concatenationyesyesyesyesyes
Kleene closureyesyesyesyesyes
IntersectionyesNOyesyesyes
ComplementyesNOyesyesNO
Intersection with a regular languageyesyesyesyesyes
Reversalyesyesyesyesyes

Two cells decide most questions on this table.

munotes.in435

The Machines and the Grammars, Side by Side

Context free languages are not closed under intersection, and the standard witness is a to the n, b to the n, c to the m intersected with a to the n, b to the m, c to the m: both are context free and their intersection is a to the n, b to the n, c to the n, which chapter 50 shows is not. Complement then fails too, because closure under complement and union would give closure under intersection.

Recursively enumerable languages are not closed under complement, and that is chapter 71's theorem in another form: a language and its complement both being recursively enumerable makes the language recursive. So the complement of the halting problem's language is not even recursively enumerable.

Context sensitive languages ARE closed under complement. That was open for twenty four years and was settled in 1988, which chapter 61 records from the paper itself.

Decision problems, across all four

Yes means an algorithm exists that always gives the right answer and halts.

Question about the languageRegularContext freeContext sensitiveRecursively enumerable
Is this string in it?yesyesyesNO
Is it empty?yesyesNONO
Is it finite?yesyesNONO
Is it everything?yesNONONO
Do these two agree?yesNONONO

Read the table as a staircase. Everything is decidable for regular languages. Membership survives to the context sensitive level. Emptiness dies one level earlier than membership. Equivalence dies first of all, at the context free level, and it is undecidable there even though membership is easy.

The membership row's last cell is only half a no. Membership in a recursively enumerable language is SEMI decidable: run the machine, and if the string is in the language it will say so. It is the no answers that never come, which is chapter 70's distinction.

Where Module 2's second half fits

The hierarchy above is about what can be done AT ALL. Module 2's complexity classes sit inside its bottom rows and ask what can be done QUICKLY.

BoundarySeparatesSettled?
finite memory against a stackregular from context freeyes, by the pumping lemma
a stack against a bounded tapecontext free from context sensitiveyes, by the pumping lemma for context free languages
a bounded tape against an unbounded onedecidable from recursively enumerableyes, by the halting problem
a machine that halts against one that may notrecursive from recursively enumerableyes, chapter 71
polynomial time against the restP from NPNO, chapter 86

Three of the subject's boundaries are proved and one is not. The unproved one is the newest, and it is the only open question in this book that a student could in principle settle.

munotes.in436

The Machines and the Grammars, Side by Side

Distinctions

Chomsky hierarchyComplexity hierarchy
Askscan it be donehow much does it cost
Resourcethe kind of memorythe amount of time or space
Levels4many
All separations provedyesno, almost none
Machine variesbetween levelsnot at all, only the bound
RegularContext freeContext sensitiveRecursively enumerable
Grammartype 3type 2type 1type 0
Machinefinite automatonpushdown automatonlinear bounded automatonTuring machine
Closed under complementyesnoyesno
Membership decidableyesyesyesno
Equivalence decidableyesnonono

What it does NOT mean

The four classes are not four separate subjects. Each contains the one before it, properly, and the witnesses in the first table prove each containment strict.

Type 2 is not literally a subset of type 1 as sets of GRAMMARS. A context free grammar with an epsilon production shortens, which type 1 forbids. The containment is between the classes of LANGUAGES, and it holds because the epsilon productions can be removed by chapter 46's method.

A linear bounded automaton is not a Turing machine with a short tape by accident. The bound is part of the definition, and it is what makes membership decidable.

Undecidable is not the same as hard. Equivalence of context free grammars is undecidable; deciding satisfiability is decidable and merely expensive. They are failures of different kinds.

The pumping lemmas do not prove a language IS regular or context free. They prove it is not. Chapter 37 gives the reason: they state a necessary condition, never a sufficient one.

Quick revision

  • Type 3 with a finite automaton, type 2 with a pushdown automaton, type 1 with a linear bounded automaton,

type 0 with a Turing machine. Each machine adds one thing: a stack, then a bounded tape, then an unbounded one.

  • The separating witnesses are a to the n b to the n, then a to the n b to the n c to the n, then the halting

problem's language.

  • Nondeterminism adds nothing to a finite automaton, adds power to a pushdown automaton, is an open question

for a linear bounded automaton and adds nothing to a Turing machine.

  • Context free languages are closed under union, concatenation, closure and intersection WITH A REGULAR

LANGUAGE, and not under intersection or complement.

  • Recursively enumerable languages are closed under everything but complement.
  • Membership is decidable down to the context sensitive level; emptiness and finiteness die one level higher;

equivalence dies at the context free level.

  • Three of the subject's four boundaries are proved. P against NP is the fourth and is open.

Test yourself

1. Draw the four by four correspondence. Type 3 with the finite automaton and the regular languages; type 2 with the pushdown automaton and the context free languages; type 1 with the linear bounded automaton and the context sensitive languages; type 0 with the Turing machine and the recursively enumerable languages.

munotes.in437

The Machines and the Grammars, Side by Side

2. Give a language at each boundary and say what it proves. A to the n b to the n is context free and not regular, so a stack beats finite memory. A to the n b to the n c to the n is context sensitive and not context free, so a bounded tape beats a stack. The halting problem's language is recursively enumerable and not recursive, so an unbounded tape reaches languages no halting machine does.

3. In which of the four machines does nondeterminism add power? Only the pushdown automaton, as far as is known. It adds nothing to the finite automaton, by the subset construction, and nothing to the Turing machine, by simulation at exponential cost. For the linear bounded automaton the question is open.

4. Show that the context free languages are not closed under intersection. Take a to the n b to the n c to the m and a to the n b to the m c to the m. Both are context free, since each matches one pair of counts and leaves the third free. Their intersection forces all three counts equal, and that language is not context free.

5. Why does closure under complement then fail as well? Because a class closed under union and complement is closed under intersection, by De Morgan. The context free languages are closed under union and not under intersection, so they cannot be closed under complement.

6. Which decision problems are undecidable for context free languages? Equivalence, whether the language is everything, and by extension whether two grammars generate the same language. Membership, emptiness and finiteness are all decidable.

7. Which boundary in this subject is still open, and what would settle it? P against NP. A polynomial time algorithm for a single NP complete problem would settle it one way, and a proof that some problem in NP has no polynomial time algorithm would settle it the other.

Contents This chapter on its own page

munotes.in438

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!