munotes®

Practical 2: Process Communication with Pipes

Get access to whole semester resourcesSemester Pass

Chapter Six

Syllabus topic Module 1, "Process Communication using Message Passing: Use message queues/pipes to solve the producer-consumer problem."

Pages 40 to 47 of 300

Aim

To pass data between processes by message passing rather than shared memory, using an anonymous pipe and a named pipe, and to solve the producer-consumer problem over a pipe.

What you need to know before you start

Shared memory, in [Practical 1: Process Communication using Shared Memory], gave two processes one piece of memory and left the order entirely to them. Message passing is the other way of doing it: the data is handed to the kernel by one process and taken from the kernel by the other, and the kernel does the waiting.

That difference is the whole of Practical 2, and it has one large consequence:

With message passing you get synchronisation for free. A reader that asks for data before there

is any simply waits, and it is the kernel that puts it to sleep and wakes it up. No semaphore is

needed.

The price is that the data is copied: once into the kernel and once out again. Shared memory copies nothing. That is why shared memory is faster and why it is harder to get right.

A pipe

A pipe is a one-way channel between two processes, held in the kernel, with a fixed amount of room in it. One end is written and the other is read. It is the oldest form of inter-process communication on Unix and it is what the shell's | is made of: ls | wc -l is two processes and one pipe.

Two kinds, and MU's exercise can be answered with either:

Anonymous pipeNamed pipe, also called a FIFO
Made bypipe()mkfifo() or the mkfifo command
Has a namenoyes, it is a file in a directory
Who can use itthis process and its childrenany process that can open the file
Lives untilthe last end is closedthe file is deleted

The five properties of a pipe that matter in the examination

  1. It is one way. Data written at one end comes out at the other and not back. Two processes

that must both talk and listen need two pipes.

  1. It is a byte stream. The pipe does not remember where one write ended and the next began.

This page proves it below.

  1. A read with nothing in the pipe blocks, which is the free synchronisation.
  2. A read when the pipe is empty and every writer has closed it returns 0, which is end of

file. That is how a reader knows to stop.

  1. A write to a pipe with no reader left kills the writer with the signal SIGPIPE, or, if

that signal is ignored, fails with EPIPE.

munotes.in40

Practical 2: Process Communication with Pipes

The program: a parent and a child, one pipe

#include <stdio.h>
#include <string.h>
#include <sys/types.h>
#include <unistd.h>
#include <sys/wait.h>

int main(void)
{
    int fd[2];

    if (pipe(fd) == -1) {
        perror("pipe");
        return 1;
    }

    pid_t pid = fork();
    if (pid == -1) {
        perror("fork");
        return 1;
    }

    if (pid == 0) {                       /* the child writes */
        close(fd[0]);                     /* it never reads: close that end */
        const char *msg = "hello down the pipe";
        write(fd[1], msg, strlen(msg) + 1);
        close(fd[1]);
        _exit(0);
    }

    close(fd[1]);                         /* the parent never writes */
    char buf[64];
    ssize_t n = read(fd[0], buf, sizeof buf);
    printf("parent read %zd bytes: [%s]\n", n, buf);
    close(fd[0]);
    wait(NULL);
    return 0;
}
parent read 20 bytes: [hello down the pipe]

Reading the program

pipe(fd) makes the pipe and fills the array with two file descriptors. The convention is fixed and worth memorising: fd[0] is the read end, fd[1] is the write end. Zero for reading because standard input is 0; one for writing because standard output is 1.

fork after pipe, never before. The child inherits the two descriptors because they were open when it was created. A pipe made after the fork exists only in the process that made it.

Each side closes the end it does not use, and this is not tidiness. The parent has fd[1] open, so as far as the kernel is concerned there is still a writer, and when the child closes its copy the pipe does not reach end of file. A parent that forgets close(fd[1]) and reads in a loop waits for ever for data from itself. This single line is the commonest cause of a pipe program that hangs.

write(fd[1], msg, strlen(msg) + 1) writes the string and the + 1 writes the terminating zero byte as well, which is why the parent can print buf as a string and why read returned 20 for a 19-character message. A pipe carries bytes and knows nothing about strings.

read(fd[0], buf, sizeof buf) returns how many bytes it actually got, which may be fewer than you asked for. It returns -1 on error and 0 at end of file.

A pipe is a byte stream, and this is the surprise

The next program has the child write three lines with three separate calls to write, and the parent read in a loop until end of file.

#include <stdio.h>
#include <sys/types.h>
#include <unistd.h>
#include <sys/wait.h>

int main(void)
{
    int fd[2];
    if (pipe(fd) == -1) {
        perror("pipe");
        return 1;
    }

    if (fork() == 0) {
        close(fd[0]);
        for (int i = 1; i <= 3; i++) {
            char line[32];
            int len = snprintf(line, sizeof line, "item %d\n", i * 10);
            write(fd[1], line, (size_t) len);       /* three writes */
        }
        close(fd[1]);
        _exit(0);
    }

    close(fd[1]);
    char buf[128];
    ssize_t n;
    while ((n = read(fd[0], buf, sizeof buf)) > 0)  /* how many reads? */
        printf("consumer got %zd bytes:\n%.*s", n, (int) n, buf);

    printf("read returned %zd, which is end of file\n", n);
    close(fd[0]);
    wait(NULL);
    return 0;
}
munotes.in41

Practical 2: Process Communication with Pipes

consumer got 24 bytes:
item 10
item 20
item 30
read returned 0, which is end of file

One read, not three. Three writes of 8 bytes each arrived as a single read of 24, on all ten runs. The pipe put the bytes in order and threw away every boundary between them.

That is what "byte stream" means, and it has three consequences a student should be able to state:

  • A reader cannot tell how many writes there were. It gets bytes, and it must decide where one

message ends. Here the newline does that job; in a real program it is a length written in front of the data, or a fixed-size record.

  • A read may also return less than one write. The opposite case happens with large writes: ask

for 128 bytes and you may get 40, because that is all that had arrived. This is why the read is in a while loop, and why a single read outside one is a bug waiting for a slow writer.

  • This is exactly what a message queue fixes. A message queue keeps each message whole, and

that is the subject of [Practical 2 continued: Message Queues, Blocking and Non-blocking].

%.*s in that printf prints exactly n characters of buf and no more. Printing buf as a plain %s would read past the bytes that arrived, because nothing in the pipe wrote a terminating zero this time.

Producer and consumer over a pipe, which is MU's own wording

Compare this with the semaphore version in [Practical 1 continued: the Race Condition, Semaphores, and Producer and Consumer]. It does the same job and there is no semaphore in it, because the pipe blocks for you.

#include <stdio.h>
#include <sys/types.h>
#include <unistd.h>
#include <sys/wait.h>

#define ITEMS 5

int main(void)
{
    int fd[2];
    if (pipe(fd) == -1) {
        perror("pipe");
        return 1;
    }

    if (fork() == 0) {                          /* the producer */
        close(fd[0]);
        for (int i = 1; i <= ITEMS; i++) {
            int item = i * i;
            if (write(fd[1], &item, sizeof item) != (ssize_t) sizeof item) {
                perror("producer write");
                _exit(1);
            }
            printf("producer: produced %d\n", item);
            fflush(stdout);
        }
        close(fd[1]);                           /* tells the consumer to stop */
        _exit(0);
    }

    close(fd[1]);                               /* the consumer */
    int item;
    ssize_t n;
    int total = 0;
    while ((n = read(fd[0], &item, sizeof item)) == (ssize_t) sizeof item) {
        printf("consumer: consumed %d\n", item);
        total = total + item;
    }
    if (n == -1)
        perror("consumer read");
    printf("consumer: the pipe is closed, the total is %d\n", total);

    close(fd[0]);
    wait(NULL);
    return 0;
}
munotes.in42

Practical 2: Process Communication with Pipes

producer: produced 1
producer: produced 4
producer: produced 9
producer: produced 16
producer: produced 25
consumer: consumed 1
consumer: consumed 4
consumer: consumed 9
consumer: consumed 16
consumer: consumed 25
consumer: the pipe is closed, the total is 55

Three things to take from that output.

The producer got all five items in before the consumer took the first one out. That is not a fault: the pipe has room, in Linux 64 kibibytes by default, and 20 bytes fits easily. The producer only waits when the pipe is full, and the consumer only waits when it is empty. Compare the one-slot semaphore version, where the two alternated strictly because the buffer held exactly one item. A pipe is a bounded buffer that the kernel provides ready-made, and its bound is the pipe's capacity.

An int was sent as four raw bytes, with write(fd[1], &item, sizeof item), and read back the same way. That works because both processes are the same program on the same machine, so the two agree about how an int is laid out. It would not be safe between machines, and it is the reason network programs convert to a defined byte order first.

close(fd[1]) in the producer is what ends the consumer's loop. Without it the consumer's read would block for ever after the fifth item, waiting for a writer that has finished but not said so. The total of 55 at the end is 1 plus 4 plus 9 plus 16 plus 25, and it is there so that the program proves nothing was lost.

A named pipe, which two unrelated programs can use

An anonymous pipe reaches only your own children. A named pipe, or FIFO, is a file in a directory, and any program that can open that file can use it. It is the pipe equivalent of giving a shared memory segment a key.

#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#include <fcntl.h>
#include <sys/stat.h>
#include <sys/types.h>
#include <sys/wait.h>

#define PATH "/tmp/munotes_fifo"

int main(void)
{
    unlink(PATH);                          /* in case one is left over */

    if (mkfifo(PATH, 0600) == -1) {
        perror("mkfifo");
        return 1;
    }
    printf("made the named pipe %s\n", PATH);
    fflush(stdout);
    system("ls -l /tmp/munotes_fifo | cut -c1-11");

    if (fork() == 0) {
        int fd = open(PATH, O_WRONLY);     /* blocks until a reader arrives */
        if (fd == -1) {
            perror("child open");
            _exit(1);
        }
        const char *m = "sent through a named pipe";
        write(fd, m, strlen(m) + 1);
        close(fd);
        _exit(0);
    }

    int fd = open(PATH, O_RDONLY);         /* blocks until a writer arrives */
    if (fd == -1) {
        perror("parent open");
        return 1;
    }
    char buf[64];
    ssize_t n = read(fd, buf, sizeof buf);
    printf("parent read %zd bytes: [%s]\n", n, buf);
    close(fd);

    wait(NULL);
    unlink(PATH);                          /* a FIFO must be deleted */
    return 0;
}
munotes.in43

Practical 2: Process Communication with Pipes

made the named pipe /tmp/munotes_fifo
prw-------
parent read 26 bytes: [sent through a named pipe]

The second line of that output is the evidence that a FIFO is a file of its own kind. ls -l prints a letter for the type in the first column, and here it is p, for pipe. An ordinary file prints -, a directory d, a symbolic link l.

Two behaviours of a FIFO are worth knowing and are the reason this program forks:

  • open for reading blocks until a writer opens it, and open for writing blocks until a reader does.

That is how the two sides find each other, and it means the writer cannot simply run first and finish. In the laboratory you run the writer in one terminal and the reader in another and watch the first one wait.

  • A FIFO is a file and has to be deleted, with unlink in a program or rm at the shell.

Like a shared memory segment, it outlives the program that made it; unlike a segment, it shows up in ls.

You can try the same thing with no C at all, which is worth doing once because it makes the idea concrete: mkfifo /tmp/f, then cat < /tmp/f in one terminal and echo hello > /tmp/f in another.

Writing to a pipe nobody is reading

The last of the five properties, and the one that produces a mysterious death.

#include <stdio.h>
#include <string.h>
#include <errno.h>
#include <signal.h>
#include <unistd.h>

int main(void)
{
    signal(SIGPIPE, SIG_IGN);              /* so we live to print the error */

    int fd[2];
    if (pipe(fd) == -1) {
        perror("pipe");
        return 1;
    }

    close(fd[0]);                          /* there is now no reader at all */

    ssize_t n = write(fd[1], "x", 1);
    printf("write returned %zd, errno says %s\n", n, strerror(errno));
    return 0;
}
write returned -1, errno says Broken pipe

Without the signal(SIGPIPE, SIG_IGN) line that program would not print anything: the kernel sends SIGPIPE to a process that writes to a pipe with no reader, and the default action for SIGPIPE is to kill it. The program simply disappears, with no message, and a student looks for a bug in the loop above.

That is also why ls | head -1 does not print an error when head stops reading: ls is killed by SIGPIPE, which is the intended design.

Pipes against shared memory

PipeShared memory
Who does the waitingthe kernel, automaticallyyou, with semaphores
Data is copiedtwice, in and outnever
Directionone way per pipeboth ways
Keeps message boundariesno, it is a byte streamthere are no messages
Related processes onlyanonymous yes, named nono, a key is enough
Cleaning upclosing the descriptors, and unlink for a FIFOshmctl with IPC_RMID
Blocks whenfull on write, empty on readnever
munotes.in44

Practical 2: Process Communication with Pipes

The full comparison, with message queues in it as well, is in [Practical 2 continued: Message Queues, Blocking and Non-blocking].

Procedure

  1. Write the first program. Compile with gcc -Wall -Wextra -o pipe1 pipe1.c and run it.
  2. Remove the parent's close(fd[1]), put the parent's read in a while loop, and run it again.

The program hangs. Press Ctrl+C and put the line back.

  1. Write the three-write program and count the reads. Note that the boundaries are gone.
  2. Write the producer-consumer program and confirm the total.
  3. Write the FIFO program. Run ls -l /tmp/munotes_fifo yourself while it is running and note the

p.

  1. At the shell: mkfifo /tmp/f, then cat < /tmp/f in one terminal and echo hello > /tmp/f in

another. Then rm /tmp/f.

Result

Two processes exchanged data through an anonymous pipe with no semaphore, because a pipe blocks on an empty read and on a full write. Three separate writes were shown to arrive as one read, which establishes that a pipe is a byte stream with no message boundaries. The producer-consumer problem was solved over a pipe and the total proved nothing was lost. A named pipe was created, seen in ls -l as type p, used between two processes and deleted.

Where marks are lost

  • Not closing the unused end. The reader keeps a write end open, end of file never arrives,

and the program hangs. This is the single commonest fault in this exercise.

  • pipe after fork. The two processes then have two different pipes.
  • Mixing up fd[0] and fd[1]. Zero reads, one writes.
  • Reading once instead of in a loop, and assuming one read returns one write.
  • Printing the buffer with %s when nothing wrote a terminating zero. Use %.*s and the

count that read returned.

  • Forgetting to unlink the FIFO, so the next run fails with EEXIST from mkfifo.
  • Not knowing why a program writing to a closed pipe dies silently. It is SIGPIPE.

For the journal

Write the aim, MU's own wording, the first program and its output, and then the producer-consumer program with its output and the total. Add the three-write experiment and one sentence saying how many reads it took and why. If your college asks for the named pipe as well, add the mkfifo program and the ls -l line showing the p. The conclusion: a pipe gives message passing with the synchronisation done by the kernel, at the cost of copying the data and of losing the boundaries between writes.

munotes.in45

Practical 2: Process Communication with Pipes

Quick revision

  • A pipe is a one-way channel in the kernel. pipe(fd) gives fd[0] to read and fd[1] to

write. Make it before the fork.

  • Each process closes the end it does not use, or end of file never arrives.
  • read returns the number of bytes it got, 0 at end of file, -1 on error.
  • A pipe is a byte stream: three writes may arrive as one read, and one write may arrive as

several reads. The reader has to find the message boundaries itself.

  • A read on an empty pipe blocks and a write to a full pipe blocks, so a pipe is a bounded buffer

the kernel provides. No semaphore is needed.

  • Writing to a pipe with no reader raises SIGPIPE, which by default kills the writer; ignoring

the signal turns it into EPIPE instead.

  • A named pipe is made with mkfifo, is a file whose ls -l type is p, works between unrelated

programs, blocks on open until the other side arrives, and must be deleted.

  • Message passing copies the data twice and synchronises for free. Shared memory copies nothing

and synchronises not at all.

Questions you should be able to answer

1. Which of fd[0] and fd[1] is the read end? fd[0]. It matches standard input being 0 and standard output being 1.

2. Why must pipe be called before fork? Because the child inherits the descriptors that were open when it was created. A pipe made afterwards exists only in the process that made it.

3. A program reads from a pipe in a loop and never finishes, although the writer has ended. Why? The reader still has the write end open, so the kernel does not report end of file. Each side must close the end it does not use.

4. Three writes of eight bytes were made. How many reads does the other side need? It cannot be known. A pipe is a byte stream with no boundaries: in the run on this page one read of 24 bytes took all three. The reader must delimit messages itself.

5. What does read return at end of file, and when does that happen? Zero, and it happens when the pipe is empty and every write end has been closed.

6. What happens when a process writes to a pipe with no reader? It is sent SIGPIPE, which by default kills it with no message. If the signal is ignored or handled, the write fails and sets errno to EPIPE.

munotes.in46

Practical 2: Process Communication with Pipes

7. Give two differences between an anonymous pipe and a named pipe. An anonymous pipe has no name and reaches only related processes; a named pipe is a file in a directory, so any program that can open it may use it, and it has to be deleted afterwards.

8. Why does the producer-consumer over a pipe need no semaphore, when the shared memory version needs two? Because the pipe itself blocks: a read on an empty pipe waits and a write to a full pipe waits, and the kernel does the sleeping and waking. Shared memory does nothing of the kind.

munotes.in47

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Issue
Done!