Practical 2: Process Communication with Pipes
Chapter Six
Syllabus topic Module 1, "Process Communication using Message Passing: Use message queues/pipes to solve the producer-consumer problem."
Pages 40 to 47 of 300
Aim
To pass data between processes by message passing rather than shared memory, using an anonymous pipe and a named pipe, and to solve the producer-consumer problem over a pipe.
What you need to know before you start
Shared memory, in [Practical 1: Process Communication using Shared Memory], gave two processes one piece of memory and left the order entirely to them. Message passing is the other way of doing it: the data is handed to the kernel by one process and taken from the kernel by the other, and the kernel does the waiting.
That difference is the whole of Practical 2, and it has one large consequence:
With message passing you get synchronisation for free. A reader that asks for data before there
is any simply waits, and it is the kernel that puts it to sleep and wakes it up. No semaphore is
needed.
The price is that the data is copied: once into the kernel and once out again. Shared memory copies nothing. That is why shared memory is faster and why it is harder to get right.
A pipe
A pipe is a one-way channel between two processes, held in the kernel, with a fixed amount of room in it. One end is written and the other is read. It is the oldest form of inter-process communication on Unix and it is what the shell's | is made of: ls | wc -l is two processes and one pipe.
Two kinds, and MU's exercise can be answered with either:
| Anonymous pipe | Named pipe, also called a FIFO | |
|---|---|---|
| Made by | pipe() | mkfifo() or the mkfifo command |
| Has a name | no | yes, it is a file in a directory |
| Who can use it | this process and its children | any process that can open the file |
| Lives until | the last end is closed | the file is deleted |
The five properties of a pipe that matter in the examination
- It is one way. Data written at one end comes out at the other and not back. Two processes
that must both talk and listen need two pipes.
- It is a byte stream. The pipe does not remember where one
writeended and the next began.
This page proves it below.
- A read with nothing in the pipe blocks, which is the free synchronisation.
- A read when the pipe is empty and every writer has closed it returns 0, which is end of
file. That is how a reader knows to stop.
- A write to a pipe with no reader left kills the writer with the signal
SIGPIPE, or, if
that signal is ignored, fails with EPIPE.
Practical 2: Process Communication with Pipes
The program: a parent and a child, one pipe
#include <stdio.h>
#include <string.h>
#include <sys/types.h>
#include <unistd.h>
#include <sys/wait.h>
int main(void)
{
int fd[2];
if (pipe(fd) == -1) {
perror("pipe");
return 1;
}
pid_t pid = fork();
if (pid == -1) {
perror("fork");
return 1;
}
if (pid == 0) { /* the child writes */
close(fd[0]); /* it never reads: close that end */
const char *msg = "hello down the pipe";
write(fd[1], msg, strlen(msg) + 1);
close(fd[1]);
_exit(0);
}
close(fd[1]); /* the parent never writes */
char buf[64];
ssize_t n = read(fd[0], buf, sizeof buf);
printf("parent read %zd bytes: [%s]\n", n, buf);
close(fd[0]);
wait(NULL);
return 0;
}parent read 20 bytes: [hello down the pipe]Reading the program
pipe(fd) makes the pipe and fills the array with two file descriptors. The convention is fixed and worth memorising: fd[0] is the read end, fd[1] is the write end. Zero for reading because standard input is 0; one for writing because standard output is 1.
fork after pipe, never before. The child inherits the two descriptors because they were open when it was created. A pipe made after the fork exists only in the process that made it.
Each side closes the end it does not use, and this is not tidiness. The parent has fd[1] open, so as far as the kernel is concerned there is still a writer, and when the child closes its copy the pipe does not reach end of file. A parent that forgets close(fd[1]) and reads in a loop waits for ever for data from itself. This single line is the commonest cause of a pipe program that hangs.
write(fd[1], msg, strlen(msg) + 1) writes the string and the + 1 writes the terminating zero byte as well, which is why the parent can print buf as a string and why read returned 20 for a 19-character message. A pipe carries bytes and knows nothing about strings.
read(fd[0], buf, sizeof buf) returns how many bytes it actually got, which may be fewer than you asked for. It returns -1 on error and 0 at end of file.
A pipe is a byte stream, and this is the surprise
The next program has the child write three lines with three separate calls to write, and the parent read in a loop until end of file.
#include <stdio.h>
#include <sys/types.h>
#include <unistd.h>
#include <sys/wait.h>
int main(void)
{
int fd[2];
if (pipe(fd) == -1) {
perror("pipe");
return 1;
}
if (fork() == 0) {
close(fd[0]);
for (int i = 1; i <= 3; i++) {
char line[32];
int len = snprintf(line, sizeof line, "item %d\n", i * 10);
write(fd[1], line, (size_t) len); /* three writes */
}
close(fd[1]);
_exit(0);
}
close(fd[1]);
char buf[128];
ssize_t n;
while ((n = read(fd[0], buf, sizeof buf)) > 0) /* how many reads? */
printf("consumer got %zd bytes:\n%.*s", n, (int) n, buf);
printf("read returned %zd, which is end of file\n", n);
close(fd[0]);
wait(NULL);
return 0;
}Practical 2: Process Communication with Pipes
consumer got 24 bytes:
item 10
item 20
item 30
read returned 0, which is end of fileOne read, not three. Three writes of 8 bytes each arrived as a single read of 24, on all ten runs. The pipe put the bytes in order and threw away every boundary between them.
That is what "byte stream" means, and it has three consequences a student should be able to state:
- A reader cannot tell how many writes there were. It gets bytes, and it must decide where one
message ends. Here the newline does that job; in a real program it is a length written in front of the data, or a fixed-size record.
- A read may also return less than one write. The opposite case happens with large writes: ask
for 128 bytes and you may get 40, because that is all that had arrived. This is why the read is in a while loop, and why a single read outside one is a bug waiting for a slow writer.
- This is exactly what a message queue fixes. A message queue keeps each message whole, and
that is the subject of [Practical 2 continued: Message Queues, Blocking and Non-blocking].
%.*s in that printf prints exactly n characters of buf and no more. Printing buf as a plain %s would read past the bytes that arrived, because nothing in the pipe wrote a terminating zero this time.
Producer and consumer over a pipe, which is MU's own wording
Compare this with the semaphore version in [Practical 1 continued: the Race Condition, Semaphores, and Producer and Consumer]. It does the same job and there is no semaphore in it, because the pipe blocks for you.
#include <stdio.h>
#include <sys/types.h>
#include <unistd.h>
#include <sys/wait.h>
#define ITEMS 5
int main(void)
{
int fd[2];
if (pipe(fd) == -1) {
perror("pipe");
return 1;
}
if (fork() == 0) { /* the producer */
close(fd[0]);
for (int i = 1; i <= ITEMS; i++) {
int item = i * i;
if (write(fd[1], &item, sizeof item) != (ssize_t) sizeof item) {
perror("producer write");
_exit(1);
}
printf("producer: produced %d\n", item);
fflush(stdout);
}
close(fd[1]); /* tells the consumer to stop */
_exit(0);
}
close(fd[1]); /* the consumer */
int item;
ssize_t n;
int total = 0;
while ((n = read(fd[0], &item, sizeof item)) == (ssize_t) sizeof item) {
printf("consumer: consumed %d\n", item);
total = total + item;
}
if (n == -1)
perror("consumer read");
printf("consumer: the pipe is closed, the total is %d\n", total);
close(fd[0]);
wait(NULL);
return 0;
}Practical 2: Process Communication with Pipes
producer: produced 1
producer: produced 4
producer: produced 9
producer: produced 16
producer: produced 25
consumer: consumed 1
consumer: consumed 4
consumer: consumed 9
consumer: consumed 16
consumer: consumed 25
consumer: the pipe is closed, the total is 55Three things to take from that output.
The producer got all five items in before the consumer took the first one out. That is not a fault: the pipe has room, in Linux 64 kibibytes by default, and 20 bytes fits easily. The producer only waits when the pipe is full, and the consumer only waits when it is empty. Compare the one-slot semaphore version, where the two alternated strictly because the buffer held exactly one item. A pipe is a bounded buffer that the kernel provides ready-made, and its bound is the pipe's capacity.
An int was sent as four raw bytes, with write(fd[1], &item, sizeof item), and read back the same way. That works because both processes are the same program on the same machine, so the two agree about how an int is laid out. It would not be safe between machines, and it is the reason network programs convert to a defined byte order first.
close(fd[1]) in the producer is what ends the consumer's loop. Without it the consumer's read would block for ever after the fifth item, waiting for a writer that has finished but not said so. The total of 55 at the end is 1 plus 4 plus 9 plus 16 plus 25, and it is there so that the program proves nothing was lost.
A named pipe, which two unrelated programs can use
An anonymous pipe reaches only your own children. A named pipe, or FIFO, is a file in a directory, and any program that can open that file can use it. It is the pipe equivalent of giving a shared memory segment a key.
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <unistd.h>
#include <fcntl.h>
#include <sys/stat.h>
#include <sys/types.h>
#include <sys/wait.h>
#define PATH "/tmp/munotes_fifo"
int main(void)
{
unlink(PATH); /* in case one is left over */
if (mkfifo(PATH, 0600) == -1) {
perror("mkfifo");
return 1;
}
printf("made the named pipe %s\n", PATH);
fflush(stdout);
system("ls -l /tmp/munotes_fifo | cut -c1-11");
if (fork() == 0) {
int fd = open(PATH, O_WRONLY); /* blocks until a reader arrives */
if (fd == -1) {
perror("child open");
_exit(1);
}
const char *m = "sent through a named pipe";
write(fd, m, strlen(m) + 1);
close(fd);
_exit(0);
}
int fd = open(PATH, O_RDONLY); /* blocks until a writer arrives */
if (fd == -1) {
perror("parent open");
return 1;
}
char buf[64];
ssize_t n = read(fd, buf, sizeof buf);
printf("parent read %zd bytes: [%s]\n", n, buf);
close(fd);
wait(NULL);
unlink(PATH); /* a FIFO must be deleted */
return 0;
}Practical 2: Process Communication with Pipes
made the named pipe /tmp/munotes_fifo
prw-------
parent read 26 bytes: [sent through a named pipe]The second line of that output is the evidence that a FIFO is a file of its own kind. ls -l prints a letter for the type in the first column, and here it is p, for pipe. An ordinary file prints -, a directory d, a symbolic link l.
Two behaviours of a FIFO are worth knowing and are the reason this program forks:
openfor reading blocks until a writer opens it, andopenfor writing blocks until a reader does.
That is how the two sides find each other, and it means the writer cannot simply run first and finish. In the laboratory you run the writer in one terminal and the reader in another and watch the first one wait.
- A FIFO is a file and has to be deleted, with
unlinkin a program orrmat the shell.
Like a shared memory segment, it outlives the program that made it; unlike a segment, it shows up in ls.
You can try the same thing with no C at all, which is worth doing once because it makes the idea concrete: mkfifo /tmp/f, then cat < /tmp/f in one terminal and echo hello > /tmp/f in another.
Writing to a pipe nobody is reading
The last of the five properties, and the one that produces a mysterious death.
#include <stdio.h>
#include <string.h>
#include <errno.h>
#include <signal.h>
#include <unistd.h>
int main(void)
{
signal(SIGPIPE, SIG_IGN); /* so we live to print the error */
int fd[2];
if (pipe(fd) == -1) {
perror("pipe");
return 1;
}
close(fd[0]); /* there is now no reader at all */
ssize_t n = write(fd[1], "x", 1);
printf("write returned %zd, errno says %s\n", n, strerror(errno));
return 0;
}write returned -1, errno says Broken pipeWithout the signal(SIGPIPE, SIG_IGN) line that program would not print anything: the kernel sends SIGPIPE to a process that writes to a pipe with no reader, and the default action for SIGPIPE is to kill it. The program simply disappears, with no message, and a student looks for a bug in the loop above.
That is also why ls | head -1 does not print an error when head stops reading: ls is killed by SIGPIPE, which is the intended design.
Pipes against shared memory
| Pipe | Shared memory | |
|---|---|---|
| Who does the waiting | the kernel, automatically | you, with semaphores |
| Data is copied | twice, in and out | never |
| Direction | one way per pipe | both ways |
| Keeps message boundaries | no, it is a byte stream | there are no messages |
| Related processes only | anonymous yes, named no | no, a key is enough |
| Cleaning up | closing the descriptors, and unlink for a FIFO | shmctl with IPC_RMID |
| Blocks when | full on write, empty on read | never |
Practical 2: Process Communication with Pipes
The full comparison, with message queues in it as well, is in [Practical 2 continued: Message Queues, Blocking and Non-blocking].
Procedure
- Write the first program. Compile with
gcc -Wall -Wextra -o pipe1 pipe1.cand run it. - Remove the parent's
close(fd[1]), put the parent's read in awhileloop, and run it again.
The program hangs. Press Ctrl+C and put the line back.
- Write the three-write program and count the reads. Note that the boundaries are gone.
- Write the producer-consumer program and confirm the total.
- Write the FIFO program. Run
ls -l /tmp/munotes_fifoyourself while it is running and note the
p.
- At the shell:
mkfifo /tmp/f, thencat < /tmp/fin one terminal andecho hello > /tmp/fin
another. Then rm /tmp/f.
Result
Two processes exchanged data through an anonymous pipe with no semaphore, because a pipe blocks on an empty read and on a full write. Three separate writes were shown to arrive as one read, which establishes that a pipe is a byte stream with no message boundaries. The producer-consumer problem was solved over a pipe and the total proved nothing was lost. A named pipe was created, seen in ls -l as type p, used between two processes and deleted.
Where marks are lost
- Not closing the unused end. The reader keeps a write end open, end of file never arrives,
and the program hangs. This is the single commonest fault in this exercise.
pipeafterfork. The two processes then have two different pipes.- Mixing up
fd[0]andfd[1]. Zero reads, one writes. - Reading once instead of in a loop, and assuming one read returns one write.
- Printing the buffer with
%swhen nothing wrote a terminating zero. Use%.*sand the
count that read returned.
- Forgetting to
unlinkthe FIFO, so the next run fails withEEXISTfrommkfifo. - Not knowing why a program writing to a closed pipe dies silently. It is
SIGPIPE.
For the journal
Write the aim, MU's own wording, the first program and its output, and then the producer-consumer program with its output and the total. Add the three-write experiment and one sentence saying how many reads it took and why. If your college asks for the named pipe as well, add the mkfifo program and the ls -l line showing the p. The conclusion: a pipe gives message passing with the synchronisation done by the kernel, at the cost of copying the data and of losing the boundaries between writes.
Practical 2: Process Communication with Pipes
Quick revision
- A pipe is a one-way channel in the kernel.
pipe(fd)givesfd[0]to read andfd[1]to
write. Make it before the fork.
- Each process closes the end it does not use, or end of file never arrives.
readreturns the number of bytes it got,0at end of file,-1on error.- A pipe is a byte stream: three writes may arrive as one read, and one write may arrive as
several reads. The reader has to find the message boundaries itself.
- A read on an empty pipe blocks and a write to a full pipe blocks, so a pipe is a bounded buffer
the kernel provides. No semaphore is needed.
- Writing to a pipe with no reader raises
SIGPIPE, which by default kills the writer; ignoring
the signal turns it into EPIPE instead.
- A named pipe is made with
mkfifo, is a file whosels -ltype isp, works between unrelated
programs, blocks on open until the other side arrives, and must be deleted.
- Message passing copies the data twice and synchronises for free. Shared memory copies nothing
and synchronises not at all.
Questions you should be able to answer
1. Which of fd[0] and fd[1] is the read end? fd[0]. It matches standard input being 0 and standard output being 1.
2. Why must pipe be called before fork? Because the child inherits the descriptors that were open when it was created. A pipe made afterwards exists only in the process that made it.
3. A program reads from a pipe in a loop and never finishes, although the writer has ended. Why? The reader still has the write end open, so the kernel does not report end of file. Each side must close the end it does not use.
4. Three writes of eight bytes were made. How many reads does the other side need? It cannot be known. A pipe is a byte stream with no boundaries: in the run on this page one read of 24 bytes took all three. The reader must delimit messages itself.
5. What does read return at end of file, and when does that happen? Zero, and it happens when the pipe is empty and every write end has been closed.
6. What happens when a process writes to a pipe with no reader? It is sent SIGPIPE, which by default kills it with no message. If the signal is ignored or handled, the write fails and sets errno to EPIPE.
Practical 2: Process Communication with Pipes
7. Give two differences between an anonymous pipe and a named pipe. An anonymous pipe has no name and reaches only related processes; a named pipe is a file in a directory, so any program that can open it may use it, and it has to be deleted afterwards.
8. Why does the producer-consumer over a pipe need no semaphore, when the shared memory version needs two? Because the pipe itself blocks: a read on an empty pipe waits and a write to a full pipe waits, and the kernel does the sleeping and waking. Shared memory does nothing of the kind.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.