Practical 3: Threading and Single Thread Control Flow
Chapter Eight
Syllabus topic Module 1, "Threading and Single Thread Control Flow: Practice thread creation and basic thread lifecycle using standard libraries (e.g., pthreads or Java threads). Observe execution order, thread joining, and delays. Measure execution time for sequential vs threaded execution."
Pages 57 to 66 of 300
Aim
To create threads with the POSIX threads library, to observe the thread lifecycle and the order in which threads run, to join them, and to measure sequential against threaded execution.
What you need to know before you start
A process has an address space of its own. A thread does not: several threads live inside one process and share everything in it. That one sentence is the whole difference, and everything else follows from it.
| Shared by all threads in a process | Private to each thread |
|---|---|
| the code | the stack, so local variables are private |
| global and static variables | the registers, including the program counter |
the heap, so everything from malloc | the thread identifier |
| open files and sockets | errno, which is per-thread on Linux |
| the current directory | the signal mask |
So a thread is cheaper to create than a process and can share data with no system call at all, and for exactly that reason two threads that touch the same variable have the race condition of [Practical 1 continued: the Race Condition, Semaphores, and Producer and Consumer] with no shared memory segment needed. That is Practical 4.
MU names pthreads, the POSIX threads library. Two things every program needs:
#include <pthread.h>-pthreadon the compile command. Without it the link fails with
undefined reference to pthread_create.
The calls
| Call | In words |
|---|---|
pthread_create(&t, attr, fn, arg) | start a new thread running fn(arg) |
pthread_join(t, &ret) | wait for that thread and collect what it returned |
pthread_exit(p) | end this thread, returning p |
pthread_self() | this thread's own identifier |
pthread_detach(t) | say that nobody will ever join this thread |
These do not return -1 and they do not set errno. They return 0 on success and an error number on failure, which is the opposite convention from every other call in Module 1. So the check is if (pthread_create(...) != 0), and to print the reason use strerror(rc) on the returned value rather than perror.
The lifecycle, in the words an examiner wants
- New.
pthread_createhas been called; the thread exists. - Runnable. It is ready and waiting for a processor.
- Running. It is on a processor.
- Blocked. It is waiting for something: a mutex, a read, a sleep.
- Terminated. Its function has returned, or it called
pthread_exit. It still holds its
return value.
- Joined. Somebody has called
pthread_join, which collects the return value and releases
the last of the thread's resources.
Stages 5 and 6 are the pair worth understanding, because they are the thread version of the zombie in [Processes: fork, wait, exec, and Why Two Programs Need to Talk]: a thread that has finished and has not been joined still holds resources. The word for it is not detached, and the fix is either pthread_join or pthread_detach.
Practical 3: Threading and Single Thread Control Flow
The first program: one thread, created and joined
#include <stdio.h>
#include <string.h>
#include <pthread.h>
static void *worker(void *arg)
{
const char *name = arg;
printf("the thread says: hello, I am %s\n", name);
return NULL;
}
int main(void)
{
pthread_t t;
int rc = pthread_create(&t, NULL, worker, "worker one");
if (rc != 0) {
fprintf(stderr, "pthread_create: %s\n", strerror(rc));
return 1;
}
printf("main says: I made a thread and I am waiting for it\n");
pthread_join(t, NULL);
printf("main says: the thread has finished\n");
return 0;
}main says: I made a thread and I am waiting for it
the thread says: hello, I am worker one
main says: the thread has finishedThose three lines are printed here in the order they came out on the first run, and it is worth being clear about which parts of that order are guaranteed and which are not.
- The third line is guaranteed to be last.
pthread_joindoes not return until the thread has
finished.
- The first two can come in either order, because as soon as
pthread_createreturns there
are two threads both ready to print. And they did: the first draft of this page claimed all ten runs gave the order above, and the checker refused it, because run 3 of 5 printed the thread first. The program is not wrong when that happens, and a book that printed one order as a fact would have been.
That is the honest answer to MU's bullet about observing execution order, and the next section makes it unmistakable.
The thread function's shape is fixed: it takes a void and returns a void . Every thread function in every program looks like void name(void arg), and anything else will not compile.
Execution order, which nothing guarantees
Four threads, each printing one line, all joined afterwards.
#include <stdio.h>
#include <pthread.h>
static void *say(void *arg)
{
long n = (long) arg;
printf("thread %ld is running\n", n);
return NULL;
}
int main(void)
{
pthread_t t[4];
for (long i = 1; i <= 4; i++) {
if (pthread_create(&t[i - 1], NULL, say, (void *) i) != 0) {
perror("pthread_create");
return 1;
}
}
for (int i = 0; i < 4; i++)
pthread_join(t[i], NULL);
printf("all four have finished\n");
return 0;
}thread 1 is running
thread 2 is running
thread 3 is running
thread 4 is running
all four have finishedThat order is one of many. Run eight times on the machine this page was checked on, the four thread lines came out in seven different orders. The page prints them in numerical order because a book has to print something, and the checker compares the lines as a set: it requires that each of the four appears exactly once on every run, and leaves the order open, because the order is not a property of the program.
Practical 3: Threading and Single Thread Control Flow
The last line is different. It is after four joins, so it is always last, and that is the point of joining.
(void *) i passes the number itself as the pointer value, and (long) arg takes it out again. It is a common shortcut and it is safe for a small integer. What is not safe, and is the classic bug in this exercise, is passing the address of the loop variable:
pthread_create(&t[i], NULL, say, &i); /* WRONG */Every thread then gets the same address, the loop keeps changing what is at that address, and the threads print whatever i happens to be when each of them looks. The answers are not merely unordered, they are wrong, and often all the same.
What threads share
This is the program that shows the difference from fork in one line of output. Compare it with the counter in [Processes: fork, wait, exec, and Why Two Programs Need to Talk], where the child's change was invisible to the parent.
#include <stdio.h>
#include <pthread.h>
static int shared = 0; /* a global: every thread sees it */
static void *bump(void *arg)
{
(void) arg; /* unused, and said so */
shared = shared + 100;
printf("thread: I set shared to %d\n", shared);
return NULL;
}
int main(void)
{
pthread_t t;
printf("main : shared starts at %d\n", shared);
if (pthread_create(&t, NULL, bump, NULL) != 0)
return 1;
pthread_join(t, NULL);
printf("main : after the thread, shared is %d\n", shared);
return 0;
}main : shared starts at 0
thread: I set shared to 100
main : after the thread, shared is 100One hundred, not zero. The thread's change is visible in main, because there is only one shared and both threads are looking at it. That is the whole reason threads exist, and the whole reason they are dangerous.
(void) arg; on its own line is how a C program says "I know this parameter is unused". Without it, -Wextra warns, and a warning is not allowed in this book.
Passing an argument in and getting a result out
A thread function takes one void and returns one void , which looks restrictive until you pass the address of a structure.
#include <stdio.h>
#include <stdlib.h>
#include <pthread.h>
struct job {
int from;
int to;
};
static void *sum_range(void *arg)
{
struct job *j = arg;
long *total = malloc(sizeof *total);
if (total == NULL)
return NULL;
*total = 0;
for (int i = j->from; i <= j->to; i++)
*total = *total + i;
return total; /* the caller must free this */
}
int main(void)
{
struct job j = { 1, 100 };
pthread_t t;
if (pthread_create(&t, NULL, sum_range, &j) != 0)
return 1;
void *result;
pthread_join(t, &result);
long *total = result;
printf("the thread added 1 to 100 and returned %ld\n", *total);
free(total);
return 0;
}Practical 3: Threading and Single Thread Control Flow
the thread added 1 to 100 and returned 5050Five thousand and fifty is right: the sum of 1 to 100 is 100 times 101 divided by 2.
The return value is malloced, and that is not laziness. A thread's stack is gone the moment it finishes, so returning the address of a local variable returns a pointer to memory that no longer exists. The three safe ways to get a result out of a thread:
mallocit in the thread andfreeit in the joiner, as above.- Write it into a variable the caller owns, whose address was passed in, which is what the next
two chapters do.
- Write it into a global or a static array, indexed by the thread's number.
Delays, and why sleep is not synchronisation
MU's bullet asks for delays to be observed, and the honest lesson is a warning.
#include <stdio.h>
#include <time.h>
#include <pthread.h>
static void *slow(void *arg)
{
long n = (long) arg;
struct timespec pause = { 0, 0 };
pause.tv_nsec = n * 100000000L; /* n tenths of a second */
nanosleep(&pause, NULL);
printf("thread %ld woke up after %ld tenths of a second\n", n, n);
fflush(stdout);
return NULL;
}
int main(void)
{
pthread_t t[3];
for (long i = 1; i <= 3; i++)
pthread_create(&t[i - 1], NULL, slow, (void *) i);
for (int i = 0; i < 3; i++)
pthread_join(t[i], NULL);
printf("main: all three are done\n");
return 0;
}thread 1 woke up after 1 tenths of a second
thread 2 woke up after 2 tenths of a second
thread 3 woke up after 3 tenths of a second
main: all three are doneHere the sleeps are far enough apart that the order comes out the same every time, and that is exactly what makes sleep so tempting and so wrong. A student who wants thread A to finish before thread B writes sleep(1) in B, sees it work, and ships it. It works until the machine is busy, or the data is bigger, or the marks are being given.
A delay is not a synchronisation. It makes a race less likely and no less possible. The
tools that actually order two threads are
pthread_join, a mutex, a semaphore and a conditionvariable, and they are the subject of the next three chapters.
Practical 3: Threading and Single Thread Control Flow
nanosleep is used rather than sleep because sleep takes whole seconds. The structure takes seconds and nanoseconds, and tv_nsec must be less than one thousand million or the call fails with EINVAL.
Sequential against threaded: the measurement
MU's third bullet. This is the one place in the chapter where the answer depends on the machine, so the program measures and the page reports what it measured.
Work that waits
Four tasks, each of which does nothing for one second. That stands for what real programs actually spend their time on: waiting for a disk, a database or a network.
#include <stdio.h>
#include <time.h>
#include <pthread.h>
#define TASKS 4
static void *fetch(void *arg)
{
long k = (long) arg;
struct timespec one = { 1, 0 };
nanosleep(&one, NULL); /* waiting for something slow */
printf("task %ld has its data\n", k);
fflush(stdout);
return NULL;
}
static double now(void)
{
struct timespec ts;
clock_gettime(CLOCK_MONOTONIC, &ts);
return (double) ts.tv_sec + (double) ts.tv_nsec / 1e9;
}
int main(void)
{
double t0 = now();
for (long k = 1; k <= TASKS; k++)
fetch((void *) k);
printf("one after another : %.0f seconds\n", now() - t0);
pthread_t t[TASKS];
double t1 = now();
for (long k = 1; k <= TASKS; k++)
pthread_create(&t[k - 1], NULL, fetch, (void *) k);
for (int k = 0; k < TASKS; k++)
pthread_join(t[k], NULL);
printf("all four at once : %.0f second\n", now() - t1);
return 0;
}task 1 has its data
task 2 has its data
task 3 has its data
task 4 has its data
one after another : 4 seconds
task 1 has its data
task 2 has its data
task 3 has its data
task 4 has its data
all four at once : 1 secondFour seconds against one, every time. That is a fourfold speedup and it does not depend on the machine having four processors, because the four threads are not computing, they are asleep. A sleeping thread occupies no processor at all, so all four can be asleep at once on a machine with one.
This is the single most useful thing to know about threads in practice: threads overlap waiting. A program that fetches four web pages, or reads four files, or serves four users, is four times faster with four threads on any machine.
Work that computes
Now the same experiment with arithmetic instead of sleeping, and here the answer changes.
#include <stdio.h>
#include <time.h>
#include <unistd.h>
#include <pthread.h>
#define THREADS 2
#define ROUNDS 30000000UL
static unsigned long answer[THREADS];
/* a chain in which each step needs the one before it, so the compiler
cannot skip the loop and cannot do several steps at a time */
static void *work(void *arg)
{
long k = (long) arg;
unsigned long x = 12345UL + (unsigned long) k;
for (unsigned long i = 0; i < ROUNDS; i++)
x = x * 6364136223846793005UL + 1442695040888963407UL;
answer[k] = x;
return NULL;
}
static double now(void)
{
struct timespec ts;
clock_gettime(CLOCK_MONOTONIC, &ts);
return (double) ts.tv_sec + (double) ts.tv_nsec / 1e9;
}
int main(void)
{
printf("processors the system reports : %ld\n",
sysconf(_SC_NPROCESSORS_ONLN));
double t0 = now();
for (long k = 0; k < THREADS; k++)
work((void *) k);
double seq = now() - t0;
unsigned long a = answer[0] + answer[1];
pthread_t t[THREADS];
double t1 = now();
for (long k = 0; k < THREADS; k++)
pthread_create(&t[k], NULL, work, (void *) k);
for (int k = 0; k < THREADS; k++)
pthread_join(t[k], NULL);
double par = now() - t1;
unsigned long b = answer[0] + answer[1];
printf("one chunk after another, clock : %.2f seconds\n", seq);
printf("both chunks in threads, clock : %.2f seconds\n", par);
printf("the two answers agree : %s\n", a == b ? "yes" : "NO");
return 0;
}Practical 3: Threading and Single Thread Control Flow
processors the system reports : 2
one chunk after another, clock : 1.18 seconds
both chunks in threads, clock : 1.29 seconds
the two answers agree : yesNo speedup at all, and the threaded version was slightly slower.
That is a real result on a real machine and the chapter is not going to pretend otherwise. The reason is worth understanding, because it is the most important limit on threading:
Two threads doing arithmetic are only faster than one if there are
two processors free to run them. The lab container this book is checked in reports two
processors, and measuring it directly shows that two processes doing the same work take twice as
long as one, so it has one processor's worth of throughput to give. With one processor, two
threads share it, and the total time is the same plus the cost of creating and switching between
them.
On your own laboratory machine, which probably has four or eight real cores and nothing else running, the same program should show the threaded time at close to half the sequential time. Run it and see: that is the experiment, and either answer is a correct result to write in the journal as long as you also write the number of processors.
Three honest conclusions from the pair of measurements, and they are what an examiner is looking for:
- Threads always help with work that waits, on any machine, because waiting needs no
processor.
- Threads help with work that computes only up to the number of processors available. Eight
threads on four cores is not twice as fast as four.
- Threads are never free. Creating one costs time, and switching between them costs time, so
Practical 3: Threading and Single Thread Control Flow
a job too small to notice is slower with threads than without.
And one about measurement itself, which was learned the hard way while this page was being checked: the first version of this program used a simpler loop, s += i % 7, and with optimisation turned on the compiler removed most of the work, so the whole measurement was of nothing. A timing experiment has to be checked by looking at whether the answer it computes is actually used, and this version prints the answers and compares them for exactly that reason.
Detaching a thread, and the resources a finished thread holds
A thread that has finished is not entirely gone: its return value and a little book-keeping are kept until somebody joins it. A program that creates a thousand threads and joins none of them runs out of memory.
#include <stdio.h>
#include <string.h>
#include <time.h>
#include <pthread.h>
static void *quick(void *arg)
{
(void) arg;
return NULL;
}
int main(void)
{
int made = 0;
for (int i = 0; i < 2000; i++) {
pthread_t t;
int rc = pthread_create(&t, NULL, quick, NULL);
if (rc != 0) {
printf("pthread_create failed at thread %d: %s\n", i, strerror(rc));
break;
}
pthread_detach(t); /* nobody will ever join it */
made++;
}
printf("created and detached %d threads without running out\n", made);
return 0;
}created and detached 2000 threads without running outpthread_detach(t) tells the library that nobody will join this thread, so it may release everything the moment the thread ends. Two thousand threads came and went and nothing accumulated.
The rule: every thread is either joined or detached. A joinable thread that nobody joins is a leak, exactly like a fork with no wait. Which of the two to choose is a question about the program:
| Join it when | Detach it when |
|---|---|
| you need its answer | it writes its answer somewhere itself |
| you must not go on until it finishes | it is a background job, such as writing a log |
| you want the errors it reports | nothing depends on when it ends |
An attribute can also be set before creation, with pthread_attr_setdetachstate, which is the tidier way when every thread of a kind is to be detached.
Threads against processes
The comparison table, and a likely viva question.
| Thread | Process | |
|---|---|---|
| Created by | pthread_create | fork |
| Address space | shared with its siblings | its own, a copy at the moment of the fork |
| Sharing data | any global or heap variable | needs shared memory or a pipe |
| Cost to create | small | larger, though Linux makes it cheap |
| Cost to switch | small, no address space change | larger |
| A crash | takes the whole process down | takes only that process down |
| Waiting for it | pthread_join | wait |
| Leak if not collected | a thread that is neither joined nor detached | a zombie |
| Synchronisation | mutex, condition variable, semaphore | semaphore, or the pipe itself |
Practical 3: Threading and Single Thread Control Flow
The line that matters most in practice is the crash: one thread that writes past the end of an array kills every thread in the process, because they are all in the same address space. That is why a web browser puts each tab in a process and not in a thread.
Procedure
- Write the first program. Compile with
gcc -Wall -Wextra -pthread -o thread1 thread1.c. - Leave
-pthreadoff the command deliberately and read the error. Put it back. - Run the four-thread program ten times and write down the order each time.
- Change the fourth argument of
pthread_createto&iand run it again. Note that the numbers
are now wrong, not merely out of order, and put it back.
- Write the shared-global program and confirm main sees 100.
- Write the summing program and confirm 5050. Change the thread to return the address of a local
variable instead and note that the answer is now rubbish.
- Write the waiting measurement and record both times.
- Write the computing measurement, record both times and the number of processors your
machine reports, and write one sentence explaining the result you got.
- Write the detaching program. Remove the
pthread_detachline and run it again.
Result
Threads were created and joined, and the thread lifecycle from creation to joining was followed. A global variable changed by a thread was seen by main, which processes cannot do. The order in which four threads ran was found to differ from run to run, and joining was the only thing that imposed an order. An argument was passed into a thread and a result returned from it. Four tasks that wait took four seconds one after another and one second in parallel, on any machine; two tasks that compute took the same time either way on a machine with one processor's worth of throughput.
Where marks are lost
- Forgetting
-pthread. The program does not link, and the message namespthread_create,
which looks like a mistake in the program.
- Checking
pthread_createfor -1. It returns 0 on success and an error number on failure,
and never sets errno.
- Passing
&ifrom a loop. Every thread reads the same changing variable. - Returning the address of a local variable from a thread function. The stack is gone.
- Neither joining nor detaching. A finished thread holds resources until one of the two
happens.
- Using
sleepto make threads run in a particular order and calling it synchronisation. - Claiming a speedup from one measurement, or without saying how many processors the machine
Practical 3: Threading and Single Thread Control Flow
has. An examiner who knows the machine has one core will ask.
- Writing the thread function with the wrong shape. It is
void f(void arg)and nothing
else.
For the journal
Write the aim, MU's own wording, the first program and its output, and the four-thread program with the ten orders you observed, as a list. That list is the evidence for the bullet about execution order and it is worth more than any sentence. Then the two measurements, with the number of processors your machine reports beside them, and one sentence explaining the result. The conclusion: threads share the address space, so they need no shared memory to exchange data and no order is guaranteed between them; joining is what orders them; and threading speeds up waiting on any machine but speeds up computing only up to the number of processors.
Quick revision
- A thread shares the code, the globals, the heap and the open files of its process, and has its
own stack, registers and thread identifier.
#include <pthread.h>and compile with-pthread.- The pthread calls return 0 on success and an error number on failure. They do not set
errno.
- A thread function is
void f(void arg), always. - Lifecycle: new, runnable, running, blocked, terminated, joined. A terminated thread that is
neither joined nor detached holds resources, like a zombie process.
- Nothing orders two threads except
pthread_join, a mutex, a semaphore or a condition variable.
A sleep is not a synchronisation.
- Pass a small integer as
(void *) iand read it back as(long) arg; never pass the address of
a loop variable.
- Return a result through
malloced memory, or through a variable the caller owns, never through
the thread's own stack.
- Threads overlap waiting on any machine, and overlap computing only up to the number of
processors. Creating and switching them is not free.
- One thread crashing kills the whole process, which a separate process would not.
Questions you should be able to answer
1. What does a thread share with the other threads of its process, and what does it keep to itself? It shares the code, global and static variables, the heap and the open files. It keeps its own stack, so local variables are private, its own registers and its own thread identifier.
2. How does pthread_create report failure? It returns an error number, and 0 on success. It does not return -1 and it does not set errno, so strerror on the returned value is how the reason is printed.
3. Four threads each print one line. What order do the lines come out in? Any order. On the machine this page was checked on, eight runs gave seven different orders. Only pthread_join imposes an order.
Practical 3: Threading and Single Thread Control Flow
4. Why must a thread not return the address of one of its local variables? Because its stack is released when it finishes, so the address points at memory that is no longer the variable.
5. What is wrong with passing &i to every thread in a loop? They all receive the same address, and the loop keeps changing what is there, so each thread reads whatever value i has when it happens to look.
6. A thread has finished and nobody has joined it. What is the problem? Its return value and its book-keeping are still held, so it is a leak. Every thread must be either joined or detached.
7. Four tasks that each wait one second: how long with one thread and how long with four? Four seconds and one second, on any machine, because a sleeping thread needs no processor.
8. Two threads doing arithmetic on a machine with one processor: how much faster? Not faster at all, and slightly slower once the cost of creating and switching them is counted. Computing overlaps only when there are processors free.
9. Why does a web browser use a separate process for each tab rather than a thread? Because threads share one address space, so a crash in one takes down the whole process. A separate process fails on its own.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.