Strings as Arrays, and the String Library
Chapter Thirty-Seven
Syllabus topic 3, "Pointer and Addresses, Pointer and Function Arguments, Pointer and Arrays."
Pages 175 to 181 of 222
In one line
A string is an array of char ending in '\0', and <string.h> is a set of functions that all find the end of a string by looking for that terminator.
What chapter 12 established, in one paragraph
A string is characters in consecutive memory with a zero byte after the last one. There is no stored length. "Anita" occupies six bytes. char s[] = "Anita"; copies it into memory you own and may change; const char *s = "Anita"; points at the literal and must not be changed. strlen counts to the terminator, sizeof gives the whole array.
Walking a string yourself
Every function in <string.h> is a loop you could write. Writing two of them once is the best way to understand the rest.
#include <stdio.h>
#include <string.h>
size_t my_strlen(const char *s)
{
size_t n = 0;
while (s[n] != '\0') {
n++;
}
return n;
}
int my_strcmp(const char *a, const char *b)
{
size_t i = 0;
while (a[i] != '\0' && a[i] == b[i]) {
i++;
}
return (unsigned char) a[i] - (unsigned char) b[i];
}
int main(void)
{
const char *words[] = {"apple", "apply", "app", "Apple"};
for (int i = 0; i < 4; i++) {
printf("%-8s my_strlen %zu, strlen %zu\n",
words[i], my_strlen(words[i]), strlen(words[i]));
}
printf("\n");
printf("my_strcmp(apple, apply) = %d, strcmp = %d\n",
my_strcmp("apple", "apply"), strcmp("apple", "apply"));
printf("my_strcmp(apply, apple) = %d, strcmp = %d\n",
my_strcmp("apply", "apple"), strcmp("apply", "apple"));
printf("my_strcmp(apple, apple) = %d, strcmp = %d\n",
my_strcmp("apple", "apple"), strcmp("apple", "apple"));
return 0;
}apple my_strlen 5, strlen 5
apply my_strlen 5, strlen 5
app my_strlen 3, strlen 3
Apple my_strlen 5, strlen 5
my_strcmp(apple, apply) = -20, strcmp = -1
my_strcmp(apply, apple) = 20, strcmp = 1
my_strcmp(apple, apple) = 0, strcmp = 0Look at the last three lines of that output before reading on. my_strcmp returned -20 and 20 where the library returned -1 and 1, and both are correct. The standard fixes only the sign of strcmp's result and whether it is zero, never the magnitude. -20 is the actual difference between 'e' and 'y'; the library's -1 is what its own faster implementation happens to produce. Further down this chapter the same library returns -160 for another pair of strings, for no reason you could predict. That is why the only correct test for equality is strcmp(a, b) == 0 and the only correct test for order is the sign.
Both are three-line loops. That is the whole of the string library's design: the terminator is the contract, and every function walks forward until it meets one.
The cast to unsigned char in the comparison is the standard's own requirement, and it is the same reason chapter 34 casts before calling a <ctype.h> function: on a machine where plain char is signed, a character above 127 would compare as negative.
Strings as Arrays, and the String Library
The functions you need
| Function | What it does | Returns |
|---|---|---|
strlen(s) | Counts characters up to '\0' | size_t, the length |
strcpy(d, s) | Copies s into d, terminator included | d |
strncpy(d, s, n) | Copies at most n characters | d |
strcat(d, s) | Appends s to the end of d | d |
strncat(d, s, n) | Appends at most n characters | d |
strcmp(a, b) | Compares | negative, 0, positive |
strncmp(a, b, n) | Compares the first n characters | negative, 0, positive |
strchr(s, c) | Finds the first c | a pointer to it, or NULL |
strstr(s, t) | Finds the first t inside s | a pointer to it, or NULL |
strcspn(s, set) | How many characters before any of set | size_t |
strcmp does not return 1 for "different". It returns a negative number if a sorts before b, zero if they are equal, and a positive number if a sorts after b. The test for equality is strcmp(a, b) == 0, and writing if (strcmp(a, b)) means "if they differ", which reads backwards and is a real source of bugs.
#include <stdio.h>
#include <string.h>
int main(void)
{
char dest[30] = "";
strcpy(dest, "Programming");
printf("after strcpy : \"%s\", length %zu\n", dest, strlen(dest));
strcat(dest, " with C");
printf("after strcat : \"%s\", length %zu\n", dest, strlen(dest));
char *found = strstr(dest, "with");
printf("strstr found \"with\" at offset %ld\n", (long) (found - dest));
char *ch = strchr(dest, 'g');
printf("strchr found the first 'g' at offset %ld\n", (long) (ch - dest));
printf("strcmp(\"abc\", \"abd\") = %d (negative: abc sorts first)\n",
strcmp("abc", "abd"));
printf("equality is tested as strcmp(a, b) == 0, which gives %d\n",
strcmp(dest, "Programming with C") == 0);
return 0;
}after strcpy : "Programming", length 11
after strcat : "Programming with C", length 18
strstr found "with" at offset 12
strchr found the first 'g' at offset 3
strcmp("abc", "abd") = -1 (negative: abc sorts first)
equality is tested as strcmp(a, b) == 0, which gives 1The danger, which is the whole reason these functions are taught carefully
strcpy and strcat do not know how big the destination is. They write until they have copied the terminator, and if the destination is too small they write past the end of it. That is undefined behaviour and it is the classic security defect in C programs.
#include <stdio.h>
#include <string.h>
int main(void)
{
char small[6];
strcpy(small, "Programming with C"); /* 18 characters into 6 bytes */
printf("%s\n", small);
return 0;
}gcc does catch this one, and says so precisely:
overflow.c: In function ‘main’:
overflow.c:8:5: warning: ‘__builtin_memcpy’ writing 19 bytes into a region of size 6 overflows the destination [-Wstringop-overflow=]
8 | strcpy(small, "Programming with C"); /* 18 characters into 6 bytes */
| ^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
overflow.c:6:10: note: destination object ‘small’ of size 6
6 | char small[6];
| ^~~~~Strings as Arrays, and the String Library
It caught it because both numbers are in the source: it can see that small is six bytes and that the literal is nineteen with its terminator. That is the limit of what a compiler can do here. Make the source string something it cannot work out, a name read from the user for instance, and the same strcpy compiles in silence and corrupts memory when it runs. Compare chapter 36, where a[100] on a three-element array drew no warning at all.
So the warning is worth having and is not a safety net. It is not run here, because what it does is undefined and any result printed would be a claim about a program the standard makes no promise about.
The safe forms:
#include <stdio.h>
#include <string.h>
int main(void)
{
char source[40];
char dest[10];
if (fgets(source, (int) sizeof source, stdin) == NULL) { return 1; }
source[strcspn(source, "\n")] = '\0';
/* strncpy: at most 9 characters, then terminate BY HAND */
strncpy(dest, source, sizeof dest - 1);
dest[sizeof dest - 1] = '\0';
printf("strncpy into char[10] : \"%s\" (length %zu)\n", dest, strlen(dest));
/* snprintf: the one that always terminates, and reports what it wanted */
char other[10];
int wanted = snprintf(other, sizeof other, "%s", source);
printf("snprintf into char[10]: \"%s\" (wanted %d characters)\n",
other, wanted);
return 0;
}Programming with Cstrncpy into char[10] : "Programmi" (length 9)
snprintf into char[10]: "Programmi" (wanted 18 characters)strncpy does not add a terminator if it ran out of room. That is the trap, and most textbooks present strncpy as the safe version without saying so. The dest[sizeof dest - 1] = '\0'; line is not optional. snprintf always terminates and returns the length it would have needed, which is why it is the better choice.
The practical: extracting a substring
MU's Practical 6(a). C has no substring function, so it is a loop or a careful strncpy.
#include <stdio.h>
#include <string.h>
/* Copy `count` characters starting at `start` of `s` into `out`.
`out` must have room for count + 1. Returns how many were copied. */
int substring(const char *s, int start, int count, char *out)
{
int len = (int) strlen(s);
int copied = 0;
if (start < 0 || start >= len || count <= 0) {
out[0] = '\0';
return 0;
}
while (copied < count && s[start + copied] != '\0') {
out[copied] = s[start + copied];
copied++;
}
out[copied] = '\0';
return copied;
}
int main(void)
{
const char *text = "Programming with C";
char part[40];
int n = substring(text, 0, 11, part);
printf("from 0, 11 characters : \"%s\" (%d copied)\n", part, n);
n = substring(text, 12, 4, part);
printf("from 12, 4 characters : \"%s\" (%d copied)\n", part, n);
n = substring(text, 12, 100, part);
printf("from 12, 100 asked for: \"%s\" (%d copied, the string ran out)\n",
part, n);
n = substring(text, 50, 4, part);
printf("from 50, out of range : \"%s\" (%d copied)\n", part, n);
return 0;
}Strings as Arrays, and the String Library
from 0, 11 characters : "Programming" (11 copied)
from 12, 4 characters : "with" (4 copied)
from 12, 100 asked for: "with C" (6 copied, the string ran out)
from 50, out of range : "" (0 copied)Three things that make this an answer rather than a sketch: the start and the count are both checked, the loop stops at the terminator as well as at the count, and the result is always terminated even when nothing was copied.
The practical: palindrome
MU's Practical 6(b). A palindrome reads the same backwards. Two pointers walk towards each other, which is chapter 28's two-counter for loop.
#include <stdio.h>
#include <string.h>
#include <ctype.h>
int is_palindrome(const char *s)
{
int left = 0;
int right = (int) strlen(s) - 1;
while (left < right) {
if (s[left] != s[right]) {
return 0;
}
left++;
right--;
}
return 1;
}
/* The version an examiner usually wants next: ignore case and anything
that is not a letter or a digit. */
int is_palindrome_loose(const char *s)
{
int left = 0;
int right = (int) strlen(s) - 1;
while (left < right) {
while (left < right && !isalnum((unsigned char) s[left])) { left++; }
while (left < right && !isalnum((unsigned char) s[right])) { right--; }
if (tolower((unsigned char) s[left]) != tolower((unsigned char) s[right])) {
return 0;
}
left++;
right--;
}
return 1;
}
int main(void)
{
const char *tests[] = {"madam", "level", "hello", "a", "",
"Madam", "Never odd or even"};
printf("%-20s %-8s %s\n", "string", "strict", "loose");
for (int i = 0; i < 7; i++) {
printf("%-20s %-8s %s\n", tests[i][0] ? tests[i] : "(empty)",
is_palindrome(tests[i]) ? "yes" : "no",
is_palindrome_loose(tests[i]) ? "yes" : "no");
}
return 0;
}string strict loose
madam yes yes
level yes yes
hello no no
a yes yes
(empty) yes yes
Madam no yes
Never odd or even no yes"Madam" is not a palindrome strictly, because M and m are different characters. Whether that is the wanted answer depends on the question, and saying so is worth a mark. The empty string and a single character are palindromes under both, because the loop never runs.
The practical: strlen and strcmp
MU's Practical 6(c), with the input read and the results shown.
#include <stdio.h>
#include <string.h>
int main(void)
{
char a[40], b[40];
printf("Enter the first string : ");
if (fgets(a, (int) sizeof a, stdin) == NULL) { return 1; }
a[strcspn(a, "\n")] = '\0';
printf("Enter the second string: ");
if (fgets(b, (int) sizeof b, stdin) == NULL) { return 1; }
b[strcspn(b, "\n")] = '\0';
printf("\nstrlen(\"%s\") = %zu\n", a, strlen(a));
printf("strlen(\"%s\") = %zu\n", b, strlen(b));
int r = strcmp(a, b);
printf("strcmp = %d, so ", r);
if (r == 0) {
printf("the two strings are equal\n");
} else if (r < 0) {
printf("\"%s\" sorts before \"%s\"\n", a, b);
} else {
printf("\"%s\" sorts after \"%s\"\n", a, b);
}
return 0;
}Strings as Arrays, and the String Library
apple
applyEnter the first string : Enter the second string:
strlen("apple") = 5
strlen("apply") = 5
strcmp = -160, so "apple" sorts before "apply"Counting things in a string
The other standard question, and it is one loop.
#include <stdio.h>
#include <ctype.h>
#include <string.h>
int main(void)
{
const char *s = "MU B.Sc. IT Semester 1";
int vowels = 0, consonants = 0, digits = 0, spaces = 0, others = 0, words = 0;
int in_word = 0;
for (int i = 0; s[i] != '\0'; i++) {
unsigned char c = (unsigned char) s[i];
if (isspace(c)) {
spaces++;
in_word = 0;
} else {
if (!in_word) { words++; }
in_word = 1;
if (isdigit(c)) {
digits++;
} else if (isalpha(c)) {
char l = (char) tolower(c);
if (l == 'a' || l == 'e' || l == 'i' || l == 'o' || l == 'u') {
vowels++;
} else {
consonants++;
}
} else {
others++;
}
}
}
printf("\"%s\"\n", s);
printf("length %zu, words %d\n", strlen(s), words);
printf("vowels %d, consonants %d, digits %d, spaces %d, other %d\n",
vowels, consonants, digits, spaces, others);
return 0;
}"MU B.Sc. IT Semester 1"
length 22, words 5
vowels 5, consonants 10, digits 1, spaces 4, other 2The in_word flag is the standard way to count words: a word begins at a non-space character that follows a space or the start of the string.
What this does NOT mean
strcmp does not return 1 or 0. It returns a negative number, zero or a positive number. Test == 0 for equality.
strlen is not sizeof. strlen walks to the terminator at run time; sizeof is the array's size at compile time.
strncpy is not simply "the safe strcpy". It does not terminate if it filled the destination. Terminate by hand, or use snprintf.
strcpy(d, s) does not check anything. The destination must already be big enough, and that is your responsibility.
A string cannot be compared with == or copied with =. Those work on addresses and on nothing respectively.
Strings as Arrays, and the String Library
strlen does not count the terminator. strlen("Anita") is 5 while the array needs 6 bytes.
gets does not exist. It was removed in C11. Use fgets.
Quick revision
- A string is a
chararray ending in'\0'; there is no stored length. strlencounts to the terminator;sizeofgives the array.strcmp(a,b) == 0means equal. Negative meansasorts first.strcpyandstrcatdo not know the destination's size; they can write past its end.strncpymay leave the destination unterminated. Terminate it yourself, or usesnprintf.- Every string function is a loop that stops at the terminator, and
my_strlenis three lines. - Substring: check the start and the count, copy while both the count and the terminator allow, and always terminate.
- Palindrome: two indexes walking towards each other; decide whether case and punctuation count.
- Count words with an
in_wordflag. - Cast a
chartounsigned charbefore any<ctype.h>call.
Test yourself
1. What does strcmp("abc", "abd") return, and what does the sign mean?
A negative number, because 'c' is less than 'd', so "abc" sorts before "abd". The exact magnitude is not specified in a useful way; only the sign and zero matter.
2. For char s[20] = "hello";, what are strlen(s) and sizeof s?
5 and 20.
3. Why is strncpy(d, s, sizeof d) still unsafe?
If s is at least as long as d, strncpy fills d with no terminator, so d is not a string. Copy at most sizeof d - 1 and set the last byte to '\0', or use snprintf.
4. Write a loop that finds the length of a string without strlen.
size_t n = 0;
while (s[n] != '\0') n++;5. Is "Madam" a palindrome?
Not if case matters, because 'M' and 'm' are different characters. It is if you compare case-insensitively. Say which convention you are using.
6. How do you test two strings for equality?
if (strcmp(a, b) == 0). if (a == b) compares addresses, and if (strcmp(a, b)) is true when they differ.
7. Why must a char be cast to unsigned char before tolower?
Because the function takes an int that must be representable as an unsigned char or be EOF, and on a machine where plain char is signed a byte above 127 would be negative, which is undefined behaviour.
What can be asked on this, and how to answer it
"Explain any five string handling functions with examples." Take strlen, strcpy, strcat, strcmp and strstr, give the prototype, one line of what it does and a worked call with its result. For strcmp give the three possible signs, because that is the one examiners probe.
Strings as Arrays, and the String Library
"Write a program to find whether a string is a palindrome." Give the two-index loop version. Mention the case and punctuation question and say which convention your program uses; that one sentence is often the difference between full and partial marks.
"Write a program to extract a portion of a string." Give the substring function from this chapter with its bounds checks, and say that C has no substring function so this is what a programmer writes.
"Write a program using strlen and strcmp." Give this chapter's program with fgets, the newline trimmed, and the three-way report on the sign of strcmp.
"What is the difference between strcpy and strncpy?" strcpy copies until the terminator with no regard for the destination's size. strncpy copies at most n characters, and if it reaches the limit it does not add a terminator. Neither is safe without care; snprintf always terminates.
"How would you count the vowels, consonants and words in a string?" One pass with <ctype.h> tests, and an in_word flag for the words. Give the program and say why the flag is needed: a word begins at a non-space that follows a space.
The rest of this subject
These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.