munotes®

Practical 3: Arrays, Basic Operations, Indexing and Slicing

Chapter Six

Syllabus topic Module 1, practical 3(a), "Write a program to perform basic operations, indexing and slicing on arrays"

Pages 27 to 32 of 297

Aim

To create arrays, to perform the basic operations on them, and to index and slice them.

First, three things called an array

This confuses more students on this paper than any other single point, because Python has three different things that a teacher may call an array.

What it isImportHolds
listPython's built in sequencenothinganything, mixed
array.arraya compact array of one numeric typefrom array import arrayone type, fixed by a type code
numpy.ndarraythe numerical array, any number of dimensionsimport numpy as npone type, with mathematics built in

A list is what you get from [1, 2, 3]. It can hold a number, a string and another list all at once, because each slot holds a reference to an object rather than the value itself. That flexibility costs memory and speed.

An array.array is a real array in the C sense: a contiguous block of one numeric type. It is part of the standard library and needs nothing installed.

A NumPy array is the one every later exercise means. It is also one contiguous block of one type, and it brings element by element arithmetic, multiple dimensions, and a large library of mathematical functions.

from array import array
import numpy as np

py_list = [3, 1, 4, 1, 5]
std_array = array('i', [3, 1, 4, 1, 5])
np_array = np.array([3, 1, 4, 1, 5])

print("list        ", py_list, type(py_list).__name__)
print("array.array ", std_array.tolist(), std_array.typecode, "itemsize", std_array.itemsize)
print("numpy array ", np_array, np_array.dtype)

print("a list can hold anything:", [1, "two", 3.0, [4]])
list         [3, 1, 4, 1, 5] list
array.array  [3, 1, 4, 1, 5] i itemsize 4
numpy array  [3 1 4 1 5] int64
a list can hold anything: [1, 'two', 3.0, [4]]

The 'i' in array('i', ...) is the type code for a signed integer. 'd' is a double, 'f' a float, 'b' a signed byte. An array.array refuses a value of the wrong type, which is the point of it:

from array import array

numbers = array('i', [1, 2, 3])
numbers.append(4.5)
TypeError: 'float' object cannot be interpreted as an integer

Creating an array

import numpy as np

print(np.array([3, 1, 4, 1, 5, 9, 2, 6]))
print(np.array([[1, 2, 3], [4, 5, 6]]))
print(np.zeros(5))
print(np.ones(4, dtype=int))
print(np.full(4, 7))
print(np.arange(0, 10, 2))
print(np.linspace(0, 1, 5))
print(np.eye(3, dtype=int))
[3 1 4 1 5 9 2 6]
[[1 2 3]
 [4 5 6]]
[0. 0. 0. 0. 0.]
[1 1 1 1]
[7 7 7 7]
[0 2 4 6 8]
[0.   0.25 0.5  0.75 1.  ]
[[1 0 0]
 [0 1 0]
 [0 0 1]]
CallWhat it makes
np.array(list)an array from a list, or a list of lists
np.zeros(n)n zeros, as floats
np.ones(n)n ones
np.full(n, v)n copies of v
np.arange(start, stop, step)like range, and the stop is excluded
np.linspace(start, stop, count)count values evenly spaced, and the stop IS included
np.eye(n)the identity matrix, ones down the diagonal
munotes.in27

Practical 3: Arrays, Basic Operations, Indexing and Slicing

The difference between arange and linspace is a favourite question: arange takes a step and excludes the stop, linspace takes a count and includes it.

Indexing

import numpy as np

a = np.array([10, 20, 30, 40, 50, 60])

print("the whole array   ", a)
print("a[0]  first       ", a[0])
print("a[3]               ", a[3])
print("a[-1] last        ", a[-1])
print("a[-2] second last ", a[-2])
print("len(a)            ", len(a))
the whole array    [10 20 30 40 50 60]
a[0]  first        10
a[3]                40
a[-1] last         60
a[-2] second last  50
len(a)             6

Indexing starts at 0, so the first item is a[0] and the last is a[len(a) - 1]. A negative index counts from the end, so a[-1] is the last item. This is worth learning properly because it removes a whole class of off by one errors: to get the last item you never need to know the length.

Going past the end raises an error rather than returning rubbish:

import numpy as np

a = np.array([10, 20, 30])
print(a[5])
IndexError: index 5 is out of bounds for axis 0 with size 3

For a two dimensional array, one pair of brackets holds both indexes, separated by a comma:

import numpy as np

m = np.array([[1, 2, 3, 4],
              [5, 6, 7, 8],
              [9, 10, 11, 12]])

print(m)
print("m[1, 2] row 1 column 2 :", m[1, 2])
print("m[0]    the whole row 0:", m[0])
print("m[:, 1] the whole col 1:", m[:, 1])
print("m[-1, -1] bottom right :", m[-1, -1])
[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]
m[1, 2] row 1 column 2 : 7
m[0]    the whole row 0: [1 2 3 4]
m[:, 1] the whole col 1: [ 2  6 10]
m[-1, -1] bottom right : 12

m[1][2] also works and gives the same value, but m[1, 2] is the NumPy way and is faster, because m[1][2] builds the whole row as an intermediate array first.

Slicing

A slice takes a piece of the array. It is written start:stop:step, and every part may be left out.

The stop is excluded. a[1:4] gives items 1, 2 and 3, not 4. This is the single most important fact about slicing, and the reason it is designed that way is worth knowing: a[:n] and a[n:] between them give the whole array exactly once, with nothing repeated and nothing missed.

import numpy as np

a = np.array([10, 20, 30, 40, 50, 60, 70])

print("a          ", a)
print("a[1:4]     ", a[1:4])
print("a[:3]      ", a[:3])
print("a[4:]      ", a[4:])
print("a[:]       ", a[:])
print("a[::2]     ", a[::2])
print("a[1::2]    ", a[1::2])
print("a[::-1]    ", a[::-1])
print("a[-3:]     ", a[-3:])
print("a[2:2]     ", a[2:2], "an empty slice, not an error")
print("a[1:100]   ", a[1:100], "an out of range stop is clipped, not an error")
munotes.in28

Practical 3: Arrays, Basic Operations, Indexing and Slicing

a           [10 20 30 40 50 60 70]
a[1:4]      [20 30 40]
a[:3]       [10 20 30]
a[4:]       [50 60 70]
a[:]        [10 20 30 40 50 60 70]
a[::2]      [10 30 50 70]
a[1::2]     [20 40 60]
a[::-1]     [70 60 50 40 30 20 10]
a[-3:]      [50 60 70]
a[2:2]      [] an empty slice, not an error
a[1:100]    [20 30 40 50 60 70] an out of range stop is clipped, not an error
SliceMeans
a[1:4]items 1, 2, 3. The stop is excluded
a[:3]from the start up to but not including 3
a[4:]from 4 to the end
a[:]the whole array
a[::2]every second item
a[::-1]the array reversed
a[-3:]the last three items

Two behaviours in that output are worth remembering because they are the opposite of indexing: an empty slice is not an error, and nor is a stop past the end. a[5] on a three item array raises IndexError; a[1:100] quietly gives what there is. That is deliberate, and it is why slicing is the safe way to take a piece of something whose length you are not sure of.

Slicing a two dimensional array takes a slice in each direction:

import numpy as np

m = np.arange(1, 13).reshape(3, 4)

print(m)
print("m[0:2, 1:3] first two rows, middle two columns")
print(m[0:2, 1:3])
print("m[:, ::2] every second column")
print(m[:, ::2])
print("m[::-1] the rows reversed")
print(m[::-1])
[[ 1  2  3  4]
 [ 5  6  7  8]
 [ 9 10 11 12]]
m[0:2, 1:3] first two rows, middle two columns
[[2 3]
 [6 7]]
m[:, ::2] every second column
[[ 1  3]
 [ 5  7]
 [ 9 11]]
m[::-1] the rows reversed
[[ 9 10 11 12]
 [ 5  6  7  8]
 [ 1  2  3  4]]

The one that will be asked: a slice is a view

Slicing a Python list gives a new list. Slicing a NumPy array gives a view, which is a window onto the same memory. Write through the window and the original changes.

import numpy as np

py = [10, 20, 30, 40, 50]
piece = py[1:4]
piece[0] = 999
print("list  :", py, "and the slice", piece)

arr = np.array([10, 20, 30, 40, 50])
window = arr[1:4]
window[0] = 999
print("array :", arr, "and the slice", window)
print("does the slice own its data?", window.base is None)
list  : [10, 20, 30, 40, 50] and the slice [999, 30, 40]
array : [ 10 999  30  40  50] and the slice [999  30  40]
does the slice own its data? False
munotes.in29

Practical 3: Arrays, Basic Operations, Indexing and Slicing

The list is untouched and the array is changed. This is the most important difference between a list and a NumPy array, and it is the whole subject of the next chapter, which is MU's own third bullet on aliasing and copying. It is also not a defect: a view costs no memory and no copying, which is why NumPy is fast on large data.

window.base is the array a view looks into, and it is None for an array that owns its own memory. That one line is the quickest way to answer "is this a view or a copy" at the table.

The basic operations MU asks for

import numpy as np

a = np.array([3, 1, 4, 1, 5, 9, 2, 6])

print("array        ", a)
print("length       ", len(a), "or a.size", a.size)
print("sum          ", a.sum())
print("smallest     ", a.min(), "at index", a.argmin())
print("largest      ", a.max(), "at index", a.argmax())
print("mean         ", a.mean())
print("sorted       ", np.sort(a))
print("the original is untouched:", a)
print("reversed     ", a[::-1])
print("is 5 present ", 5 in a)
print("where a > 3  ", np.where(a > 3)[0])
print("cumulative   ", a.cumsum())
print("unique       ", np.unique(a))
print("append 7     ", np.append(a, 7))
print("insert 0 at 3", np.insert(a, 3, 0))
print("delete idx 2 ", np.delete(a, 2))
print("and still    ", a)
array         [3 1 4 1 5 9 2 6]
length        8 or a.size 8
sum           31
smallest      1 at index 1
largest       9 at index 5
mean          3.875
sorted        [1 1 2 3 4 5 6 9]
the original is untouched: [3 1 4 1 5 9 2 6]
reversed      [6 2 9 5 1 4 1 3]
is 5 present  True
where a > 3   [2 4 5 7]
cumulative    [ 3  4  8  9 14 23 25 31]
unique        [1 2 3 4 5 6 9]
append 7      [3 1 4 1 5 9 2 6 7]
insert 0 at 3 [3 1 4 0 1 5 9 2 6]
delete idx 2  [3 1 1 5 9 2 6]
and still     [3 1 4 1 5 9 2 6]

Two things in that run are the marks.

np.sort(a) returns a sorted copy and leaves a alone. a.sort(), with no np. in front, sorts in place and returns None. Writing a = a.sort() therefore destroys the array, which is a classic mistake.

append, insert and delete all return a new array. They do not change the original, as the last line proves. A NumPy array has a fixed size: there is no such thing as growing one. If you need to add items repeatedly, build a Python list and convert it once at the end.

munotes.in30

Practical 3: Arrays, Basic Operations, Indexing and Slicing

Procedure

  1. Save the program as practical3a.py and import numpy as np.
  2. Create arrays with np.array, np.zeros, np.arange and np.linspace and print each.
  3. Index the first, the last and a middle item, using a negative index for the last.
  4. Take at least six slices, including a[::2] and a[::-1], and print each with a label.
  5. Slice a two dimensional array in both directions.
  6. Show that a list slice is a copy and an array slice is a view, by writing through both.
  7. Run the basic operations: sum, min, max, mean, sort, search, append, insert, delete.
  8. Record the dtype and itemsize your own machine printed.

Result

Arrays were created five ways, indexed from both ends, sliced eleven ways including reversal and step, and sliced in two dimensions. Writing through a list slice left the list unchanged; writing through an array slice changed the array, proving the slice is a view. Every basic operation ran and np.sort, append, insert and delete all left the original array untouched.

Where marks are lost

  • Not saying which kind of array you used. Name it: a NumPy array, or array.array.
  • Expecting a[1:4] to include item 4. The stop is always excluded.
  • a = a.sort(), which sets a to None because the in place sort returns nothing.
  • Expecting np.append to change the array. It returns a new one.
  • Trying to grow an array in a loop. Build a list, then convert once.
  • Not knowing a slice is a view. This is the question on this exercise.
  • Using m[1][2] and not knowing that m[1, 2] is the NumPy form.

For the journal

The aim in MU's words. A short table of the three things called an array in Python, with one line on each. The program, with every print labelled so the output can be read against the code. The full output. Then the list against array comparison, with both printed before and after writing through the slice, and one sentence: a list slice is a copy, a NumPy slice is a view onto the same memory. The conclusion: indexing a single item raises an error when out of range, while slicing clips silently, and a NumPy array has a fixed size so append, insert and delete all return new arrays.

Quick revision

  • Three arrays in Python: list (anything, flexible), array.array (one numeric type,

standard library, type code such as 'i'), numpy.ndarray (one type, n dimensions, the mathematics). MU's exercises mean NumPy.

  • Create: np.array, np.zeros, np.ones, np.full, np.arange, np.linspace,

np.eye.

  • arange takes a step and excludes the stop; linspace takes a count and

includes it.

  • Index from 0. a[-1] is the last. Out of range raises IndexError.
  • Two dimensions: m[row, col]. m[:, 1] is a whole column.
  • Slice start:stop:step, stop excluded. a[::-1] reverses, a[::2] steps.
  • An empty slice and a stop past the end are not errors.
  • A list slice is a copy. A NumPy slice is a VIEW. x.base is None tells them apart.
  • np.sort(a) returns a copy; a.sort() sorts in place and returns None.
  • np.append, np.insert, np.delete all return new arrays. Array size is fixed.
munotes.in31

Practical 3: Arrays, Basic Operations, Indexing and Slicing

Questions you should be able to answer

1. Name the three things in Python that get called an array, and what each holds. A list, which holds anything; an array.array, a compact array of one numeric type from the standard library; and a NumPy ndarray, one type with n dimensions and mathematics built in.

2. What does a[1:4] give? Items at indexes 1, 2 and 3. The stop is excluded.

3. How do you reverse an array in one expression? a[::-1].

4. a[5] on a three item array raises an error, but a[1:100] does not. Why? Indexing asks for one item that must exist. Slicing asks for whatever is in a range and clips to what there is, so that code can take a piece of something of unknown length.

5. What is the difference between slicing a list and slicing a NumPy array? A list slice is a new list, so writing to it leaves the original alone. A NumPy slice is a view onto the same memory, so writing to it changes the original.

6. How do you find out whether an array is a view or owns its data? x.base is the array it looks into, and it is None when the array owns its own memory.

7. What is wrong with a = a.sort()? a.sort() sorts in place and returns None, so a becomes None. Use b = np.sort(a) for a sorted copy, or a.sort() on its own line.

8. Why does np.append not grow the array? A NumPy array is one fixed block of memory. np.append makes a new, bigger array and copies into it, which is why appending in a loop is slow and building a list first is right.

9. What is the difference between arange(0, 1, 0.25) and linspace(0, 1, 5)? The first takes a step of 0.25 and stops before 1, giving four values. The second asks for five values evenly spaced and includes 1.

munotes.in32

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Report or request
Done!