munotes®

Wireless and Sensor Networks Notes | B.Sc. (Computer Science) Semester 5 | Mumbai University | munotes

Get access to whole semester resourcesSemester Pass

Official Notes munotes.in

Wireless and Sensor Networks

B.SC. (COMPUTER SCIENCE) · SEMESTER 5

Strictly as per the University of Mumbai NEP syllabus in force for B.Sc. (Computer Science)

For B.Sc. (Computer Science) students of the University of Mumbai and all its affiliated colleges

Open the book ↓

munotes.in Third Year

Wireless and Sensor Networks

Copyright © 2026 munotes.in. All rights reserved.

Written and first published by munotes.in, 2026.

This book is free for individual students to read at munotes.in. No part of it may be reproduced, distributed, stored, translated or used for institutional or classroom purposes in any form without a prior written licence from munotes.in.

Licensing and permissions: contact@munotes.in

The text of statutes and of judgments reproduced in this book is in the public domain under section 52(1)(q) of the Copyright Act 1957. The commentary, arrangement, examples and questions are the original work of munotes.in.

munotes.in is an independent study resource for MU students. It is not affiliated with, endorsed by, or officially connected to the University of Mumbai. Course names and university references describe the students and syllabus the material relates to.

munotes.in

Contents

Module I Wireless sensor networks: the node and the network, operating systems and ad hoc networks, medium access control, routing, transport and middleware

  1. What a Wireless Sensor Network Is 1
  2. The Architectural Elements of a Sensor Network 8
  3. The Sensor Network Protocol Stack and Its Three Planes 15
  4. The Advantages of Wireless Sensor Networks 20
  5. The Challenges of Wireless Sensor Networks 25
  6. Applications of Wireless Sensor Networks 34
  7. Inside a Sensor Node: The Five Units 39
  8. The Radio, the Sensors and the Power Supply of a Node 45
  9. How Long a Node Lasts: The Energy Budget Worked Out 52
  10. Sensor Taxonomy 58
  11. The Operating Environment and the Design Factors 64
  12. Radio Technology in WSNs: The Sensor Radio and Its Link Budget 70
  13. The Wireless Technologies a Sensor Network Can Use 76
  14. Network Architecture: Sources, Sinks, Hops and Mobility 82
  15. Single Hop or Multiple Hops: The Energy Argument Worked Out 87
  16. Optimization Goals: Quality of Service, Energy Efficiency and Lifetime 93
  17. Figures of Merit: Scalability, Robustness and Measuring a Network 98
  18. Deployment and Coverage: Random Against Grid 102
  19. Design Principles: Distributed Organisation and In-network Processing 107
  20. Design Principles: Data Centricity, Location, Activity and Heterogeneity 112
  21. Service Interfaces of a WSN 117
  22. Gateway Concepts 121
  23. Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP 126
  24. Why a Sensor Node Needs an Operating System 133
  25. Event-driven or Multithreaded: The Two Execution Models 137
  26. TinyOS: Components, Tasks and the Scheduler 142
  27. Commands, Events and Split-phase Operation 147
  28. nesC: Modules, Configurations, Interfaces and Wiring 152
  29. Blink: A TinyOS Application Read Line by Line 157
  30. TOSSIM: Simulating Motes, Radio Gain and Packet Loss 161
  31. Contiki, RIOT and the Other Sensor Operating Systems 174
  32. Ad Hoc Networks: MANETs, and How a Sensor Network Differs 185
  33. The Characteristics and Challenges of Ad Hoc Networks in a WSN 193
  34. Time Synchronisation and Localisation 202
  35. Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid 212
  36. AODV: Route Discovery and Route Maintenance 220
  37. DSR, and What Its Routes Cost Against AODV 229
  38. Measuring a MANET Protocol: Throughput, Delivery Ratio and Delay 237
  39. Energy Efficiency in Ad Hoc Networks: Where the Energy Goes 243
  40. Energy-aware Routing 250
  41. Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks 257
  42. Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood 265
  43. Keys and Link Security: Key Predistribution, SPINS and 802.15.4 273
  44. Privacy in Ad Hoc and Sensor Networks 281
  45. MAC Protocols for Sensor Networks: The Job and Where the Energy Goes 289
  46. Contention: ALOHA and CSMA 298
  47. Hidden and Exposed Terminals, and RTS and CTS 309
  48. CSMA/CA Worked Step by Step 317
  49. TDMA and Schedule-based MAC 325
  50. Duty Cycling: Preamble Sampling, B-MAC and X-MAC 332
  51. S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step 341
  52. S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing 348
  53. S-MAC: Latency, Adaptive Listening and the Energy Saved 355
  54. Routing Challenges and Design Issues in WSNs 363
  55. Routing Strategies in WSNs: A Map 370
  56. Flooding, Gossiping and the Broadcast Storm 377
  57. SPIN: Negotiating Before Sending 385
  58. Directed Diffusion and Rumour Routing 392
  59. LEACH: Clusters That Take Turns 401
  60. PEGASIS, TEEN and the Other Hierarchical Protocols 409
  61. Geographic Routing: Greedy Forwarding and GPSR 417
  62. Routing Tables and What Happens When the Topology Changes 425
  63. IEEE 802.15.4: The Standard, Its Devices and Its Topologies 433
  64. The 802.15.4 Physical Layer 440
  65. The 802.15.4 Superframe and Guaranteed Time Slots 454
  66. CSMA-CA, Data Transfer and Frames in 802.15.4 466
  67. Built on 802.15.4: Zigbee Routing, Security and the Later Amendments 479
  68. Traditional Transport Control Protocols: TCP and UDP 492
  69. Why a Sensor Network Cannot Simply Run TCP 502
  70. Transport Protocol Design Issues in WSNs 511
  71. Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST 521
  72. WSN Middleware: Why It Is Needed, and Its Architecture 531
  73. Middleware Approaches, and TinyDB 540

Module II Wireless transmission: frequencies, signals, antennas, propagation, multiplexing, modulation and spread spectrum, then cellular systems and GSM, DECT, TETRA and UMTS, and satellite and broadcast systems

  1. Frequencies for Radio Transmission 550
  2. Signals: Amplitude, Frequency and Phase 559
  3. Antennas: Radiators, Dipoles and Radiation Patterns 567
  4. Directional Antennas, Sectorisation, Diversity and Spatial Reuse 576
  5. Signal Propagation: Ranges, Path Loss and How a Signal Travels 585
  6. Multipath, Fading and the Doppler Effect 594
  7. Multiplexing: Space, Frequency, Time and Code 602
  8. Modulation: ASK, FSK and PSK 611
  9. Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM 619
  10. Spread Spectrum and Direct Sequence 628
  11. Frequency Hopping Spread Spectrum 636
  12. Cellular Systems: Cells, Clusters and Frequency Reuse 644
  13. Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing 653
  14. GSM and Its Mobile Services 662
  15. The GSM System Architecture 670
  16. The GSM Radio Interface: Carriers, the TDMA Frame and Bursts 678
  17. GSM Logical Channels and the Frame Hierarchy 686
  18. GSM Protocols 694
  19. Localization and Calling in GSM 701
  20. Handover in GSM 709
  21. GSM Security 716
  22. New Data Services: HSCSD, GPRS and EDGE 724
  23. DECT: System Architecture and Protocol Architecture 732
  24. TETRA 740
  25. UMTS and IMT-2000 748
  26. The UMTS System Architecture: UTRAN and the Core Network 755
  27. The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover 762
  28. A Web Request Over a Cellular Network 770
  29. The History of Satellite Systems 778
  30. Applications of Satellite Systems 785
  31. Satellite Basics: Orbits, Periods, Elevation and Footprints 793
  32. GEO: The Geostationary Orbit 803
  33. LEO and MEO 813
  34. Routing in Satellite Systems 822
  35. Localization and Handover in Satellite Systems 832
  36. Broadcast Systems: Cyclic Repetition, DAB and DVB 840
  37. Across the Two Modules: The Comparisons Question 3 Draws On 851
munotes.in

Module I

Wireless sensor networks: the node and the network, operating systems and ad hoc networks, medium access control, routing, transport and middleware

munotes.in

Chapter One

What a Wireless Sensor Network Is

Syllabus topic Module 1, "Introduction and Overview of WSNs"

In one line

A wireless sensor network is a large number of small, battery-powered devices that measure something about the physical world and pass their readings to one collection point by radio, usually hopping through each other on the way.

In the wording a student can write in an examination: a wireless sensor network (WSN) is a network of a large number of sensor nodes, each combining a sensing unit, a processing unit, a radio transceiver and a power source in one small low-cost device, densely deployed inside or very close to the phenomenon being observed. The nodes organise themselves, cooperate by processing data locally, and deliver it over a multi-hop, infrastructureless wireless network to a sink (base station), from which it reaches the task manager and the user over the Internet or a satellite link.

The survey that gave the subject its standard vocabulary puts the central idea in one phrase: the nodes are "densely deployed either inside the phenomenon or very close to it". Everything else in this book follows from taking that phrase seriously.

Why such a network exists at all

Suppose you need to know the soil moisture of a vineyard, the temperature inside two hundred grain silos, or whether a bridge is vibrating more than it did last year. There were always three ways to do it, and all three are bad at scale.

Send a person. Somebody walks the field with an instrument and a notebook. It is cheap for one reading and hopeless for a reading every fifteen minutes at forty places, day and night, for a season.

Wire the sensors. Industrial plants do this, and it works, but every sensor needs a cable for power and a cable for data, and the cabling costs more than the sensor. A cable cannot be run through a forest, across a glacier or into a bird's burrow.

Put one expensive sensor far away and look at the whole scene. A weather station at the edge of the farm, or a camera on a tower. It sees the average and misses the detail, and the detail (the dry corner, the one hot silo) is usually what matters.

A wireless sensor network is the fourth way: many cheap sensors, placed right where the phenomenon is, each with its own battery and radio, so nothing has to be wired and nobody has to visit. The survey lists the property that makes this possible: the position of the nodes "need not be engineered or predetermined", so they can be scattered in terrain nobody can reach, and the network sorts itself out afterwards.

The four things every node does

A sensor node is a small computer with four jobs, and the next several chapters take them one at a time.

munotes.in1

What a Wireless Sensor Network Is

  1. It senses. A sensor turns a physical quantity (temperature, moisture, light, vibration, sound) into an electrical signal, which is converted into a number.
  2. It processes. A microcontroller, a very small and very frugal processor, decides what to do with the number: store it, compare it with a threshold, combine it with others, or send it.
  3. It communicates. A low-power radio sends the result to a neighbouring node, and also receives and forwards other nodes' packets. A packet is one small parcel of data sent over the network in one go.
  4. It survives on its own energy. Almost always a pair of batteries, sometimes helped by a small solar cell. Nobody is expected to change them for months.

By 2000 a Berkeley group could already build such a node from shop-bought parts, "on the scale of a square inch in size and a fraction of a watt in power", and wrote a whole operating system for it, TinyOS, which the chapter [TinyOS: Components, Tasks and the Scheduler] takes apart.

The picture: the sensor field, the sink and the way out

The sensor field, the multi-hop path from an event to the sink, and the route from the sink to the user

Figure 1.1 A wireless sensor network: nodes in the sensor field, a multi-hop path to the sink, and the way out to the user

Read the figure from the event outwards, because that is the direction the data goes.

  • The sensor field is the area being watched: the vineyard, the forest, the factory floor. The small circles are sensor nodes scattered over it.
  • An event happens near node A: the soil there has dried out. Node A senses it.
  • A cannot reach the sink directly, because a low-power radio reaches tens of metres, not the whole field. So A sends its reading to its neighbour B, B to C, and so on through D and E to the sink. Each of those steps is a hop, and the whole route is a multi-hop path. Every node is both a source of its own data and a router for other nodes' data.
  • The sink, also called the base station, is the collection point. It is usually better equipped: mains power or a big battery, more memory, and a second way out.
  • The sink reaches the task manager node over the Internet or a satellite link. The task manager is where the network is told what to do (what to measure, how often, what to report) and where the results arrive.
  • The user is the person or program that actually wants the answer: the farmer, the biologist, the control room.

Notice that there are no routers, cables or towers inside the field. That is what "infrastructureless" means: the nodes themselves are the network. The only infrastructure is at the edge, at the sink and beyond it.

munotes.in2

What a Wireless Sensor Network Is

How a WSN differs from the networks you already know

A student who has studied the Internet brings expectations that are wrong here, and each wrong expectation costs marks. Six differences, each with the reason.

  1. Energy, not bandwidth, is the scarce resource. A laptop's network card is designed for speed. A sensor node's radio is designed to last a year on two AA cells, and it does so by being switched off almost all the time. Every protocol in this book is judged first on the energy it costs.
  2. The network exists for data, not for conversations between named machines. On the Internet you ask a particular address for something. In a WSN the user asks a question about the world (which part of the field is below 20 per cent moisture?) and does not care which node answers. This is called data-centric operation, and the chapter [Design Principles: Data Centricity, Location, Activity and Heterogeneity] develops it.
  3. The traffic flows mostly one way, towards the sink. Many sources, one destination. The technical name for this many-to-one pattern is convergecast. An Internet router expects traffic between any pair of hosts; a sensor network is shaped like a funnel.
  4. The nodes cooperate and compute inside the network. Rather than send every raw reading to the sink, nodes can combine readings on the way (the average of ten neighbours sent once instead of ten readings sent ten times). The survey calls this the "cooperative effort of sensor nodes": they "transmit only the required and partially processed data".
  5. Nodes are many, cheap, unattended and expected to fail. A deployment may have hundreds or thousands of nodes. Some will die, be eaten, be washed away or run flat, and the network is expected to carry on. Designing for failure is normal here, not an afterthought.
  6. Every network is built for one application. The Internet carries everything for everybody. A WSN is designed, often down to its radio schedule, for the one job it was deployed to do. That is why there are dozens of WSN protocols: each suits a different kind of task.

Where the idea came from

Dargie and Poellabauer, the first reference book on MU's list, trace the history, and it is worth knowing in outline.

  • Military research first. In 1978 the United States Defense Advanced Research Projects Agency (DARPA) organised a Distributed Sensor Nets Workshop, on networking, signal processing and distributed algorithms for sensors. DARPA then ran a Distributed Sensor Networks (DSN) programme in the early 1980s, followed by the Sensor Information Technology (SensIT) programme.
  • Putting a whole node on a chip. The University of California at Los Angeles, with the Rockwell Science Center, proposed Wireless Integrated Network Sensors (WINS); one result, in 1996, was a single CMOS chip carrying sensors, interface circuits, signal processing, a radio and a microcontroller.
  • Smaller still. Berkeley's Smart Dust project set out to show that a complete sensor system could be built into a device perhaps the size of a grain of sand. Berkeley's PicoRadio project and MIT's microAMPS project worked on nodes frugal enough to be powered by their surroundings.
  • Commercial motes. From the Berkeley work came the motes and TinyOS, and companies selling ready-made nodes.
munotes.in3

What a Wireless Sensor Network Is

A real one: Great Duck Island, 2002

The deployment that showed the idea working outside a laboratory was on Great Duck Island, off the coast of Maine in the United States. Biologists wanted to study a seabird, Leach's Storm Petrel, which nests in underground burrows and is likely to abandon a burrow if it is disturbed. People walking about the colony were the problem the study was trying to avoid.

In July 2002 a team from Intel Research, the University of California at Berkeley and the College of the Atlantic deployed thirty-two motes, nine of them inside burrows, measuring temperature, relative humidity, barometric pressure, light and infrared (which shows whether a warm bird is at home). Readings went from the motes to a gateway beside the patch, across a longer radio link to a base station with a database and a wide-area connection, and from there to researchers over the Internet. The design goal was a network that runs for nine months on non-rechargeable batteries, so that nobody needed to visit during the breeding season.

Every element of the figure above is there: the sensor field (the colony), the nodes (the motes), the sink and gateway (the island's base station), the way out (the Internet), and the user (the biologists). And the reason for using a WSN at all is exactly the one this chapter gave: the measurement had to be taken where the phenomenon was, without people.

Worked example: a vineyard near Nashik

A grape grower near Nashik wants to irrigate only when the soil needs it. The farm's engineer, Meera, plans a WSN.

The task. Measure soil moisture and temperature at 40 points across the vineyard every 15 minutes, and tell the pump controller when any zone falls below the moisture threshold.

Step 1: what each node produces. A reading every 15 minutes is

24 × 60 / 15 = 96

readings per node per day, so the whole network produces 40 × 96 = 3,840 readings a day.

Step 2: how much data that is. Each reading, with the node's identity and a timestamp, fits in 12 bytes, so the day's data is 3,840 × 12 = 46,080 bytes, about 45 kilobytes. That is less than one photograph on a phone. Bandwidth is not the problem, and it never will be here. The problem is doing this for a season without anyone changing batteries.

munotes.in4

What a Wireless Sensor Network Is

Step 3: how the data travels. The farthest row of vines is 600 metres from the pump house, where the sink sits with mains power. The node radios reliably reach about 100 metres through the vines, so a reading from the far end needs at least 600 / 100 = 6 hops. Each node therefore forwards its neighbours' readings as well as sending its own, and the nodes nearest the pump house carry the most traffic.

Step 4: what the sink does. It collects the readings, notices that zone 7 has dropped below the threshold, and sends one message to the pump controller and one to the task manager, a small server in Nashik that the grower checks on his phone.

Step 5: what is already visible. The nodes near the sink will run out of energy first, because every other node's data passes through them. Meera will have to deal with that, and the chapters [Single Hop or Multiple Hops: The Energy Argument Worked Out] and [Energy-aware Routing] show how.

That is a whole WSN, in miniature, and it already contains the three ideas the rest of the module is about: the node's energy, the multi-hop network, and the sink as the bridge to the outside world.

Distinctions

Wireless sensor networkThe InternetA wired sensor system
Main purposeDeliver readings about the physical worldCarry any traffic between any machinesDeliver readings from fixed points
Scarcest resourceEnergy in each nodeBandwidth and latencyCabling and installation cost
AddressingData-centric: ask about the dataAddress-centric: ask a named hostEach sensor on its own wire
Traffic patternMany sources to one sink (convergecast)Any to anyEvery sensor to the controller
InfrastructureNone inside the fieldRouters, links, serversCables and junction boxes
Failure of a unitExpected, the network carries onHandled by reroutingThat point goes dark
Built forOne applicationEvery applicationOne installation

What it does not mean

A WSN is not Wi-Fi with a thermometer on it. Wi-Fi is built for tens of megabits per second and a power supply. Sensor node radios are built for a few hundred kilobits per second at most and years on batteries, and they are asleep nearly all the time. The chapter [The Wireless Technologies a Sensor Network Can Use] compares the real options.

"Wireless" does not mean every node talks to the sink. Most nodes cannot reach it. Multi-hop forwarding through neighbours is the normal case, and it is why a node spends as much energy on other nodes' data as on its own.

munotes.in5

What a Wireless Sensor Network Is

The sink is not the user. The sink is a machine at the edge of the field. The user is a person or program somewhere else, reached through the task manager.

"No infrastructure" is about the field, not the whole system. The sink, the gateway, the Internet and the task manager are all infrastructure. What a WSN does without is infrastructure among the nodes.

A WSN is not the same thing as the Internet of Things. The Internet of Things is any device on the Internet, from a smart television to a car. A WSN is one particular kind of network of sensing devices, which may or may not be joined to the Internet. The chapter [Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP] shows how the two meet.

Quick revision

  • WSN: many small battery-powered sensor nodes, densely deployed inside or near the phenomenon, which sense, process and send data over a multi-hop, infrastructureless radio network to a sink.
  • A node has four jobs: sense, process, communicate, survive on its own energy.
  • The path: sensor field, nodes, multi-hop route, sink (base station), Internet or satellite, task manager node, user.
  • Nodes need not be placed by design, so the network must self-organise.
  • Six differences from the Internet: energy is scarce; data-centric; convergecast to one sink; nodes compute and cooperate; nodes are many, cheap and fail; one network per application.
  • History (Dargie and Poellabauer): DARPA workshop 1978, DSN programme early 1980s, then SensIT; UCLA WINS; Berkeley Smart Dust and motes.
  • First real deployment usually cited: Great Duck Island, July 2002, 32 motes, 9 in burrows, designed to run 9 months on batteries.
  • Vineyard: 96 readings a node a day, 3,840 in all, 46,080 bytes: bandwidth is not the problem, energy is.

Test yourself

1. Define a wireless sensor network. A WSN is a network of a large number of low-cost, low-power sensor nodes, each with sensing, processing, radio and power units, densely deployed inside or close to a phenomenon. The nodes self-organise, process data locally and send it over a multi-hop, infrastructureless wireless network to a sink, which passes it through the Internet or a satellite link to a task manager and the user.

2. Draw and explain the elements of the network from an event to the user. The figure: sensor nodes in the sensor field; a node near the event senses it; the reading hops node to node to the sink; the sink sends it over the Internet or satellite to the task manager node; the user reads it there. Each intermediate node acts as a router.

munotes.in6

What a Wireless Sensor Network Is

3. Why do WSN nodes use multi-hop forwarding instead of sending straight to the sink? Because their radios are low-power and reach only tens of metres, far less than the size of the field. Sending a long distance directly would also cost far more energy than several short hops in most cases, which is examined in [Single Hop or Multiple Hops: The Energy Argument Worked Out].

4. State four ways a WSN differs from the Internet. Energy rather than bandwidth is the scarce resource; operation is data-centric rather than address-centric; traffic converges on one sink rather than flowing between any pair of hosts; and nodes cooperate and process data inside the network. Also: nodes are many, cheap and expected to fail, and each network is built for one application.

5. In the vineyard, 40 nodes each report every 15 minutes with 12-byte readings. How much data does the network produce a day, and what does that tell you? 96 readings a node a day, 3,840 readings, 46,080 bytes, about 45 kilobytes. It tells you that the design problem is energy and lifetime, not data rate.

6. What did the Great Duck Island deployment show? That a WSN could monitor a phenomenon (a seabird colony) where people's presence would itself disturb it: 32 motes, nine in burrows, deployed in July 2002, sending data through an island gateway to researchers over the Internet, designed to run nine months on batteries.

Contents This chapter on its own page

munotes.in7

Chapter Two

The Architectural Elements of a Sensor Network

Syllabus topic Module 1, "Introduction and Overview of WSNs: Basic sensor network architectural elements"

In one line

A sensor network is built from five kinds of element: the sensor nodes, the wireless links between them, the sink or base station that collects their data, the gateway that joins the network to the outside world, and the task manager and user who ask the questions and read the answers.

In the wording a student can write in an examination: the basic architectural elements of a wireless sensor network are (1) sensor nodes, which sense, process and relay data, some of them taking special roles as cluster heads or actuators; (2) the interconnecting wireless network of radio, infrared or optical links, arranged as a star, mesh or cluster tree; (3) the sink or base station, a better-resourced node where data is collected and from which tasks are issued; (4) the gateway, which connects the sensor network to other networks such as the Internet, a satellite or a cellular network; and (5) the task manager and back-end computing resources, through which the user tasks the network and receives its results.

Why the elements are worth naming separately

Because each one has a different job, a different budget and a different failure, and every design decision in this book belongs to one of them. A routing protocol is a property of the links and the nodes. A query language belongs to the task manager. Deciding where the data is stored depends on the sink. A student who writes the sensors send data to the server has merged four elements into two words and cannot then say where anything goes wrong.

Element 1: the sensor nodes

A sensor node is the small device in the field. The survey that fixed the vocabulary gives it four basic units: a sensing unit (a sensor and an analogue-to-digital converter), a processing unit with a little storage, a transceiver unit (the radio) and a power unit. Some nodes carry up to three optional units, depending on the application: a location finding system, a power generator such as a solar cell, and a mobilizer that can move the node. The chapter [Inside a Sensor Node: The Five Units] takes each unit apart.

In the network, nodes play one or more of these roles:

  • Source node. It senses the phenomenon and produces data. In most networks every node is a source.
  • Relay node. It forwards other nodes' packets towards the sink. In a multi-hop network every node is also a relay, which is why a node near the sink carries far more traffic than it produces.
  • Cluster head. In a hierarchical network one node in a group collects its neighbours' data, combines it, and forwards the result. It works much harder than the others, so either it has more energy or the role is rotated, which is the whole idea of the protocol in [LEACH: Clusters That Take Turns].
  • Actuator node. Some networks do not only sense but also act: they open a valve, sound an alarm, switch a pump. A network with both is called a wireless sensor and actuator network.
munotes.in8

The Architectural Elements of a Sensor Network

Element 2: the interconnecting wireless network

The nodes are joined by wireless links. The survey names three media: radio, which almost every real node uses; infrared, which needs no licence and resists electrical interference but needs a line of sight; and optical links, which the Berkeley Smart Dust design used, again needing a clear line of sight. Radio wins because it does not need the two ends to see each other.

How the links are arranged is the network's topology, and three arrangements cover almost every real deployment.

Star, mesh and cluster tree topologies, stacked

Figure 2.1 The three topologies a sensor network is built in

Star. Every node talks directly, in one hop, to the base station, and nodes never relay for each other. It is the simplest to build and gives the lowest delay, and the nodes can sleep whenever they are not sending. Its limit is range: every node must be within radio reach of the base station, and if the base station fails, everything stops.

Mesh. Nodes talk to their neighbours and relay for each other, so data reaches the base station over several hops, by whichever path is working. It covers a large area with short-range radios and survives the failure of any single node, because another path exists. It costs energy and delay at every hop, and the relays near the base station carry the heaviest load.

Cluster tree, or hierarchical. Nodes are grouped into clusters; each cluster has a cluster head that collects its members' data and forwards it, over one or more hops of cluster heads, to the base station. It combines the two: members do short one-hop transmissions like a star, cluster heads form a backbone like a mesh, and the cluster head can aggregate (combine) its members' readings so that less data travels the long distance.

Element 3: the sink, or base station

The sink is where the data ends up inside the sensor network. The survey draws it at the edge of the sensor field, and the word "sink" (the place data drains into) and "base station" are used for the same element.

What makes it different from an ordinary node:

  • More resources. Usually mains power or a large battery, more memory and a faster processor, because it receives everybody's traffic and cannot be allowed to run flat.
  • Two roles. It is the destination of the collected data and the origin of the tasks and queries sent into the network. Traffic in a sensor network flows in both directions through the sink: tasks out, data in.
  • Not necessarily one. Large deployments use several sinks so that no single one is a bottleneck, and some networks use a mobile sink, a vehicle or a person carrying a receiver, that visits the nodes. The chapter [Network Architecture: Sources, Sinks, Hops and Mobility] sets out those scenarios.
munotes.in9

The Architectural Elements of a Sensor Network

Element 4: the gateway

The gateway is the element that connects the sensor network to a different network: the Internet, a satellite link, a cellular network or a company's own network. It has to speak both sides: the sensor network's low-power radio protocol on one side and Internet protocols on the other, translating between them.

In a small deployment the sink and the gateway are the same box. In a large one they are separated. On Great Duck Island in 2002 each patch of sensors had its own gateway, which passed data across a longer "transit" radio link to a base station that held a database and the wide-area connection to the mainland. That arrangement, gateway then transit network then base station, is a real system's version of the elements in this chapter. The gateway's design problems (how an Internet host can address a node, and how a node's data can be made to look like an Internet service) are the chapter [Gateway Concepts].

Element 5: the task manager and the user

The task manager node, in the survey's figure, is where the network is told what to do and where its results are stored and analysed. It may be a server in an office or a cloud service. Behind it are the computing resources that turn raw readings into something useful: storage, analysis that spots trends and correlations, answering the user's queries, and raising alarms.

The user is the person or program that wants the information: a farmer checking a phone, a control room, a researcher's database. The user almost never deals with individual nodes. The user asks about the phenomenon (the moisture in zone 7, the temperature in cold store 3), and the lower elements decide which nodes answer.

The two directions of flow

Every element takes part in two flows, and a complete answer mentions both.

  1. Tasking, outward. The user's request goes to the task manager, through the gateway and the sink, and is spread (disseminated) into the field: which quantity to measure, how often, and what to report.
  2. Collection, inward. Readings go from source nodes, hop by hop or through cluster heads, to the sink, and from there through the gateway to the task manager and the user.
munotes.in10

The Architectural Elements of a Sensor Network

The elements as MU's first text book sets them out

MU's printed label, "Basic sensor network architectural elements", is the title of section 1.2.1 of Sohraby, Minoli and Znati, the first text book on MU's list. That section describes the elements not only as boxes on a diagram but as the technologies a sensor network is built from, and an examiner who follows the text book may expect them. There are nine, and every one has a home later in this book.

Element in the text bookWhat it coversWhere this book teaches it
Sensor types and technologyThe node, its sensing units (single or array), processing, communication and power units, and the sensor field with its sink[Inside a Sensor Node: The Five Units]
Software: operating systems and middlewareComponent-based operating systems such as TinyOS, with an event-driven model that allows fine-grained power management[Why a Sensor Node Needs an Operating System]
Standards and the protocol stackA generic stack with power, mobility and task management planes, and the lower-layer standards that can carry it[The Sensor Network Protocol Stack and Its Three Planes]
Routing and data disseminationData-centric routing, directed diffusion, aggregation, energy-aware and location-based routing[Routing Strategies in WSNs: A Map]
Network organisation and trackingSelf-organisation, group management, and detecting, classifying and tracking targets, with coverage and detectability[Deployment and Coverage: Random Against Grid]
ComputationAggregation, fusion and analysis inside the network, near the source of the data[Design Principles: Distributed Organisation and In-network Processing]
Data managementWhere data is stored and how it is queried: centrally, or distributed through the network[Middleware Approaches, and TinyDB]
SecurityConfidentiality, integrity and availability[Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks]
Network design issuesReliable transport, limited bandwidth and power, self-configuration, and the design factors[The Operating Environment and the Design Factors]

The text book places all of these in what it calls the C1WSN environment, the large multi-hop kind of sensor network, and lists the conditions that make them hard: a very large sensor population, large streams of data, incomplete or uncertain data, frequent failure of nodes and of links, limited power and processing, a multi-hop topology, no node with global knowledge of the network, and often little administrative support. Those conditions are the challenges of [The Challenges of Wireless Sensor Networks] seen from the architect's side.

Worked example: fire detection in a twelve-storey office building

An office building in Andheri has 12 floors. The facilities manager, Farhan, installs a sensor network for early fire detection.

The elements.

ElementIn this building
Sensor nodes20 smoke and temperature nodes on each floor, battery powered, mounted on the ceiling
Cluster headsOne mains-powered node per floor, beside the electrical riser
Links and topologyCluster tree: ceiling nodes reach their floor's cluster head in one hop; cluster heads relay up the riser to the base station
Sink or base stationA panel in the ground-floor security room
GatewayA 4G router in the same panel, connected to the monitoring company's server
Task manager and userThe monitoring company's server; the security guard and the fire officer
munotes.in11

The Architectural Elements of a Sensor Network

Why the hierarchy matters, in numbers. Each ceiling node reports a "still healthy" status every 10 minutes, which is 6 × 24 = 144 reports a day. The building has 12 × 20 = 240 nodes, so if every report travelled all the way to the base station there would be 240 × 144 = 34,560 reports a day.

With cluster heads, each floor's 20 reports are combined into one summary (floor 7: 20 nodes healthy, highest temperature 31 degrees), so the base station receives 12 × 144 = 1,728 summaries a day. That is 34,560 / 1,728 = 20 times fewer messages crossing the building, and the saving is exactly the number of nodes each cluster head serves.

What happens in a fire. A node on floor 9 detects smoke. It does not wait for its ten-minute slot: it sends an alarm at once to its cluster head, which forwards it immediately to the base station, and the gateway sends it to the monitoring company and the fire brigade. The summarising that saves energy in normal times is bypassed for an alarm, because an alarm must never wait. Deciding which traffic is periodic and which is urgent is part of every sensor network's design, and the chapter [Applications of Wireless Sensor Networks] names the four patterns.

Distinctions

Sink or base stationGatewayTask manager
WhereAt the edge of the sensor fieldBetween the sensor network and another networkIn the outside world, often an office or cloud server
DoesCollects data, issues tasks into the fieldTranslates between protocols, forwards across networksStores and analyses data, takes the user's requests
SpeaksThe sensor network's own protocolsBoth sidesInternet protocols
Often the same box asThe gatewayThe sinkNeither
Source nodeRelay nodeCluster head
Produces its own dataYesPossiblyYes
Forwards others' dataNoYesYes, its whole cluster's
Combines dataNoRarelyUsually
Energy loadLowestRises with nearness to the sinkHighest, so it is rotated or better powered
StarMeshCluster tree
Hops to the base stationOneManyOne to the head, then more
CoverageLimited to one radio rangeLargeLarge
If one node failsOnly that node is lostTraffic takes another pathOnly that cluster is at risk, unless the head fails
If the base station failsEverything stopsEverything stopsEverything stops
EnergyLowest for membersHighest for relays near the sinkConcentrated in cluster heads
munotes.in12

The Architectural Elements of a Sensor Network

What it does not mean

The sink and the gateway are not always two boxes. They are two functions. In most small networks one device does both, and an answer should say so.

A cluster head is not a different kind of hardware. It can be an ordinary node that has taken on the role for a while. Rotation is the point of several protocols.

The base station failing is not survived by a mesh. A mesh survives the loss of relays. It does not survive the loss of the one place all the data is going, which is why critical networks use more than one sink.

Topology is not the same as geography. Two nodes a metre apart may have no link between them if a steel beam is in the way, and two nodes forty metres apart may have a good one. The topology is the graph of links that actually work.

Quick revision

  • Five elements: sensor nodes, wireless links, sink or base station, gateway, task manager and user.
  • Node units: sensing, processing, transceiver, power; optional location finding system, power generator, mobilizer.
  • Node roles: source, relay, cluster head, actuator.
  • Media: radio (used almost everywhere), infrared and optical (both need line of sight).
  • Topologies: star (one hop, simple, short range), mesh (many hops, robust, costly relays), cluster tree (members to heads, heads to the base station, aggregation at heads).
  • Two flows: tasks out, data in, both through the sink.
  • The text book's nine elements (Sohraby and colleagues, section 1.2.1): sensor types and technology; software; standards and the stack; routing and dissemination; organisation and tracking; computation; data management; security; network design issues.
  • Building example: 34,560 reports a day without cluster heads, 1,728 with them, 20 times fewer; alarms bypass the summarising.

Test yourself

1. List and explain the basic architectural elements of a WSN. Sensor nodes, which sense, process and relay; the wireless links arranged in a topology; the sink or base station that collects data and issues tasks; the gateway that connects the network to the Internet, satellite or cellular networks and translates protocols; and the task manager with its computing resources, through which the user tasks the network and receives results.

2. Compare the star, mesh and cluster tree topologies. Star: one hop to the base station, simple, low delay, limited range, single point of failure. Mesh: multi-hop relaying, large coverage, robust to node failure, extra energy and delay per hop. Cluster tree: members send one hop to cluster heads, which aggregate and forward over a backbone, combining the star's cheap members with the mesh's reach.

munotes.in13

The Architectural Elements of a Sensor Network

3. What is the difference between the sink and the gateway? The sink is the collection point of the sensor network and the source of tasks sent into it. The gateway connects the sensor network to another network and translates between their protocols. They are two functions that are often, but not always, in the same device.

4. Why does a cluster head need more energy than an ordinary node, and what are the two answers to that? It receives and forwards its whole cluster's data, often over a longer link. Either it is given more energy (mains power, a bigger battery) or the role is rotated among the nodes so the load is shared.

5. In the office building, why does the base station receive twenty times fewer messages with cluster heads? Because each cluster head combines its 20 members' status reports into one summary. 240 nodes at 144 reports a day is 34,560 reports; 12 cluster heads at 144 summaries a day is 1,728, and 34,560 / 1,728 = 20.

Contents This chapter on its own page

munotes.in14

Chapter Three

The Sensor Network Protocol Stack and Its Three Planes

Syllabus topic Module 1, "Introduction and Overview of WSNs: Basic sensor network architectural elements"

In one line

The protocol stack of a sensor network is the usual five layers, from physical to application, crossed by three planes that manage power, mobility and tasks across all of them.

In the wording a student can write in an examination: the protocol stack used by the sink and the sensor nodes consists of the physical layer, data link layer, network layer, transport layer and application layer, together with the power management plane, mobility management plane and task management plane. The stack "combines power and routing awareness, integrates data with networking protocols, communicates power efficiently through the wireless medium, and promotes cooperative efforts of sensor nodes". The planes help the nodes coordinate the sensing task and lower the overall power consumption.

Why a sensor network needs its own stack

A student who has learnt the TCP/IP model may ask why the Internet's layers are not simply reused. Two reasons, and they are the reasons for the planes.

The Internet's layers do not care about energy. An Internet router forwards a packet the same way whether its battery is full or nearly empty, because it has no battery. A sensor node must take its remaining energy into account in almost every decision: whether to listen, whether to relay, whether to take a turn at sensing.

The Internet's layers keep to themselves. Each layer is meant to know nothing about the others. In a sensor network, energy, movement and the sharing of the sensing task are concerns of every layer at once. A layer cannot solve them alone, so the survey adds planes: management functions that run alongside all the layers and coordinate them.

Five layers crossed by three management planes

Figure 3.1 The sensor network protocol stack: five layers and three management planes

The five layers, bottom to top

Physical layer

The survey gives its job as "frequency selection, carrier frequency generation, signal detection, modulation, and data encryption". It turns bits into a radio signal and back.

What is different in a sensor network: the energy cost of distance. The power needed to send a signal over a distance d grows as d to the power n, where n lies between 2 and 4, and the survey notes that n is "closer to four for low-lying" antennas and near-ground channels, which is exactly where sensor nodes sit, on the ground or on a post. So doubling the distance can cost up to sixteen times the power. This one fact is why sensor networks prefer short hops and simple, robust modulation. At the time the survey was written, the 915 MHz industrial, scientific and medical band had been "widely suggested"; today most nodes use the 2.4 GHz band of IEEE 802.15.4, taught in [The 802.15.4 Physical Layer].

munotes.in15

The Sensor Network Protocol Stack and Its Three Planes

Data link layer

Its job is "the multiplexing of data streams, data frame detection, medium access and error control". It gets a frame reliably from one node to its neighbour.

Two parts, both different in a sensor network:

  • Medium access control (MAC) decides when a node may use the shared radio channel. In a sensor network it "must be power-aware and able to minimize collision with neighbors' broadcasts", and above all it must let the radio sleep, because a radio left listening wastes most of a node's energy. The chapters from [MAC Protocols for Sensor Networks: The Job and Where the Energy Goes] onwards are this sublayer.
  • Error control repairs or detects damaged frames. The survey names the two modes: forward error correction (FEC), which adds redundant bits so the receiver can correct errors itself, and automatic repeat request (ARQ), which resends a frame that did not arrive intact. A sensor node needs simple codes, because decoding a complex one costs energy too.

Network layer

It routes the data from the node that produced it to the sink, across many hops. The survey lists four design principles for this layer in a sensor network:

  1. Power efficiency is always an important consideration.
  2. Sensor networks are mostly data-centric: what matters is the data, not which node sent it.
  3. Data aggregation is useful only when it does not hinder the collaborative effort of the nodes.
  4. An ideal sensor network has attribute-based addressing and location awareness: a packet is sent to "the nodes in region A" or to the nodes reading above 70°F, not to node number 145.

The whole of [Routing Strategies in WSNs: A Map] and the chapters after it are this layer.

Transport layer

It "helps to maintain the flow of data if the sensor networks application requires it". Note the condition: many sensor applications tolerate lost readings (another reading comes in a minute), so they need little from this layer.

The survey's key point is where TCP stops. It suggests splitting the connection at the sink: the user talks to the sink over ordinary TCP or UDP through the Internet or satellite, and the sink talks to the nodes with a light UDP-type protocol, "because each sensor node has limited memory" and "acknowledgments are too costly for sensor networks". The chapter [Why a Sensor Network Cannot Simply Run TCP] works out why.

Application layer

It carries the software of the particular application: the monitoring, tracking or alarm program. The survey also proposed three application-layer protocols, which it called open research issues:

  • Sensor Management Protocol (SMP), through which administrators manage the network: introducing rules for aggregation, naming and clustering; time synchronisation; turning nodes on and off; querying and reconfiguring the network; and authentication and key distribution.
  • Task Assignment and Data Advertisement Protocol (TADAP), for spreading a user's interest (what the user wants to know) into the network, or for nodes to advertise what data they have so users can ask for it.
  • Sensor Query and Data Dissemination Protocol (SQDDP), the interface through which an application issues queries and collects replies. Queries are attribute-based ("the locations of the nodes that sense temperature higher than 70°F") or location-based ("temperatures read by the nodes in region A"), never addressed to a named node.
munotes.in16

The Sensor Network Protocol Stack and Its Three Planes

The three planes

The planes are the part of the stack an examiner is really asking about, because they are what the TCP/IP model does not have.

Power management plane

It "manages how a sensor node uses its power". The survey gives two examples, and both are worth writing in an answer:

  • A node turns off its receiver after receiving a message from one of its neighbours, so as not to receive the same message again from another neighbour.
  • When its power is low, a node broadcasts to its neighbours that it is low in power and cannot take part in routing, and keeps what is left for its own sensing.

Mobility management plane

It "detects and registers the movement of sensor nodes, so a route back to the user is always maintained". Knowing which nodes are its neighbours also lets a node balance its power and its tasks with theirs. In a network where nothing moves this plane does little; on a network of nodes carried by animals, vehicles or water currents, it is essential.

Task management plane

It "balances and schedules the sensing tasks given to a specific region". Not every node in a region needs to sense at the same time, so some nodes take on more of the task than others, depending on their remaining power. Ten nodes watching one field can take turns, and the field is watched just as well at a fraction of the energy.

Worked example: one reading, all the way down

Back to the vineyard of [What a Wireless Sensor Network Is]. Node 23, in zone 7, has a moisture reading of 18 per cent and it is time to report. Follow the reading down its stack and across the planes.

StepLayer or planeWhat happens
1Task management planeOf the three nodes in zone 7, node 23 is the one scheduled to sense this hour; the other two stay asleep
2Application layerThe monitoring program forms the message zone 7, moisture 18 per cent, 14:15
3Transport layerNo connection is set up. The reading goes as a single UDP-type datagram; if it is lost, the next reading in 15 minutes replaces it
4Network layerThe routing table says node 23's next hop towards the sink is node 17, the neighbour on the best path
5Data link layer, MACThe MAC waits until the channel is free and node 17 is awake, then sends the frame
6Data link layer, error controlThe frame carries a checksum; node 17 acknowledges it, and without the acknowledgement node 23 would resend (ARQ)
7Physical layerThe frame is modulated onto the 2.4 GHz radio and transmitted at the lowest power that reaches node 17
8Power management planeThe frame sent, node 23 switches its radio off until its next scheduled slot
9Mobility management planeNothing to do here: vines do not move. A node on the tractor would report its new position here
munotes.in17

The Sensor Network Protocol Stack and Its Three Planes

Node 17 receives the frame, passes it up only to its network layer (it is a relay, so the application layer never sees it), and sends it down its own stack towards the next hop. At the sink the reading climbs all five layers and is handed to the gateway.

Distinctions

A layerA plane
RunsAbove one layer and below anotherAcross all five layers
Deals withOne step of communicationA concern every step shares: energy, movement, sharing the task
ExampleThe network layer chooses the next hopThe power plane stops a low-battery node from routing
In TCP/IPYesNo equivalent
Sensor network stackTCP/IP stack
Designed aroundEnergy and cooperationThroughput and generality
AddressingAttribute-based and location-basedGlobal IP addresses
TransportLight, UDP-type inside the field; TCP split at the sinkTCP end to end
Cross-layer managementThree planesNone
Intermediate nodesMay process and combine dataOnly forward it

What it does not mean

The planes are not extra layers. They sit beside the stack, not in it. Drawing them as three more boxes on top of the application layer is the commonest mistake in this answer.

"Transport layer" does not mean TCP. The survey says the opposite: TCP stops at the sink, and inside the field the transport layer is light, UDP-type, or absent.

The stack is a reference model, not an implementation. Real systems merge and split the layers freely. TinyOS has no layers at all in this sense, only components, and IEEE 802.15.4 defines the physical and MAC layers together. The model is for understanding and for answers; [Design Principles: Data Centricity, Location, Activity and Heterogeneity] explains why real designs deliberately cut across layers.

Encryption at the physical layer is the survey's wording, not a rule. Most real sensor networks encrypt at the data link layer (802.15.4 does), which the chapter [Keys and Link Security: Key Predistribution, SPINS and 802.15.4] covers.

munotes.in18

The Sensor Network Protocol Stack and Its Three Planes

Quick revision

  • Stack (Akyildiz and colleagues, 2002): physical, data link, network, transport, application, plus power, mobility and task management planes.
  • Physical: frequency selection, carrier generation, signal detection, modulation, encryption; power to send over distance d goes as d to the n, n from 2 to 4, near 4 close to the ground.
  • Data link: multiplexing, frame detection, MAC (power-aware, few collisions) and error control (FEC and ARQ).
  • Network: power efficiency, data-centric, aggregation where it helps, attribute-based addressing and location awareness.
  • Transport: only if needed; TCP split at the sink, UDP-type inside.
  • Application: SMP, TADAP, SQDDP.
  • Power plane: receiver off after a message; announce low power, stop routing. Mobility plane: track movement and neighbours, keep a route to the user. Task plane: schedule and balance sensing so not every node senses at once.

Test yourself

1. Draw and explain the WSN protocol stack. Five layers (physical, data link, network, transport, application) crossed by three planes (power, mobility and task management). Physical: frequency selection, carrier generation, signal detection, modulation, encryption. Data link: multiplexing, frame detection, medium access, error control. Network: data-centric, energy-efficient routing to the sink. Transport: maintains flow where needed. Application: the application software and management protocols. Planes: coordinate energy, movement and the sensing task across all layers.

2. What are the three management planes, and give an example of each. Power management: a node turns off its receiver after receiving a message to avoid duplicates, and announces low power so it is not used for routing. Mobility management: records node movement so a route to the user is always kept and neighbours are known. Task management: schedules sensing in a region so that nodes take turns according to their power.

3. Why does a sensor network's stack need planes at all? Because energy, movement and the sharing of the sensing task concern every layer at once. No single layer can manage them, so they are handled by functions that run across all the layers.

4. What does the survey recommend for the transport layer, and why? Splitting the connection at the sink: TCP or UDP between the user and the sink over the Internet or satellite, and a light UDP-type protocol between the sink and the nodes, because nodes have little memory and acknowledgements cost too much energy.

5. State the four design principles of the network layer. Power efficiency; data-centric operation; aggregation only where it does not hinder collaboration; and attribute-based addressing with location awareness.

6. Name the three application-layer protocols the survey proposes. Sensor Management Protocol (SMP), Task Assignment and Data Advertisement Protocol (TADAP), and Sensor Query and Data Dissemination Protocol (SQDDP).

Contents This chapter on its own page

munotes.in19

Chapter Four

The Advantages of Wireless Sensor Networks

Syllabus topic Module 1, "Introduction and Overview of WSNs: Advantages and challenges of WSNs"

In one line

A wireless sensor network measures the world where things happen, at many points at once, without cables and without people, and keeps working when some of its parts fail.

In the wording a student can write in an examination, the advantages of wireless sensor networks are: (1) no cabling, so deployment is quick, cheap and possible in existing buildings; (2) sensing close to the phenomenon, giving stronger signals and finer spatial detail; (3) deployment where people and cables cannot go, including random deployment from the air; (4) fault tolerance through redundancy; (5) self-organisation, with no infrastructure in the field; (6) scalability, since nodes can be added at any time; (7) in-network processing, which reduces the data sent; (8) continuous, unattended monitoring at low running cost; (9) low cost per node; and (10) flexibility, since the network can be moved, extended or retasked.

Advantages over what?

An examiner who asks for advantages wants to know that you understand what a WSN replaced. There were three ways of measuring a physical phenomenon before, and each advantage below beats at least one of them.

  • Wired sensors, each with a cable for power and a cable for data, as in a factory control system.
  • Manual observation, somebody visiting with an instrument and a notebook.
  • A single instrument at a distance, one weather station for a whole district, or one camera on a tower.

The advantages, one by one

1. No cabling

Every node carries its own battery and radio, so nothing has to be wired. Installation becomes a matter of fixing nodes in place, and it can be done in a building that is already occupied, in a field that is being ploughed, or on a structure where drilling is not allowed. Removing or moving a node is as easy as installing it. Against wired sensors, this is the headline advantage, and the worked example below puts numbers on it.

2. Sensing close to the phenomenon

The defining phrase of a WSN is that its nodes are deployed "inside the phenomenon or very close to it". That matters for two reasons.

The signal is stronger. Many physical signals weaken with distance. The intensity of sound from a small source in open air falls with the square of the distance (the inverse-square law), so a cheap microphone at 10 metres receives a far stronger signal than an expensive one at 100 metres. The worked example calculates how much.

The detail is finer. One weather station reports one temperature for a whole farm. Forty nodes report forty, and the one dry corner or hot silo, which is usually what the user cares about, appears in the data instead of being averaged away.

munotes.in20

The Advantages of Wireless Sensor Networks

3. Deployment where people and cables cannot go

Because node positions "need not be engineered or predetermined", nodes can be scattered rather than placed. The survey lists the ways: "dropping from a plane", delivered in an artillery shell, rocket or missile, or placed one by one by a person or a robot. A forest canopy, a glacier, a flooded area, a battlefield or a chemically contaminated site can be monitored without anyone standing in it.

4. Fault tolerance through redundancy

A dense network has more nodes than the task strictly needs, so the loss of a few does not stop it. A dead node's neighbours are still sensing the same area, and a mesh has other paths for the data. A wired sensor that fails leaves that point dark; a manual reading that is missed is simply missing. The design factor behind this, with its reliability formula, is in [The Operating Environment and the Design Factors].

5. Self-organisation, and no infrastructure in the field

The nodes discover their neighbours and build their own routes after they are deployed. There are no routers, towers or access points to install inside the field, only the sink at its edge. This is what makes a WSN possible in places no infrastructure could be built.

6. Scalability and extensibility

The survey names a "redeployment of additional nodes phase": nodes can be added at any time, to replace failed ones or because the task has changed. A network that watches ten fields this year can watch twenty next year by adding nodes, without redesigning the part that already exists.

7. In-network processing

Each node has a processor, so nodes can compute before they send. Instead of forwarding every raw reading, they "use their processing abilities to locally carry out simple computations and transmit only the required and partially processed data". An average of twenty readings sent once instead of twenty readings sent twenty times saves energy and bandwidth everywhere downstream. The principle is taught in [Design Principles: Distributed Organisation and In-network Processing].

8. Continuous, unattended monitoring at low running cost

A network samples every few seconds or minutes, day and night, for months, with nobody present. Against manual observation, it gives far more readings at a fraction of the running cost, and readings arrive as they are taken, so an alarm can be raised at once rather than at the next visit.

9. Low cost per node

A sensor node is built from mass-produced chips: a microcontroller, a radio and a sensor. Because a network uses many of them, the design goal from the start has been to keep each node cheap. The survey goes as far as to say that if a sensor network costs more than deploying traditional sensors, "the sensor network is not cost-justified".

munotes.in21

The Advantages of Wireless Sensor Networks

10. Flexibility

The same network can be given a new task (report every minute instead of every hour, or report only when a threshold is crossed) by sending it a new instruction. Nodes can be carried on vehicles or animals, and the sink itself can move. A wired system is fixed to where its cables run.

Worked example: two sums behind the advantages

The sensing distance

A forest department wants to detect chainsaws. It can mount one sensitive microphone on a watchtower 100 metres from a stand of trees, or scatter cheap microphone nodes so that one sits 10 metres from each tree.

By the inverse-square law, the sound intensity at the tower is less than at the scattered node by the square of the ratio of the distances:

(100 / 10)² = 100

The node receives 100 times the sound intensity the tower does. Expressed in decibels, a ratio of 100 in intensity is 20 decibels (a factor of 10 is 10 decibels, and there are two of them). A tower microphone would need to be 100 times more sensitive to do as well, and it would still hear only one direction. This is the advantage of being "very close to" the phenomenon, in one number.

The cabling

A cold-chain company has 200 storage rooms spread across a warehouse complex and wants temperature readings every 5 minutes. The prices below are assumptions made for this exercise, not market prices.

Wired option. Assume the average cable run from a room to the control panel is 60 metres, and cable plus installation costs Rs 250 a metre. The cabling alone costs

200 × 60 × 250 = 30,00,000

rupees, before a single sensor is bought, and every future room adds another cable run.

Wireless option. Assume a node costs Rs 3,000 and its batteries Rs 100 a year. The nodes cost 200 × 3,000 = 6,00,000 rupees once, and batteries 200 × 100 = 20,000 rupees a year.

Manual option. A person visiting every room is cheap to start, but at one visit a day the company gets one reading per room per day instead of 288 (which is 24 × 60 / 5), and a compressor failure at night is discovered in the morning.

On these assumptions the wireless network costs a fifth of the cabling alone, gives the full reading rate, and can be extended one room at a time. Change the assumptions and the numbers change, but the shape of the answer is the one this chapter describes: the wireless network wins where cabling is long, awkward or impossible, and where readings are needed often.

munotes.in22

The Advantages of Wireless Sensor Networks

Distinctions

Wireless sensor networkWired sensorsManual observationOne distant instrument
InstallationFix nodes, no cablesCable to every pointNoneOne site
ReadingsContinuous, many pointsContinuous, many pointsOccasional, many pointsContinuous, one point
Distance to the phenomenonVery closeCloseClose, when presentFar
Survives failuresYes, by redundancyNo, a point goes darkNot applicableNo
Works in inaccessible placesYes, even air-droppedRarelyNoSometimes
Main running costBatteriesMaintenance of cablingStaffMaintenance

What it does not mean

These advantages are not free. Every one of them is paid for in energy, and the next chapter, [The Challenges of Wireless Sensor Networks], is the bill. An answer that lists advantages and never mentions the battery has told half the story.

"Low cost" is per node, not per network. A thousand cheap nodes are not cheap. The claim is that the whole network costs less than the alternative for the same coverage.

A WSN is not always the right answer. A factory that already has power and data cabling at every machine gains little by going wireless and loses the reliability of a wire. The advantages are strongest where cabling is long, awkward or impossible.

Fault tolerance has a limit. Redundant nodes cover for each other; nothing covers for a failed sink unless there is a second one.

Quick revision

  • Advantages against cabling, manual observation and one distant instrument.
  • No cabling: quick, cheap, works in occupied buildings, easy to move.
  • Close to the phenomenon: stronger signal (inverse square: 10 m against 100 m is 100 times, 20 decibels) and finer spatial detail.
  • Inaccessible places: random deployment, dropped from a plane.
  • Fault tolerance by redundancy; self-organisation, no infrastructure in the field; scalability by redeployment.
  • In-network processing: send partially processed data, not raw readings.
  • Continuous, unattended, cheap to run; low cost per node; flexible and retaskable.
  • Silo example (assumed prices): cabling Rs 30,00,000 against nodes Rs 6,00,000 plus Rs 20,000 a year; 288 readings a day against one.

Test yourself

1. State five advantages of wireless sensor networks. Any five of: no cabling; sensing close to the phenomenon; deployment in inaccessible places; fault tolerance through redundancy; self-organisation with no infrastructure in the field; scalability by adding nodes; in-network processing; continuous unattended monitoring; low cost per node; flexibility and retasking.

2. Why does being close to the phenomenon improve sensing? Many signals weaken with distance, sound in open air with the square of the distance, so a nearby cheap sensor receives a stronger signal than a distant expensive one; and many nearby sensors give spatial detail that one distant instrument averages away.

munotes.in23

The Advantages of Wireless Sensor Networks

3. A sensor 10 metres from a sound source is compared with one at 100 metres. By what factor does the intensity differ? By (100 / 10)² = 100, which is 20 decibels, in favour of the nearer sensor.

4. How does a WSN achieve fault tolerance? By redundancy: it is deployed densely, so a failed node's area is still sensed by its neighbours, and in a mesh the data finds another path. Only the loss of the sink is not covered, unless there are several sinks.

5. When would a wired sensor system be the better choice? Where power and data cabling already reach every sensing point, as at the machines of a factory, so the wireless network's saving on cabling disappears and the reliability of a wire is worth more.

Contents This chapter on its own page

munotes.in24

Chapter Five

The Challenges of Wireless Sensor Networks

Syllabus topic Module 1, "Introduction and Overview of WSNs: Advantages and challenges of WSNs"

In one line

A sensor network must work for months on a small battery, with a tiny processor, over radio links that come and go, in places nobody visits, and at a scale where nothing can be configured by hand.

In the wording a student can write in an examination, the challenges of wireless sensor networks are: (1) limited and irreplaceable energy; (2) limited processing power and memory; (3) unreliable, time-varying wireless links; (4) unattended operation and the need to self-configure; (5) large scale and high density; (6) a dynamic topology caused by failures, sleeping and movement; (7) application-specific quality of service in place of throughput; (8) the need to be data-centric and to process data in the network; (9) security in exposed, capturable nodes on a broadcast medium; (10) time synchronisation and localisation; (11) programming, deployment and maintenance of thousands of nodes in the field; and (12) cost.

Why the challenges come first

A network designer does not choose a routing protocol or a MAC protocol because it is elegant. It is chosen because it survives these challenges better than the alternatives. So learn the challenges well: in the rest of the book, whenever a protocol does something strange (sleeps for 90 per cent of the time, refuses to acknowledge a packet, forgets which node sent a reading), the reason is on this list.

The challenges, one by one

1. Limited, irreplaceable energy

A node runs on a small battery that, in most deployments, nobody will ever replace. When the battery is flat the node is dead, and when enough nodes are dead the network is dead. Energy is the first constraint on every decision.

The surprise for a student is where the energy goes. Look at the two chips that many research nodes were built from: the Texas Instruments CC2420 radio and the MSP430F1611 microcontroller.

Part and stateCurrent drawnSource
Radio receiving (or just listening)18.8 mACC2420 datasheet, 6.10
Radio transmitting at 0 dBm17.4 mACC2420 datasheet, 6.10
Radio idle, oscillator running0.426 mACC2420 datasheet, 6.10
Radio powered down0.02 mACC2420 datasheet, 6.10
Processor active at 1 MHz, 2.2 V0.33 mAMSP430F1611 datasheet
Processor in standby0.0011 mAMSP430F1611 datasheet

Three lessons, and each shapes a later chapter.

  • The radio, not the processor, is the energy problem. Listening costs 18.8 / 0.33, about 57 times, what computing costs. So it pays to compute a lot in order to send a little.
  • Listening costs as much as sending. 18.8 mA to receive, 17.4 mA to transmit. A radio that is merely waiting for a packet that never comes, which is called idle listening, burns energy as fast as one that is talking. This is why a sensor node's radio must be switched off most of the time, and why [MAC Protocols for Sensor Networks: The Job and Where the Energy Goes] treats idle listening as the main enemy.
  • Sleeping is almost free. Powered down, the radio draws 20 microamperes, nearly a thousand times less than listening. The whole art of low-power networking is to be in that state as often as possible without missing anything.
munotes.in25

The Challenges of Wireless Sensor Networks

Worked: a node that never sleeps. Suppose the node runs on a battery pack rated 2,500 mAh (an assumption for the exercise; the chapter [How Long a Node Lasts: The Energy Budget Worked Out] uses real battery figures). A radio left listening draws 18.8 mA, so the battery lasts 2,500 / 18.8, about 133 hours, which is about five and a half days. A network designed to last a season is dead in a week. Keeping the radio off for most of the time is not an optimisation; it is the only way the network can exist.

2. Limited processing power and memory

The MSP430F1611 has 10 KB of RAM and 48 KB of flash for program code. A phone has millions of times more. A protocol that keeps a large table, a long queue or a big buffer simply does not fit, and an operating system designed for a desktop is out of the question. The first TinyOS, by comparison, fitted in 178 bytes of memory. Everything that runs on a node must be small and simple, which is why [Why a Sensor Node Needs an Operating System] exists.

3. Unreliable, time-varying wireless links

A radio link is not a wire. Its quality changes with the weather, with people and vehicles moving about, and with other radios: the 2.4 GHz band that many sensor nodes use is shared with Wi-Fi, Bluetooth and microwave ovens. Links can be asymmetric (A hears B but B does not hear A), and between the range where every packet arrives and the range where none does lies a wide band where some arrive and some do not, which the chapter [TOSSIM: Simulating Motes, Radio Gain and Packet Loss] explains. A protocol that assumes a link either works or does not will fail in the field.

4. Unattended operation and self-configuration

Nobody stands beside each node to give it an address, tell it who its neighbours are or set up its routes. After deployment the nodes must discover their neighbours, build routes, and repair them when something changes, all by themselves. A network of a thousand nodes cannot be configured by hand even if someone were there.

5. Large scale and high density

The survey expects deployments of hundreds or thousands of nodes, sometimes more, with densities as high as 20 nodes in a cubic metre. At that scale a node cannot know the whole network. Every algorithm must be local: a node decides from what its neighbours tell it, never from a global picture. And with so many nodes in one radio range, their transmissions collide unless the MAC protocol keeps them apart.

munotes.in26

The Challenges of Wireless Sensor Networks

6. A dynamic topology

The survey divides the topology's life into three phases: deployment, post-deployment (when the topology changes because nodes move, links are jammed or blocked, batteries run down, or nodes malfunction) and redeployment of additional nodes. On top of that, nodes that sleep to save energy are, from their neighbours' point of view, temporarily absent. So the set of working links changes all the time, and routes must adapt.

7. Application-specific quality of service

On the Internet, quality of service means throughput and delay. A sensor network's user wants something else: that every fire is detected, that the location of a vehicle is estimated within five metres, that the average temperature is right to half a degree. These are measures of the information, not of the bits, and they differ from one application to another. Designing for them is the subject of [Optimization Goals: Quality of Service, Energy Efficiency and Lifetime].

8. Being data-centric, and processing in the network

The user asks about the phenomenon, not about node number 145, and nodes may not even have globally unique identities. The network has to route, name and store data by what it is and where it came from, and combine it on the way. That is a different way of building a network from the one students learn first, and it is a challenge because the familiar tools (IP addresses, end-to-end connections) do not apply.

9. Security

Nodes lie in the open, where an attacker can pick one up, read its memory and reprogram it. The radio is a broadcast medium that anyone can hear and anyone can transmit on. And the node cannot afford the cryptography a server uses. Keeping data secret and authentic, and keeping routing honest, is hard, which is why the book gives it four chapters starting from [Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks].

10. Time synchronisation and localisation

A reading is useless without when and where. Each node has its own clock, which drifts, and giving every node a satellite receiver to learn its position costs too much money and energy. So nodes must agree on the time and work out their positions among themselves, which is [Time Synchronisation and Localisation].

munotes.in27

The Challenges of Wireless Sensor Networks

11. Programming, deployment and maintenance

A bug found after deployment must be fixed on a thousand nodes scattered across a forest, without collecting them. The network has to be reprogrammable over its own radio, tested before it goes out, and monitored while it runs. Tools for programming a whole network, rather than one node, are part of the answer, and [WSN Middleware: Why It Is Needed, and Its Architecture] describes them.

12. Cost

Because a network needs many nodes, each must be cheap. The survey, writing in 2002, went as far as to say a node should cost "much less than US$1" for a sensor network to be feasible, a target far below what real nodes cost then or now. Cheap hardware means small batteries, simple radios and little memory, which feeds straight back into challenges 1 to 3.

The same challenges, as MU's second text book organises them

Karl and Willig, the second text book on MU's list, split the challenges into two lists, and an answer that uses their headings is well organised.

The characteristic requirements, the properties most applications demand:

  1. Type of service. A WSN is not there to move bits but to give meaningful information about a task; they quote a Berkeley engineer's line, "People want answers, not numbers". Interactions are scoped to regions and time intervals.
  2. Quality of service. Bounded delay and minimum bandwidth often do not matter; what matters is the amount and quality of information extracted, such as reliable detection of events or the accuracy of a temperature map. They add that the packet delivery ratio "is an insufficient metric".
  3. Fault tolerance, achieved by redundant deployment: more nodes than would be needed if every node worked.
  4. Lifetime, a "very important figure of merit", whose precise definition depends on the application ([Optimization Goals: Quality of Service, Energy Efficiency and Lifetime] gives the definitions).
  5. Scalability to large numbers of nodes.
  6. A wide range of densities, varying between applications and within one network over time and space.
  7. Programmability: nodes must be reprogrammable during operation as tasks change.
  8. Maintainability: the network must monitor its own health and adapt, for example by giving lower quality when energy runs short.

The required mechanisms, the techniques that meet those requirements:

  1. Multi-hop wireless communication, because direct communication over long distances needs prohibitively high power.
  2. Energy-efficient operation, including avoiding "hotspots" where energy consumption concentrates.
  3. Auto-configuration, including nodes working out their own positions ("self-location"), tolerating failed nodes and integrating new ones.
  4. Collaboration and in-network processing, such as combining readings as they travel to find the highest or average temperature.
  5. Data-centric operation, asking for values rather than for particular nodes.
  6. Locality: each node keeps state only about its direct neighbours, so the network can scale.
  7. Exploiting trade-offs, such as energy against accuracy, or the lifetime of the whole network against the lifetime of individual nodes.
munotes.in28

The Challenges of Wireless Sensor Networks

The twelve challenges above and these two lists describe the same ground. The twelve are organised by problem; Karl and Willig organise by what the application needs and what the network must do about it.

How the challenges pull against each other

The challenges are not independent, and the hard part of design is that solving one makes another worse.

  • Energy against delay. Sleeping saves energy, but a packet that arrives for a sleeping node must wait. Every duty-cycling MAC protocol trades one for the other, and [S-MAC: Latency, Adaptive Listening and the Energy Saved] measures the trade.
  • Energy against reliability. Acknowledging and retransmitting packets makes delivery reliable, and costs energy each time.
  • Cost against lifetime. A bigger battery lasts longer and costs more.
  • Density against collisions. More nodes give better coverage and more redundancy, and more radios contend for the same channel.
  • Security against energy. Every byte of authentication code added to a packet is transmitted, and transmission is the most expensive thing a node does.

The challenges, as numbers

Every challenge listed above is a number underneath, and the numbers are worth seeing together, because they are not independent: each one is bought with another.

# Each challenge, given the one number that makes it a challenge.
import math

V, CAPACITY_MAH = 3.0, 2500.0          # two AA alkaline cells, a round figure
JOULES = CAPACITY_MAH / 1000 * 3600 * V
I_RX, I_TX, I_SLEEP = 18.8e-3, 17.4e-3, 20e-6
RATE = 250e3
DAY = 24 * 3600.0

print("Every challenge in this chapter is a number. Here they are.")
print()
print("ENERGY. Two AA cells hold about %.0f joules." % JOULES)
for name, cur in (("listening, radio always on", I_RX),
                  ("listening one second in ten", I_RX * 0.1 + I_SLEEP * 0.9),
                  ("listening one second in a hundred", I_RX * 0.01 + I_SLEEP * 0.99),
                  ("listening one second in a thousand", I_RX * 0.001 + I_SLEEP * 0.999)):
    life = JOULES / (cur * V)
    print("  %-36s %8.1f days" % (name, life / DAY))
print("  a node that is simply left switched on lasts under a week; the same node at a")
print("  thousandth of a duty cycle runs for over seven years, which is about as long as")
print("  the cell will sit on a shelf anyway. The challenge")
print("  is not the battery. It is the radio being on.")

print()
print("BANDWIDTH. One channel of %.0f kbit/s, shared." % (RATE / 1000))
for n in (10, 100, 1000):
    share = RATE / n
    print("  %4d nodes sharing it: %8.0f bit/s each before any collision at all"
          % (n, share))
print("  and that is the ideal. Contention takes a large part of it back, which is why")
print("  the MAC chapters spend so long on collisions.")

print()
print("SCALE. A flood reaches everyone by making everyone shout.")
for n in (10, 100, 1000):
    print("  %4d nodes: one flood is %4d transmissions, and a query answered by all of"
          % (n, n))
    print("       them is another %4d coming back" % n)
print("  the traffic grows with the number of nodes, and so does the chance two of them")
print("  speak at once. Nothing about a protocol that works for ten is evidence it works")
print("  for a thousand.")

print()
print("LATENCY. Sleeping costs time as surely as listening costs energy.")
for duty, frame_ms in ((0.1, 100), (0.01, 1000), (0.001, 10000)):
    wait = frame_ms / 2.0
    for h in (5, 10):
        print("  at %5.1f%% duty with a %4d ms cycle, a %2d hop path waits about %6.0f ms"
              % (duty * 100, frame_ms, h, h * wait))
print("  a network that sleeps to survive is a network that answers slowly, and the two")
print("  cannot both be optimised. That is the trade every MAC in this book is making.")

print()
print("COST AND UNATTENDED OPERATION.")
for per_node in (300, 1000, 3000):
    for n in (100, 1000):
        print("  %4d nodes at Rs %4d each: Rs %9d of hardware" % (n, per_node, n * per_node))
print("  and one visit by one person to replace one battery, at a day's wage and a day's")
print("  travel, can cost more than the node. That is why the design target is a node")
print("  that is never visited, and why every other challenge reduces to energy.")
munotes.in29

The Challenges of Wireless Sensor Networks

Every challenge in this chapter is a number. Here they are.

ENERGY. Two AA cells hold about 27000 joules.
  listening, radio always on                5.5 days
  listening one second in ten              54.9 days
  listening one second in a hundred       501.3 days
  listening one second in a thousand     2686.1 days
  a node that is simply left switched on lasts under a week; the same node at a
  thousandth of a duty cycle runs for over seven years, which is about as long as
  the cell will sit on a shelf anyway. The challenge
  is not the battery. It is the radio being on.

BANDWIDTH. One channel of 250 kbit/s, shared.
    10 nodes sharing it:    25000 bit/s each before any collision at all
   100 nodes sharing it:     2500 bit/s each before any collision at all
  1000 nodes sharing it:      250 bit/s each before any collision at all
  and that is the ideal. Contention takes a large part of it back, which is why
  the MAC chapters spend so long on collisions.

SCALE. A flood reaches everyone by making everyone shout.
    10 nodes: one flood is   10 transmissions, and a query answered by all of
       them is another   10 coming back
   100 nodes: one flood is  100 transmissions, and a query answered by all of
       them is another  100 coming back
  1000 nodes: one flood is 1000 transmissions, and a query answered by all of
       them is another 1000 coming back
  the traffic grows with the number of nodes, and so does the chance two of them
  speak at once. Nothing about a protocol that works for ten is evidence it works
  for a thousand.

LATENCY. Sleeping costs time as surely as listening costs energy.
  at  10.0% duty with a  100 ms cycle, a  5 hop path waits about    250 ms
  at  10.0% duty with a  100 ms cycle, a 10 hop path waits about    500 ms
  at   1.0% duty with a 1000 ms cycle, a  5 hop path waits about   2500 ms
  at   1.0% duty with a 1000 ms cycle, a 10 hop path waits about   5000 ms
  at   0.1% duty with a 10000 ms cycle, a  5 hop path waits about  25000 ms
  at   0.1% duty with a 10000 ms cycle, a 10 hop path waits about  50000 ms
  a network that sleeps to survive is a network that answers slowly, and the two
  cannot both be optimised. That is the trade every MAC in this book is making.

COST AND UNATTENDED OPERATION.
   100 nodes at Rs  300 each: Rs     30000 of hardware
  1000 nodes at Rs  300 each: Rs    300000 of hardware
   100 nodes at Rs 1000 each: Rs    100000 of hardware
  1000 nodes at Rs 1000 each: Rs   1000000 of hardware
   100 nodes at Rs 3000 each: Rs    300000 of hardware
  1000 nodes at Rs 3000 each: Rs   3000000 of hardware
  and one visit by one person to replace one battery, at a day's wage and a day's
  travel, can cost more than the node. That is why the design target is a node
  that is never visited, and why every other challenge reduces to energy.
munotes.in30

The Challenges of Wireless Sensor Networks

Distinctions: each challenge, and where the book answers it

ChallengeThe mechanism that answers itWhere
EnergyDuty cycling, short hops, in-network processingMAC chapters; single hop against multiple hops
Processing and memoryA tiny, event-driven operating systemTinyOS and the operating system chapters
Unreliable linksLink estimation, retransmission, alternative pathsTOSSIM chapter; routing tables chapter
Unattended operationSelf-organisation and auto-configurationAd hoc network chapters
Scale and densityLocal algorithms, clusteringRouting strategies; LEACH
Dynamic topologyAdaptive, on-demand routingAODV, DSR and the routing tables chapter
Application QoSNew figures of meritOptimization goals chapter
Data-centric operationAttribute-based naming, aggregationDesign principles chapters; directed diffusion
SecurityLink-layer security, key predistributionSecurity chapters
Time and positionSynchronisation and localisation protocolsTime synchronisation and localisation
Programming and maintenanceMiddleware, network reprogrammingMiddleware chapters
CostSimple hardware, which feeds the first threeNode technology chapters
munotes.in31

The Challenges of Wireless Sensor Networks

What it does not mean

Energy being limited does not mean the answer is a bigger battery. A bigger battery makes the node bigger and dearer and only postpones the problem by a constant factor. The real answer is to change what the node does, above all to keep its radio off.

The processor is not the energy problem. It is natural to assume that computing is what uses the power. On these chips the radio listening uses about 57 times as much as the processor running.

An unreliable link is not a broken radio. Links in the band between perfect and dead are normal in every real deployment, and protocols must be designed for them, not tested only where every packet arrives.

Challenges are not the same as design factors, but they overlap. The survey's design factors (fault tolerance, scalability, cost, environment, topology, hardware, medium and power) are the properties a designer must take into account; the challenges are the problems those factors create. Both are asked, and [The Operating Environment and the Design Factors] teaches the factors with their formulas.

Quick revision

  • Twelve challenges: energy; processing and memory; unreliable links; unattended self-configuration; scale and density; dynamic topology; application QoS; data-centric operation; security; time and position; programming and maintenance; cost.
  • CC2420: receive 18.8 mA, transmit 17.4 mA, idle 0.426 mA, power down 0.02 mA. MSP430F1611: active 0.33 mA at 1 MHz, standby 0.0011 mA; 10 KB RAM, 48 KB flash.
  • Listening costs as much as sending; the radio costs about 57 times the processor; sleeping is almost free.
  • A 2,500 mAh battery (assumed) with the radio always listening: 2,500 / 18.8, about 133 hours, five and a half days.
  • Densities up to 20 nodes per cubic metre; hundreds to thousands of nodes; the first TinyOS in 178 bytes.
  • Karl and Willig: requirements type of service, QoS, fault tolerance, lifetime, scalability, range of densities, programmability, maintainability; mechanisms multi-hop, energy efficiency, auto-configuration, collaboration and in-network processing, data-centric, locality, exploiting trade-offs.
  • The challenges pull against each other: energy against delay, reliability, security; density against collisions; cost against lifetime.

Test yourself

1. List eight challenges of wireless sensor networks. Any eight of: limited, irreplaceable energy; limited processing and memory; unreliable, time-varying links; unattended operation and self-configuration; large scale and high density; dynamic topology; application-specific QoS; data-centric operation; security; time synchronisation and localisation; programming and maintenance; cost.

2. Using the CC2420 and MSP430F1611 figures, explain why the radio is the main energy problem. The radio draws 18.8 mA receiving and 17.4 mA transmitting, while the processor draws 0.33 mA active at 1 MHz. Listening therefore costs about 57 times as much as computing, so a node should compute to reduce what it sends and should keep its radio off whenever it can.

munotes.in32

The Challenges of Wireless Sensor Networks

3. What is idle listening, and why does it matter? Keeping the receiver on while waiting for a packet that may never come. On the CC2420 it draws the full 18.8 mA of receiving, as much as transmitting, so a node that idles with its radio on wastes energy as fast as one that is talking.

4. A node's radio listens continuously on a 2,500 mAh battery. Roughly how long does it last? 2,500 / 18.8, about 133 hours, or about five and a half days.

5. Why must sensor network algorithms be local? Because with hundreds or thousands of nodes, densities up to 20 a cubic metre and a constantly changing topology, no node can know the whole network. Each node must decide from what its neighbours tell it.

6. Give two pairs of challenges that pull against each other. Energy against delay (sleeping saves energy but delays packets for sleeping nodes), and density against collisions (more nodes improve coverage but contend for the channel). Others: energy against reliability, security against energy, cost against lifetime.

Contents This chapter on its own page

munotes.in33

Chapter Six

Applications of Wireless Sensor Networks

Syllabus topic Module 1, "Introduction and Overview of WSNs: Applications of WSNs"

In one line

Sensor networks are used wherever something physical has to be watched closely, over a wide area or for a long time, without people or cables: in the environment, on farms, in hospitals, factories, buildings and cities, on battlefields and after disasters.

In the wording a student can write in an examination: applications of WSNs include environmental monitoring (habitat, volcano, forest fire, flood and pollution monitoring), precision agriculture, healthcare (patient and elderly monitoring), industrial monitoring (machine condition, process control, predictive maintenance), buildings and homes (energy, safety, comfort), structural health monitoring of bridges and buildings, military uses (surveillance, target detection and tracking), disaster management, logistics and inventory, smart cities and smart grids. Every application is built from four kinds of sensing task: event detection, periodic measurement, function approximation and edge detection, and tracking.

Why the applications decide the network

A sensor network is built for one job, which is one of the six differences from the Internet in [What a Wireless Sensor Network Is]. So the application decides almost everything: how many nodes, how often they sample, how fast a report must arrive, how long the batteries must last and how much loss is tolerable. Two networks that look identical in a photograph can need completely different protocols because one watches seabirds and the other watches a volcano. Learn the applications with their demands, not as a list of names.

The applications, sector by sector

Environmental monitoring

This is where sensor networks began. The first famous deployment, on Great Duck Island off Maine in July 2002, watched a seabird colony: 32 motes, nine of them inside nesting burrows, reporting temperature, humidity, pressure, light and infrared so that biologists could study the birds without disturbing them. The data was periodic and slow, and the demand was lifetime: nine months on batteries.

The opposite extreme came in August 2005 at Reventador, an active volcano in Ecuador. Computer scientists from Harvard, working with seismologists from the Universities of New Hampshire and North Carolina, deployed 16 nodes for 19 days, each sampling seismic and acoustic sensors 100 times a second. The nodes sat 200 to 400 metres apart and sent their data by multi-hop radio to a gateway, which relayed it to an observatory 4.6 km away. Because the raw data was far too much to send continuously, each node ran an event-detection algorithm and the network collected high-resolution data only around interesting activity; it recorded 229 earthquakes, eruptions and other events. Here the demand was data fidelity: every sample of an event, accurately timed.

Other environmental uses follow the same two patterns: forest fire detection (events), flood and river-level warning (events with periodic background readings), and air and water quality monitoring (periodic).

munotes.in34

Applications of Wireless Sensor Networks

Agriculture

Soil moisture, temperature and humidity nodes let a farmer irrigate or spray only where and when it is needed, which is called precision agriculture. The vineyard in [What a Wireless Sensor Network Is] is this application: periodic readings, low data volume, a whole season on batteries. Cold stores and greenhouses use the same idea. MU's own description of this paper names smart agriculture first among its applications.

Healthcare

Small sensors worn on the body measure heart rate, blood oxygen, temperature or movement and report to a phone or a bedside unit; this is a body area network. The survey's introduction already names it: nodes "deployed to monitor patients and assist disabled patients". A hospital ward can monitor many patients without wires tying them to machines, and an elderly person living alone can be watched for a fall. The demands are unusual for a sensor network: reports must be reliable and quick, and the data is private.

Industry

Vibration and temperature nodes on motors, pumps and compressors detect wear before a machine fails, which is predictive maintenance. Nodes on pipelines, tanks and valves report pressures and levels for process control. The attraction is the one [The Advantages of Wireless Sensor Networks] priced: a plant already built cannot easily be rewired. MU's description calls this "industrial IoT". The demands are reliability and, for control, bounded delay.

Buildings and homes

Temperature, occupancy and light nodes let a building run its air conditioning and lighting only where people are. Smoke, gas and water-leak nodes raise alarms; the office building in [The Architectural Elements of a Sensor Network] is this application. In homes the same nodes drive "smart home" automation.

Structural health monitoring

Accelerometers and strain sensors on bridges, towers, dams and tall buildings measure how the structure vibrates and whether that is changing over the years, and they record the response to an earthquake or a heavy load. The data rates are high during an event and low otherwise, like the volcano.

Military

The survey lists the military uses first: rapid deployment, self-organisation and fault tolerance make sensor networks suited to "command, control, communications, computing, intelligence, surveillance, reconnaissance, and targeting systems". In practice that means detecting and tracking vehicles or people crossing a boundary, and monitoring an area for chemical or biological agents. Nodes may be dropped from the air, and must resist capture and jamming.

Disaster management

After an earthquake, flood or industrial accident, nodes scattered in the affected area can report temperatures, toxic gases or the presence of survivors where rescuers cannot yet go. The survey names "monitoring disaster areas" among its commercial applications. The demand is rapid deployment with no infrastructure at all, because whatever infrastructure existed may be destroyed.

munotes.in35

Applications of Wireless Sensor Networks

Logistics and inventory

Nodes on containers, pallets or shelves report location, temperature and shock, so that a consignment of vaccines, for example, can be proven to have stayed cold all the way. The survey names "managing inventory" and "monitoring product quality".

Smart cities and smart grids

Street-level nodes report parking spaces, air quality, noise and street-light faults. On the electricity network, sensor nodes and smart meters report consumption and faults so that supply can follow demand. MU's description names both.

The four kinds of sensing task

Karl and Willig, the second text book on MU's list, sort all these applications into four patterns of interaction between the network and the phenomenon. Every application in the sectors above is one of them or a mixture.

Event detection, periodic measurement, function approximation and edge detection, and tracking

Figure 6.1 The four kinds of sensing task every application is built from

1. Event detection. The network reports only when something happens: a fire, an intruder, a gas leak, a machine vibrating abnormally. Most of the time nothing is sent. When an event happens, the report must arrive quickly and reliably, and the network may need to classify the event (a car or a lorry? a small tremor or an eruption?), often by nodes combining their readings.

2. Periodic measurement. Every node reports its reading at regular intervals: soil moisture every 15 minutes, a burrow's temperature every 5 minutes. The traffic is steady and predictable, individual readings matter little, and the network can plan its sleeping around the schedule. The demand is lifetime.

3. Function approximation and edge detection. The user wants a picture of how a quantity varies across space, the temperature map of a forest or the moisture across a field. The nodes sample the underlying function at points, and the sink reconstructs it. A special case is finding the edge of a region where the quantity crosses a threshold, such as the front of a fire or the boundary of a contaminated area. Only nodes near the edge need to report.

4. Tracking. The network follows a moving target: a vehicle, an animal, a person. Nodes near the target detect it, estimate its position between them, and hand the task on to the nodes ahead of it as it moves. The demands are timeliness and accuracy of position.

Worked example: classifying ten applications

The table classifies ten applications by their sensing task and by what each demands of the network, which is the analysis an examiner is looking for.

ApplicationSensing taskData rateDelay allowedMain demand
Seabird burrows, Great Duck IslandPeriodicVery lowHoursLifetime of months
Volcano, ReventadorEvent detection, then bulk dataVery high during eventsMinutesFidelity and exact timing
Vineyard irrigationPeriodicVery lowMinutesLifetime of a season
Forest fireEvent detection, then edge detectionLow, a burst in a fireSecondsReliable, fast alarms
Patient monitoringPeriodic plus event alarmsModerateSecondsReliability and privacy
Motor vibrationPeriodic plus event alarmsHigh while samplingMinutesAccuracy
Office fire alarmEvent detectionVery lowSecondsNo missed alarms
Contamination boundaryEdge detectionLowMinutesAccuracy of the boundary
Vehicle at a borderTrackingModerate near the targetSecondsPosition accuracy
Cold-chain consignmentPeriodic plus event alarmsVery lowHoursProof of the record
munotes.in36

Applications of Wireless Sensor Networks

Two things follow from the table, and both are worth writing in an answer.

The same hardware can serve very different applications, but not with the same protocols. The seabird network used Mica motes and the volcano network TMote Sky nodes, both small battery-powered motes of the same general kind. One slept almost all the time and trickled readings out; the other sampled a hundred times a second and moved bursts of data over multiple hops.

Most real applications mix tasks. A fire network does event detection to raise the alarm and then edge detection to map the fire front. A patient monitor sends periodic readings and immediate alarms. The protocol has to carry both kinds of traffic, which is why [Optimization Goals: Quality of Service, Energy Efficiency and Lifetime] measures quality by the task, not by throughput.

Distinctions

Event detectionPeriodic measurementFunction approximationTracking
Who reportsNodes that detect the eventEvery nodeEnough nodes to reconstruct the field, or those at the edgeNodes near the target
WhenWhen it happensOn a scheduleOn request or on a scheduleContinuously while the target is near
TrafficRare burstsSteadyModerateMoves with the target
Critical measureDetection probability and delayLifetimeAccuracy of the reconstructionPosition accuracy and timeliness

What it does not mean

A list of sectors is not an application answer. Military, health, environment earns little. Say what the network senses, which task it performs, and what that demands.

Environmental monitoring is not one kind of application. Seabird burrows and an erupting volcano are both environmental and are opposite in every demand that matters.

Healthcare body sensors are not quite a sensor network in the classic sense. They are few, close together and usually one hop from a phone. They share the energy and radio problems but not the multi-hop scale.

Tracking is not the same as event detection. Detecting that a vehicle has entered is an event; following where it goes is tracking, and it needs the network to hand the target from node to node.

Quick revision

  • Sectors: environment, agriculture, healthcare, industry, buildings and homes, structural health, military, disaster management, logistics, smart cities and grids.
  • Two real deployments: Great Duck Island, July 2002 (32 motes, periodic, lifetime) and Reventador volcano, August 2005 (16 nodes, 19 days, 100 samples a second, 229 events, event-triggered bulk data).
  • Four tasks (Karl and Willig): event detection, periodic measurement, function approximation and edge detection, tracking.
  • The application decides the data rate, the delay allowed and the main demand; most real applications mix tasks.
munotes.in37

Applications of Wireless Sensor Networks

Test yourself

1. Explain five applications of wireless sensor networks. Any five, each with what is sensed and what is demanded: habitat monitoring (periodic, lifetime); volcano monitoring (event-triggered high-rate data, fidelity); precision agriculture (periodic soil readings, a season on batteries); patient monitoring (periodic plus alarms, reliability and privacy); industrial predictive maintenance (vibration, accuracy); building safety (event alarms, no misses); military surveillance and tracking; disaster management; logistics; smart cities and grids.

2. Name and explain the four kinds of sensing task. Event detection: report only when something happens, quickly and reliably. Periodic measurement: every node reports on a schedule; lifetime matters most. Function approximation and edge detection: reconstruct how a quantity varies over space, or find the boundary where it crosses a threshold. Tracking: follow a moving target, handing it from node to node.

3. Compare the Great Duck Island and Reventador deployments. Both were environmental. Great Duck Island (2002, 32 motes) sent slow periodic readings and was designed for months of lifetime. Reventador (2005, 16 nodes, 19 days) sampled at 100 Hz, used event detection to decide when to collect, and cared most about capturing every sample of an event accurately.

4. Which sensing task does a forest fire network perform? Event detection to raise the alarm, then edge detection to map where the fire front is.

5. Why do sensor network protocols differ from one application to another? Because the application decides the data rate, the delay that is acceptable, the loss that is tolerable and the lifetime required, and a protocol that suits one set of demands, such as sleeping almost always, fails another, such as moving bursts of high-rate data.

Contents This chapter on its own page

munotes.in38

Chapter Seven

Inside a Sensor Node: The Five Units

Syllabus topic Module 1, "Introduction and Overview of WSNs: Sensor node technology"

In one line

A sensor node is five units on one small board: a sensing unit that turns the physical world into numbers, a microcontroller that decides what to do with them, a radio that talks to the neighbours, a memory that holds the program and the data, and a power unit that keeps the other four alive.

In the wording a student can write in an examination: the hardware of a sensor node consists of (1) a sensing unit, made of one or more sensors and an analogue-to-digital converter (ADC); (2) a processing unit, usually a microcontroller, which runs the node's software and controls the other units; (3) a transceiver, the radio that sends and receives packets; (4) memory, the microcontroller's RAM and flash and often an external flash chip for logged data; and (5) a power unit, a battery with a voltage regulator, sometimes helped by an energy-scavenging power generator. Application-dependent units may be added: a location finding system and a mobilizer. Akyildiz and colleagues count four basic units, because they treat memory as part of the processing unit.

Why every unit is chosen for energy first

The node's job is to last. In [The Challenges of Wireless Sensor Networks] the radio alone, left listening, emptied a battery in under a week. So the question asked of every part on the board is not how fast is it? but how little can it draw, how deeply can it sleep, and how quickly can it wake? The same questions run through each unit below, and the practical's phrase, "energy-constrained environments", is the reason.

The five units of a sensor node, with the three optional units dashed

Figure 7.1 The block diagram of a sensor node

Unit 1: the sensing unit

What it is. A sensor (also called a transducer) converts a physical quantity into an electrical signal: a thermistor's resistance changes with temperature, a photodiode's current with light, a microphone's voltage with sound. Most sensors produce an analogue signal, a voltage that varies continuously. The ADC samples that voltage and turns it into a number the microcontroller can use. Between the two there is usually some signal conditioning, an amplifier or a filter that brings the signal into the ADC's range.

Two properties of the ADC that matter.

  • Resolution, the number of bits in each result. The MSP430F1611's ADC has 12 bits, so it divides its input range into 2 to the power 12, which is 4,096, steps. Over a range of 0 to 2.5 volts each step is 2.5 / 4,096 of a volt, about 0.61 millivolts.
  • Sampling rate, how many readings a second. A burrow's temperature needs one a minute; the volcano network of the previous chapter needed 100 a second.
munotes.in39

Inside a Sensor Node: The Five Units

Its energy. Most simple sensors draw little, but some draw a lot: a gas sensor has to heat its element, and an active sensor such as an ultrasonic range finder transmits a pulse. The rule the software follows is the same as for the radio: power the sensor only when a reading is due.

Unit 2: the processing unit

What it is. A microcontroller (MCU) is a small processor with its memory and its input and output devices on the same chip: timers, an ADC, serial interfaces (UART, SPI and I2C, the three ways chips on a board talk to each other). It runs the node's operating system and application, decides when to sense, what to send and when to sleep, and runs the network protocols.

Why a microcontroller and not something else. A desktop processor is far too hungry. A digital signal processor is faster at arithmetic on signals but draws more. An FPGA (a reconfigurable chip) or an ASIC (a chip designed for one job) can be very efficient but is costly to design and hard to change. The microcontroller wins for most nodes because it is cheap, reprogrammable and above all sleeps deeply and wakes fast.

A real one. The Telos mote, designed at Berkeley and described in 2005, chose the Texas Instruments MSP430 after comparing parts from several makers, because it had the lowest power in both sleep and active modes. The MSP430F1611 on it is a 16-bit processor with a 125-nanosecond instruction cycle, 10 kB of RAM and 48 kB of flash. It moves from standby, at about 1 microampere, to full speed in no more than 6 microseconds.

Why wake-up time matters. A node spends almost all its time asleep and wakes for a few milliseconds at a time. If waking takes as long as the work, half the energy goes on waking. The Telos authors put the whole design principle in one sentence: the node "is asleep for the majority of the time, wakes up quickly on an event, processes, and returns to sleep".

Why the lowest working voltage matters too. Two AA cells in series start at about 3 volts and are exhausted at about 1.8 volts. The MSP430 runs down to 1.8 volts, so it uses the whole battery. The ATmega128 used on the earlier Mica2 motes stops at 2.7 volts, which the Telos paper says leaves "almost 50% of the AA batteries unused".

Unit 3: the transceiver

What it is. The radio: it modulates bits onto a carrier to send them and demodulates received signals back into bits (both explained in [Modulation: ASK, FSK and PSK]). It has an antenna, often printed on the circuit board itself, as on Telos.

munotes.in40

Inside a Sensor Node: The Five Units

A real one. The Texas Instruments (originally Chipcon) CC2420 is an IEEE 802.15.4 radio for the 2.4 GHz band, sending 250 kbps and receiving signals as weak as -95 dBm (the unit is explained in [Radio Technology in WSNs: The Sensor Radio and Its Link Budget]). It has 128-byte buffers for one received and one outgoing frame, so the microcontroller can hand over a whole packet and go back to sleep while the radio sends it.

Its energy. The transceiver is the most expensive unit on the board: on the CC2420 it draws 18.8 mA receiving and 17.4 mA transmitting, against a third of a milliampere for the processor. Its states and the cost of moving between them are the subject of the next chapter.

Unit 4: memory

What it is. Three kinds, each for a different job.

  • RAM, inside the microcontroller: fast, lost when power is lost, and tiny (10 kB on the MSP430F1611). It holds the running program's variables, the packet being built, the routing table.
  • Program flash, inside the microcontroller: keeps its contents without power, holds the program (48 kB on the MSP430F1611).
  • External flash, a separate chip for logged readings, and on some motes for a new program image received over the air. Telos carries an ST M25P80, which holds 1,024 kB. The volcano network in [Applications of Wireless Sensor Networks] stored its samples in flash in 256-byte blocks.

Its energy. Reading flash is cheap; writing and erasing it costs far more, and flash wears out after many erase cycles. So a node logs data in large blocks rather than byte by byte.

Unit 5: the power unit

What it is. The battery (almost always), a voltage regulator or converter that gives the other units the steady voltage they need, and sometimes a power generator that scavenges energy from the surroundings, such as a small solar cell. The survey calls the power unit "one of the most important components of a sensor node".

Its role. Everything else is designed around it. A node's lifetime is its battery's energy divided by its average power, so the power unit sets the lifetime and the other four units' job is to keep the average power low. Batteries, their real capacity and energy scavenging are in the next chapter, and the lifetime arithmetic in [How Long a Node Lasts: The Energy Budget Worked Out].

The optional units

The survey names three units that some applications add.

  • Location finding system, for example a GPS receiver, because many routing schemes and most sensing tasks need to know where a reading came from. It is expensive in money and energy, so most nodes estimate their position from their neighbours instead, as [Time Synchronisation and Localisation] explains.
  • Power generator, a solar cell or other scavenger.
  • Mobilizer, which moves the node, for example a robot base. Rare, and used when the network must reposition itself to carry out its task.
munotes.in41

Inside a Sensor Node: The Five Units

Worked example: the block diagram of the vineyard node

Meera's vineyard node, from [What a Wireless Sensor Network Is], measures soil moisture and temperature every 15 minutes. Here is its block diagram as a table of the five units, and what each is doing in each part of its 15-minute cycle.

UnitPart chosenDuring the 60 ms it is awakeDuring the rest of the 15 minutes
Sensing unitMoisture probe (analogue) and temperature sensor, 12-bit ADCPowered, one reading eachSwitched off
MicrocontrollerMSP430-class, 10 kB RAMReads, builds the packet, runs the protocolStandby, timer running
Transceiver802.15.4 radio at 250 kbpsSends the packet, waits for the acknowledgement, forwards a neighbour's packetPowered down
MemoryRAM for the packet; flash logs every readingOne write to flashRetains contents
Power unitTwo AA cells, regulatorSupplies peak currentSupplies microamperes

The analysis the practical asks for, unit by unit.

  • Sensing unit: its energy is small because it is on for milliseconds; a heated gas sensor would change that completely.
  • Microcontroller: its active current matters less than its standby current and its wake-up time, because it is in standby for more than 99 per cent of the time.
  • Transceiver: the largest current by far; everything the protocols do is aimed at shortening the time it is on.
  • Memory: the reason the node can log data it could not send at once, for example while a neighbour is down.
  • Power unit: sets the lifetime; it must work down to the lowest voltage the other units accept, so no energy is stranded in the battery.

The fraction of time awake. Awake for 60 milliseconds in every 15 minutes, which is 900,000 milliseconds, the node is awake for 60 / 900,000 of the time, about 0.0067 per cent. That fraction is called the duty cycle, and it is what makes a season on two AA cells possible.

A short history of real nodes

The Telos paper traces the Berkeley family, and the direction of travel is the lesson.

MoteYearMicrocontrollerRAMRadio, data rateModulation
WeC1998AT90LS85350.5 kBTR1000, 10 kbpsOOK
Mica2001Atmel, 128 kB flash4 kBTR1000, 40 kbpsASK
Mica22002ATmega1284 kBCC1000, 38.4 kbpsFSK
Telos2004TI MSP43010 kBCC2420, 250 kbpsO-QPSK

Memory grew, radios got faster and moved to the standard 802.15.4, and the sleep current and wake-up time fell. The Mica2 took up to 4 milliseconds to wake; Telos takes 6 microseconds. MicaZ, the Mica family's successor, kept the Mica2 design but replaced its radio with the CC2420. Many later designs put the microcontroller and the radio on one piece of silicon, but the five units are still there, only closer together.

munotes.in42

Inside a Sensor Node: The Five Units

Distinctions

MicrocontrollerTransceiver
DoesComputes, controls, runs the protocolsSends and receives bits over the air
Current on the Telos partsAbout 0.33 mA active17.4 to 18.8 mA
What matters mostStandby current and wake-up timeTime spent on, and switching cost
RAMProgram flashExternal flash
Keeps data without powerNoYesYes
HoldsVariables, packets, tablesThe programLogged readings, new program images
Size on Telos10 kB48 kB1,024 kB

What it does not mean

Four units or five is not a contradiction. The survey's four put memory inside the processing unit; the practical's five list it separately. Say which you are using and name every part either way.

The sensor is not the whole sensing unit. The ADC, and usually a conditioning circuit, belong to it too.

The fastest processor is not the best one. A node's processor is judged by how deeply it sleeps and how quickly it wakes, because it is asleep nearly all the time.

A location finding system is not standard equipment. It is optional, and most networks avoid GPS on every node because of its cost and energy.

Quick revision

  • Five units: sensing unit (sensor, conditioning, ADC), microcontroller, transceiver, memory, power unit; optional location finding system, power generator, mobilizer. The survey counts four (memory inside processing).
  • 12-bit ADC: 4,096 steps; over 2.5 V, about 0.61 mV a step.
  • MSP430F1611: 16-bit, 10 kB RAM, 48 kB flash, standby about 1 microampere, wake-up within 6 microseconds, runs down to 1.8 V.
  • CC2420: 802.15.4, 2.4 GHz, 250 kbps, -95 dBm, 128-byte buffers; 18.8 mA receive, 17.4 mA transmit.
  • Design principle (Telos): asleep most of the time, wake quickly, process, sleep again.
  • The vineyard node's duty cycle: 60 ms in 15 minutes, about 0.0067 per cent.

Test yourself

1. Draw the block diagram of a sensor node and explain each unit. The figure: the sensing unit (sensor and ADC) converts the physical quantity to a number; the microcontroller processes it and controls the node; the transceiver sends and receives packets; memory holds the program, variables and logged data; the power unit supplies all of them. Optional units: location finding system, power generator, mobilizer.

2. Why is a microcontroller used as the processing unit? It is cheap, reprogrammable, has memory and interfaces on one chip, and above all sleeps at microamperes and wakes within microseconds, which matters more than speed for a node that is asleep nearly all the time.

munotes.in43

Inside a Sensor Node: The Five Units

3. A node's 12-bit ADC measures 0 to 2.5 volts. What is the size of one step? 2.5 / 4,096 volts, about 0.61 millivolts.

4. Why did the Telos designers care that the MSP430 runs down to 1.8 volts? Because two AA cells in series are exhausted at about 1.8 volts. A processor that stops at 2.7 volts, like the ATmega128, leaves almost half of the batteries' energy unused.

5. Which unit uses the most energy, and what follows for the design? The transceiver: on the CC2420 about 17 to 19 mA when on, against about 0.33 mA for the processor. So the node must keep the radio off as much as possible and compute more in order to send less.

Contents This chapter on its own page

munotes.in44

Chapter Eight

The Radio, the Sensors and the Power Supply of a Node

Syllabus topic Module 1, "Introduction and Overview of WSNs: Sensor node technology"

In one line

A node's radio moves between sleeping, idling, listening and sending, each at a very different current and each change taking time; its sensors cost little unless they have to heat or transmit; and its battery delivers less energy the harder it is worked.

In the wording a student can write in an examination: a sensor node's transceiver has four operating states, transmit, receive, idle and sleep (power down), with receive and transmit drawing far more current than idle, and idle far more than sleep. Changing state costs time and energy, so sleeping pays only if the sleep is long enough. Sensors are passive or active and may need warm-up time; actuators convert a command into physical action and may draw large currents. The power supply is usually a primary battery, whose usable capacity falls at high discharge currents (the rate-capacity effect) and whose voltage falls as it empties; it may be helped by energy scavenging from light, vibration, heat or radio waves.

The radio's four states

Every low-power radio has the same four states, and the CC2420 shows them with real numbers.

The CC2420's states, their currents and the time to move between them

Figure 8.1 The radio's states and the time to move between them (CC2420)

  1. Transmit. The transmitter and power amplifier are on and a frame is being sent. 17.4 mA at an output power of 0 dBm.
  2. Receive, which includes listening. The receiver is on, whether or not anything arrives. 18.8 mA. There is no cheaper "just listening" mode: a radio waiting for a packet is in the receive state.
  3. Idle. The crystal oscillator is running, so the radio is ready to switch to receive or transmit quickly, but neither is on. 0.426 mA.
  4. Sleep, or power down. The oscillator is off. 20 microamperes with the voltage regulator on, and 0.02 microamperes with it off as well.

Moving between states costs time

The radio cannot jump from sleep to sending. It has to come up through the states, and each step takes time during which it draws current but cannot communicate.

  • Turning the voltage regulator on: up to 0.6 ms.
  • Starting the crystal oscillator, from power down to idle: 1.0 ms (Telos, with a carefully chosen crystal, measured 580 microseconds).
  • Locking the frequency synthesiser, from idle to receive or transmit: 192 microseconds.
  • Turning around between receive and transmit: also 192 microseconds.

The datasheet draws the conclusion itself: "Due to the very fast start-up time, CC2420 can remain in Power Down until a transmission session is requested." For an older, slower radio the answer is different. The Telos paper lists the CC1000 of the Mica2 at 2 ms to turn on and another radio at 3 ms, and for such radios very short sleeps are not worth taking.

munotes.in45

The Radio, the Sensors and the Power Supply of a Node

When sleeping pays: the break-even rule

Going to sleep saves energy only if the energy saved while asleep is more than the energy spent waking up. So there is a break-even time: sleep for longer than that and you save; sleep for less and you lose.

Take the worst case for the CC2420. Suppose that waking cost as much as listening, 18.8 mA, for the whole 1.0 ms of the crystal start-up. The energy spent waking is then what 1.0 ms of listening costs, and staying asleep instead of listening saves 18.8 mA less 0.02 mA, almost all of it. So even in this pessimistic case any sleep longer than about 1 ms pays for itself. Because a real start-up draws less than full receive current, the true break-even time is shorter still.

The rule in general: break-even time = energy to wake up / (power while awake - power while asleep), where "awake" means whatever the radio would otherwise be doing, listening in this example. A radio with a slow, costly start-up has a long break-even time; one with a fast start-up can sleep between every packet. This is why the Telos designers cared so much about wake-up times, and why [Duty Cycling: Preamble Sampling, B-MAC and X-MAC] works at all.

Transmit power is not transmit current

The CC2420 lets the software choose its output power, and the datasheet gives the current at each setting.

Output powerCurrent drawn
0 dBm17.4 mA
-5 dBm14 mA
-10 dBm11 mA
-15 dBm9.9 mA
-25 dBm8.5 mA

Going from 0 dBm to -25 dBm cuts the radiated power by a factor of 10 to the power 2.5, about 316 (each 10 dB is a factor of 10). But the current only falls from 17.4 mA to 8.5 mA, about half. Most of a low-power radio's current goes into its electronics (the oscillator, synthesiser, mixer and baseband), not into the power actually radiated. So turning the transmit power down saves less than a student would guess, and receiving, which has no power amplifier at all, costs more than transmitting. Both facts shape the argument in [Single Hop or Multiple Hops: The Energy Argument Worked Out].

Energy per bit: why a faster radio can be cheaper

What matters is not the current but the energy to move one bit, which is power divided by data rate. The Telos paper's table gives both for two generations of mote. In milliwatts divided by kilobits per second, which comes out in microjoules per bit:

  • Mica2, CC1000 radio, transmitting at 42 mW and 38.4 kbps: 42 / 38.4 = 1.09375 microjoules per bit.
  • Telos, CC2420 radio, transmitting at 35 mW and 250 kbps: 35 / 250 = 0.14 microjoules per bit.
munotes.in46

The Radio, the Sensors and the Power Supply of a Node

So the newer radio sends a bit for 1.09375 / 0.14 = 7.8125 times less energy, although it draws almost the same power, because it finishes about six and a half times sooner and then sleeps. The Telos authors put it simply: "The higher data rate allows shorter active periods further reducing energy consumption."

Sensors and actuators

Sensors. The kinds of sensor, and how they are classified, are the subject of [Sensor Taxonomy]. What matters for the node's hardware is:

  • Passive sensors only observe (a thermometer, a light sensor, a microphone) and typically draw little.
  • Active sensors emit something to measure (an ultrasonic or radar range finder sends a pulse and times the echo) and can draw more than the radio while they work.
  • Some sensors need warming up: a gas sensor heats its element before it reads correctly, which can take seconds and a lot of current, so it cannot simply be switched on for a millisecond.
  • Analogue or digital output: an analogue sensor goes through the node's ADC; a digital sensor has its own converter and talks to the microcontroller over a serial interface (I2C or SPI).

Actuators do the opposite of a sensor: they turn a command into action, a relay, a valve, a buzzer, a motor. In a sensor and actuator network the node that detects a dry field may also open the irrigation valve. Actuators often need far more current than the node's battery can supply, so they are usually powered separately and only switched by the node.

The power supply

Batteries

A primary battery is used once and thrown away; a secondary (rechargeable) battery can be recharged, which matters only if something can recharge it, such as a solar cell. Most nodes use primary cells, because they hold more energy for their size and lose less of it while they sit unused.

Two real AA cells from their datasheets:

Energizer E91Energizer L91
ChemistryAlkaline, zinc and manganese dioxideLithium and iron disulfide
Nominal voltage1.5 V1.5 V
Weight23.0 grams15 grams
Operating temperature-18 to 55 degrees C-40 to 60 degrees C
Shelf life at 21 degrees C10 years25 years

For a node left outdoors for years, the lithium cell's wider temperature range and longer shelf life can be worth its higher price.

The battery never gives its label

Three effects mean a battery delivers less than its rated capacity.

The rate-capacity effect. The harder a cell is worked, the less charge it delivers in total. The E91 datasheet's chart of capacity at continuous discharge (to 0.8 volts, at 21 degrees C) reads, approximately, 3,000 mAh at 25 mA, 2,500 mAh at 100 mA, 2,000 mAh at 250 mA and 1,500 mAh at 500 mA. Halve the capacity by drawing twenty times the current. A sensor node draws tiny average currents, which is good for this effect, but its radio draws its 18 mA in bursts.

munotes.in47

The Radio, the Sensors and the Power Supply of a Node

The voltage falls as the cell empties. An AA cell starts at about 1.5 volts and is considered exhausted at about 0.9 volts. Two in series therefore end at about 1.8 volts. The electronics must keep working down to that voltage or the rest of the energy is stranded, which is exactly the argument the Telos designers made for the MSP430 (runs to 1.8 V) against the ATmega128 (stops at 2.7 V).

A converter is not free either. A DC-DC converter can hold the voltage steady as the cell empties, but it wastes some energy in doing so and draws current even when the node sleeps. The original Mica had a boost converter; the Mica2 discarded it.

Energy in a cell, worked. Energy is voltage times charge. An AA cell at 1.5 volts delivering 3,000 mAh holds 1.5 × 3,000 = 4,500 milliwatt-hours, which is 4.5 watt-hours. A watt-hour is 3,600 joules, so the cell holds 4.5 × 3,600 = 16,200 joules, and two cells about 32,400 joules. Every protocol in this book is spending from that one budget.

Energy scavenging

A node can top up its battery by scavenging (or harvesting) energy from its surroundings.

  • Light, with a small solar cell: by far the most useful outdoors in daylight, and much weaker under indoor lighting.
  • Vibration, from machines, vehicles or bridges, through a piezoelectric or electromagnetic generator.
  • Temperature differences, through a thermoelectric generator, for example between a hot pipe and the air.
  • Airflow or water flow, through a small turbine.
  • Radio waves, collected by an antenna, which yields very little unless a transmitter is close.

Scavenged energy is irregular (no sun at night, no vibration when the machine stops), so it charges a rechargeable cell or a supercapacitor, and the node spends from that store. The goal is energy-neutral operation: over each day the node uses no more energy than it collects, and then it can in principle run for as long as its hardware lasts.

Where the day's energy actually goes

The three subsystems have now been described one at a time. Put together over a whole day, they do not share the bill evenly, and which one dominates is not fixed.

# Where a day's energy goes across the three subsystems, at four duty cycles.
V = 3.0                      # volts, two AA cells
I_TX, I_RX, I_IDLE = 17.4e-3, 18.8e-3, 0.426e-3     # CC2420, amperes
I_SLEEP = 20e-6                                      # power down, regulator on
I_MCU_ON, I_MCU_SLEEP = 0.5e-3, 1.1e-6               # a representative low power MCU
I_SENSOR, T_SENSOR = 1.0e-3, 0.05                    # a representative sensor: 1 mA for 50 ms
DAY = 24 * 3600.0

print("A day in the life of a node, in joules, at four radio duty cycles.")
print("Two AA cells at %.1f V; the radio is the CC2420's own figures; the MCU and the" % V)
print("sensor are representative parts, and the sensor is read once a minute.")
print()
print("  radio on   radio        MCU      sensor      sleep       total   radio's share")
samples = DAY / 60.0
e_sensor = samples * T_SENSOR * I_SENSOR * V
for duty in (1.0, 0.1, 0.01, 0.001):
    t_on = DAY * duty
    # half the on time listening, half sending, which is generous to the radio
    e_radio = t_on * (I_RX * 0.5 + I_TX * 0.5) * V
    e_mcu = t_on * I_MCU_ON * V + (DAY - t_on) * I_MCU_SLEEP * V
    e_sleep = (DAY - t_on) * I_SLEEP * V
    total = e_radio + e_mcu + e_sensor + e_sleep
    print("  %7.1f%% %8.1f J %8.2f J %8.2f J %8.2f J %9.1f J %8.0f%%"
          % (duty * 100, e_radio, e_mcu, e_sensor, e_sleep, total,
             100 * e_radio / total))
print()
print("At a hundred per cent the radio is everything and nothing else is worth measuring.")
print("At a tenth of a per cent the sensor and the sleep current are the bill, and making")
print("the radio more efficient would change almost nothing. The subsystem to optimise is")
print("not a property of the node: it is a property of the duty cycle.")

# The crossing point: the duty cycle at which the radio stops dominating.
lo, hi = 1e-6, 1.0
for _ in range(80):
    mid = (lo + hi) / 2
    t_on = DAY * mid
    e_radio = t_on * (I_RX * 0.5 + I_TX * 0.5) * V
    rest = (t_on * I_MCU_ON * V + (DAY - t_on) * I_MCU_SLEEP * V + e_sensor
            + (DAY - t_on) * I_SLEEP * V)
    if e_radio > rest:
        hi = mid
    else:
        lo = mid
print()
print("The radio stops being the largest share below a duty cycle of %.3f per cent," % (hi * 100))
print("which is %.1f seconds of radio in a whole day." % (DAY * hi))
munotes.in48

The Radio, the Sensors and the Power Supply of a Node

A day in the life of a node, in joules, at four radio duty cycles.
Two AA cells at 3.0 V; the radio is the CC2420's own figures; the MCU and the
sensor are representative parts, and the sensor is read once a minute.

  radio on   radio        MCU      sensor      sleep       total   radio's share
    100.0%   4691.5 J   129.60 J     0.22 J     0.00 J    4821.3 J       97%
     10.0%    469.2 J    13.22 J     0.22 J     4.67 J     487.3 J       96%
      1.0%     46.9 J     1.58 J     0.22 J     5.13 J      53.8 J       87%
      0.1%      4.7 J     0.41 J     0.22 J     5.18 J      10.5 J       45%

At a hundred per cent the radio is everything and nothing else is worth measuring.
At a tenth of a per cent the sensor and the sleep current are the bill, and making
the radio more efficient would change almost nothing. The subsystem to optimise is
not a property of the node: it is a property of the duty cycle.

The radio stops being the largest share below a duty cycle of 0.124 per cent,
which is 107.5 seconds of radio in a whole day.
munotes.in49

The Radio, the Sensors and the Power Supply of a Node

Distinctions

Receive (listen)TransmitIdleSleep
CC2420 current18.8 mA17.4 mA at 0 dBm0.426 mA0.02 mA
Can hear a packetYesNoNoNo
Time to reach send or receiveTurnaround 192 microsecondsTurnaround 192 microseconds192 microsecondsAbout 1 ms more
Primary batteryRechargeable battery with scavenging
EnergyFixed, spent onceReplenished, irregularly
Lifetime set byCapacity divided by average powerWhether harvest covers use, and cycle life
Typical useMost nodesOutdoor nodes with solar cells

What it does not mean

Listening is not cheap. On the CC2420 it costs more than sending. A node "just waiting" with its receiver on is spending at full rate.

Lower transmit power is not proportionally lower current. A 316-fold cut in radiated power roughly halves the current, because the electronics dominate.

A higher data rate is not a higher energy cost. It is usually lower per bit, because the radio finishes sooner and sleeps.

Sleeping is not always a saving. A sleep shorter than the break-even time costs more to wake from than it saves. With fast radios like the CC2420 that time is around a millisecond; with slower ones it is longer.

A 3,000 mAh label is not 3,000 mAh in service. It depends on the current drawn, the temperature and the cut-off voltage of the electronics.

Quick revision

  • Radio states (CC2420): transmit 17.4 mA, receive 18.8 mA, idle 0.426 mA, power down 0.02 mA.
  • Transitions: regulator up to 0.6 ms, crystal 1.0 ms (Telos 580 microseconds), idle to receive or transmit and turnaround 192 microseconds.
  • Break-even time = energy to wake / (awake power - sleep power); for the CC2420 against listening, around 1 ms even in the worst case.
  • Transmit current: 8.5 mA at -25 dBm to 17.4 mA at 0 dBm: a factor of about 316 in power for about 2 in current.
  • Energy per bit: Mica2 1.09375, Telos 0.14 microjoules, 7.8125 times less for the faster radio.
  • Sensors: passive or active, some need warm-up; actuators usually powered separately.
  • AA cells: E91 alkaline and L91 lithium, 1.5 V; E91 capacity falls from about 3,000 mAh at 25 mA to about 1,500 mAh at 500 mA (rate-capacity effect); cut-off about 0.9 V a cell.
  • One AA: 1.5 × 3,000 = 4,500 mWh, 16,200 joules. Scavenging: light, vibration, heat, flow, radio; aim for energy-neutral operation.
munotes.in50

The Radio, the Sensors and the Power Supply of a Node

Test yourself

1. Name the four states of a sensor node's radio and give the CC2420's current in each. Transmit (17.4 mA at 0 dBm), receive or listen (18.8 mA), idle with the oscillator running (0.426 mA), and sleep or power down (20 microamperes, or 0.02 microamperes with the regulator off).

2. What is the break-even time for sleeping, and why does it matter? The shortest sleep for which the energy saved exceeds the energy spent waking up: energy to wake divided by the difference between the power while awake and the power while asleep. A shorter sleep wastes energy, so a MAC protocol must not put the radio to sleep for gaps shorter than this.

3. The CC2420's output power is cut from 0 dBm to -25 dBm. What happens to the radiated power and to the current, and what does that teach? The radiated power falls by a factor of about 316; the current falls only from 17.4 mA to 8.5 mA, about half. Most of the current goes into the radio's electronics, so cutting transmit power saves less than expected.

4. Compare the energy per bit of the Mica2 and Telos radios. Mica2: 42 mW at 38.4 kbps, 1.09375 microjoules per bit. Telos: 35 mW at 250 kbps, 0.14 microjoules per bit, 7.8125 times less, because the faster radio finishes its transmission sooner.

5. What is the rate-capacity effect? Give the figures from the E91 datasheet. A battery delivers less total charge when discharged at a higher current. The E91 gives roughly 3,000 mAh at 25 mA, 2,500 at 100 mA, 2,000 at 250 mA and 1,500 at 500 mA.

6. How much energy does one 1.5 V AA cell of 3,000 mAh hold? 1.5 × 3,000 = 4,500 milliwatt-hours, 4.5 watt-hours, which is 4.5 × 3,600 = 16,200 joules.

7. List four sources a node can scavenge energy from, and say why scavenged energy needs a store. Light, vibration, temperature differences, air or water flow, and radio waves. The supply is irregular, so it charges a rechargeable cell or supercapacitor that the node draws on steadily.

Contents This chapter on its own page

munotes.in51

Chapter Nine

How Long a Node Lasts: The Energy Budget Worked Out

Syllabus topic Module 1, "Introduction and Overview of WSNs: Sensor node technology"

In one line

A node's lifetime is the charge its battery can deliver divided by the average current it draws, and the average current is each part's current weighted by the fraction of time it is on.

In the wording a student can write in an examination: if a node spends a fraction d of its time awake, drawing a current I(awake), and the rest asleep, drawing I(sleep), its average current is

I(avg) = d × I(awake) + (1 - d) × I(sleep)

and a battery of usable capacity C (in milliampere-hours) lasts C / I(avg) hours. The fraction d is the node's duty cycle. Lifetime is therefore set by three things: the duty cycle, the awake current (dominated by the radio) and, at low duty cycles, the sleep current.

Why this sum is worth learning

Because it turns every energy argument in the book into a number. S-MAC saves energy by sleeping is an assertion; S-MAC at a 10 per cent duty cycle lasts about ten times as long as an always-on radio is an answer. And the sum is the tool a designer actually uses to decide whether a network will last a season.

The method, in four steps

  1. List every state the node can be in and its current, from the datasheets.
  2. Find the fraction of time spent in each state, from how often the node wakes and how long it stays awake.
  3. Add up current times fraction to get the average current.
  4. Divide the battery's usable capacity by the average current, then check the answer against the battery's shelf life, because a battery also runs down while doing nothing.

Worked by hand: a 1 per cent duty cycle

The node is the Telos pair of parts, running from two AA cells at about 3 volts.

Step 1, the currents. Awake, the radio listens at 18.8 mA and the processor runs at 0.5 mA, so I(awake) = 18.8 + 0.5 = 19.3 mA. Asleep, the radio is powered down at 0.02 mA and the processor is in standby at 0.002 mA, so I(sleep) = 0.02 + 0.002 = 0.022 mA.

Step 2, the fractions. Awake 1 per cent of the time, d = 0.01, and asleep for the remaining 0.99.

Step 3, the average current.

I(avg) = 0.01 × 19.3 + 0.99 × 0.022

= 0.193 + 0.02178

= 0.21478

so the node draws 0.21478 mA on average.

Step 4, the lifetime. Two AA cells in series give 3 volts, but their capacity in mAh does not add: the same charge flows through both. The E91 alkaline cell's chart gives about 3,000 mAh at 25 mA; allowing for cold nights, for the cut-off voltage and for the cells' age, we allow ourselves 2,500 mAh. The lifetime is 2,500 / 0.21478, about 11,640 hours, which is about 485 days, or 1.3 years.

munotes.in52

How Long a Node Lasts: The Energy Budget Worked Out

Compare that with the 5.4 days the same node lasts with its radio always on, which is the first line of the program's output below. Sleeping 99 per cent of the time has bought about ninety times the lifetime.

The same sum, run across duty cycles

# Lifetime of a Telos-style node (CC2420 radio, MSP430F1611 processor)
# against its duty cycle. Currents in mA, from the two datasheets at 3 V.
RADIO_ON = 18.8      # CC2420 receiving (transmitting is 17.4, a little less)
RADIO_SLEEP = 0.02   # CC2420 power down, voltage regulator still on
CPU_ON = 0.5         # MSP430F1611 active at 1 MHz and 3 V
CPU_SLEEP = 0.002    # MSP430F1611 standby (LPM3) at 3 V and 25 degrees C
BATTERY = 2500       # mAh we allow ourselves from two AA cells in series

def average_current(duty):
    """Radio and processor awake for a fraction `duty` of the time."""
    awake = RADIO_ON + CPU_ON
    asleep = RADIO_SLEEP + CPU_SLEEP
    return duty * awake + (1 - duty) * asleep

print("duty cycle   average mA    lifetime")
for duty in (1, 0.1, 0.01, 0.001, 0.0001, 0):
    mA = average_current(duty)
    days = BATTERY / mA / 24
    if days < 365:
        life = "%.1f days" % days
    else:
        life = "%.1f years" % (days / 365)
    print("%9.2f %%   %10.4f   %s" % (duty * 100, mA, life))
duty cycle   average mA    lifetime
   100.00 %      19.3000   5.4 days
    10.00 %       1.9498   53.4 days
     1.00 %       0.2148   1.3 years
     0.10 %       0.0413   6.9 years
     0.01 %       0.0239   11.9 years
     0.00 %       0.0220   13.0 years

Read the lifetime column from the top, and three things stand out.

Always on is hopeless. 5.4 days, which is what [The Challenges of Wireless Sensor Networks] found with a rougher sum.

At first, every tenfold cut in duty cycle buys almost tenfold life. From 100 to 10 per cent, and from 10 to 1 per cent, the lifetime grows by nearly ten each time, because the awake current is almost the whole average.

Then the curve flattens. From 0.1 per cent to 0.01 per cent the lifetime does not grow tenfold, and even a node that never wakes at all (the last line) lasts only 13 years. That ceiling is set by the sleep current alone: 2,500 mAh divided by 0.022 mA. Below about 0.1 per cent, making the node sleep even more hardly helps; making its sleep deeper does.

A real cycle, item by item

Meera's vineyard node wakes every 15 minutes for 60 milliseconds. What it does in that time is an assumption about the software, stated here so the reader can change it: the sensors are powered for 20 ms at 1 mA, the radio listens for 40 ms (waiting for the channel, for the acknowledgement and for a neighbour's packet to forward) and transmits for 4 ms. The listing adds up the charge of each item, in milliampere-milliseconds, and compares two ways of putting the radio to sleep.

munotes.in53

How Long a Node Lasts: The Energy Budget Worked Out

# One 15-minute cycle of the vineyard node, item by item.
# Charge is current (mA) times time (ms); dividing the total by the cycle
# length gives the average current the battery must supply.
CYCLE_MS = 15 * 60 * 1000
AWAKE_MS = 60

def budget(radio_sleep):
    items = [
        ("processor awake",     0.5,   AWAKE_MS),
        ("sensors powered",     1.0,   20),
        ("radio listening",     18.8,  40),
        ("radio transmitting",  17.4,  4),
        ("processor asleep",    0.002, CYCLE_MS - AWAKE_MS),
        ("radio asleep",        radio_sleep, CYCLE_MS - AWAKE_MS),
    ]
    total = sum(mA * ms for _, mA, ms in items)
    for name, mA, ms in items:
        share = 100 * mA * ms / total
        print("  %-19s %7.3f mA × %6d ms = %9.1f   %5.1f %%"
              % (name, mA, ms, mA * ms, share))
    average = total / CYCLE_MS
    years = 2500 / average / 24 / 365
    print("  total %.1f mA ms; average %.5f mA; lifetime %.1f years"
          % (total, average, years))

print("Radio powered down between cycles (20 microamperes):")
budget(0.02)
print()
print("Radio's voltage regulator off as well (0.02 microamperes):")
budget(0.00002)
Radio powered down between cycles (20 microamperes):
  processor awake       0.500 mA ×     60 ms =      30.0     0.1 %
  sensors powered       1.000 mA ×     20 ms =      20.0     0.1 %
  radio listening      18.800 mA ×     40 ms =     752.0     3.6 %
  radio transmitting   17.400 mA ×      4 ms =      69.6     0.3 %
  processor asleep      0.002 mA × 899940 ms =    1799.9     8.7 %
  radio asleep          0.020 mA × 899940 ms =   17998.8    87.1 %
  total 20670.3 mA ms; average 0.02297 mA; lifetime 12.4 years

Radio's voltage regulator off as well (0.02 microamperes):
  processor awake       0.500 mA ×     60 ms =      30.0     1.1 %
  sensors powered       1.000 mA ×     20 ms =      20.0     0.7 %
  radio listening      18.800 mA ×     40 ms =     752.0    28.0 %
  radio transmitting   17.400 mA ×      4 ms =      69.6     2.6 %
  processor asleep      0.002 mA × 899940 ms =    1799.9    66.9 %
  radio asleep          0.000 mA × 899940 ms =      18.0     0.7 %
  total 2689.5 mA ms; average 0.00299 mA; lifetime 95.5 years

The first budget is the surprise of this chapter. The radio's listening and transmitting, the part everyone worries about, is about 4 per cent of the charge. The radio asleep, at 20 microamperes, is 87 per cent, because it is asleep for 899,940 of every 900,000 milliseconds. At this duty cycle the job is no longer to sleep more but to sleep deeper.

munotes.in54

How Long a Node Lasts: The Energy Budget Worked Out

The second budget switches the radio's voltage regulator off as well, the CC2420's 0.02-microampere state, which costs only a slightly longer wake-up (up to 0.6 ms, from [The Radio, the Sensors and the Power Supply of a Node]). The average current falls from about 0.023 mA to about 0.003 mA, nearly eight times less, and now the processor's standby is the largest item.

But neither lifetime can be believed as it stands. 12.4 years and 95.5 years are longer than the batteries themselves last on a shelf: the E91 datasheet gives its shelf life as 10 years at 21 degrees C, the L91 lithium cell's as 25 years. When the electronics draw this little, the battery's own chemistry, not the node, sets the lifetime. A designer who stopped at the program's last line would promise the farmer a century.

What the model leaves out

A good answer names the limits of its own sum.

  • Forwarding. This node sends only its own packet. A node near the sink also receives and forwards everyone else's, so it listens and transmits many times longer and dies first. The network's lifetime is set by those nodes, which is why [Optimization Goals: Quality of Service, Energy Efficiency and Lifetime] defines lifetime in several ways.
  • Retransmissions and collisions, which add radio time whenever the channel is busy or a packet is lost.
  • The battery's behaviour: the rate-capacity effect on the radio's 18 mA bursts, the effect of heat and cold, and self-discharge over the years.
  • Start-up energy, the milliseconds of oscillator start-up at each wake-up; small here, larger for a node that wakes very often.
  • The cut-off voltage, which decides how much of the battery the electronics can actually use.

Distinctions

Node lifetimeNetwork lifetime
IsHow long one node's battery lastsHow long the network does its job
Set byThat node's duty cycle and currentsThe nodes that work hardest, usually those near the sink
Worked hereYesIn the optimization goals chapter
Duty cycleWhat limits the lifetime
High, 1 per cent and aboveThe awake current, above all the radio listening
Low, around 0.1 per centBoth the awake and the sleep currents
Very low, below 0.01 per centThe sleep current, then the battery's shelf life

What it does not mean

Halving the radio's active current does not double the lifetime at low duty cycles. Once the sleep current dominates, the active current hardly matters, as the vineyard budget showed.

Two cells in series do not double the mAh. They double the voltage; the charge that flows is the same through both.

munotes.in55

How Long a Node Lasts: The Energy Budget Worked Out

A computed lifetime longer than the battery's shelf life is not a real lifetime. It means the chemistry will expire first.

The node's lifetime is not the network's. A relay near the sink may last a tenth as long as a node at the edge of the field.

Quick revision

  • I(avg) = d × I(awake) + (1 - d) × I(sleep); lifetime = C / I(avg).
  • Telos parts at 3 V: I(awake) = 18.8 + 0.5 = 19.3 mA; I(sleep) = 0.02 + 0.002 = 0.022 mA.
  • 1 per cent duty cycle: 0.21478 mA; 2,500 mAh lasts about 1.3 years. Always on: 5.4 days.
  • Every tenfold cut in duty cycle buys nearly tenfold life until the sleep current takes over; with no waking at all, the ceiling is 13 years.
  • Vineyard cycle: radio asleep is 87 per cent of the charge; switching its regulator off too cuts the average about eight times.
  • Then the battery's shelf life (E91 10 years, L91 25 years) is the real limit.

Test yourself

1. A node draws 19.3 mA awake and 0.022 mA asleep, and is awake 1 per cent of the time. Find its average current and its lifetime on 2,500 mAh. 0.01 × 19.3 + 0.99 × 0.022 = 0.193 + 0.02178 = 0.21478 mA. Lifetime 2,500 / 0.21478, about 11,640 hours, about 1.3 years.

2. The same node is left with its radio always on. How long does it last, and what does the comparison show? 2,500 / 19.3, about 130 hours, 5.4 days. The 1 per cent duty cycle buys about ninety times the lifetime, which is why sensor network radios are switched off almost all the time.

3. Why does cutting the duty cycle from 0.01 per cent to 0.001 per cent hardly change the lifetime? Because at such low duty cycles the sleep current makes up almost all of the average current, and it is not affected by how often the node wakes. Only a deeper sleep state helps.

4. In the vineyard budget, what fraction of the charge goes on the radio while it is asleep, and why? About 87 per cent. The radio draws only 20 microamperes asleep, but it is asleep for 899,940 of every 900,000 milliseconds, while it listens and transmits for only 44.

5. The program predicts 95.5 years. What is wrong with that answer? It ignores the battery's own limits. An E91 alkaline cell has a shelf life of about 10 years at 21 degrees C, so the chemistry, not the node's consumption, sets the lifetime. The prediction also ignores forwarding, retransmissions, temperature and the cut-off voltage.

munotes.in56

How Long a Node Lasts: The Energy Budget Worked Out

6. Two AA cells of 2,500 mAh are connected in series. What voltage and what capacity does the node see? About 3 volts, and still 2,500 mAh: series connection adds voltages, not charge.

Contents This chapter on its own page

munotes.in57

Chapter Ten

Sensor Taxonomy

Syllabus topic Module 1, "Introduction and Overview of WSNs: Sensor taxonomy"

In one line

Sensor taxonomy is the classification of sensors (by what they measure and how) and of sensor networks (by how data is delivered, what moves, and how the network is built).

In the wording a student can write in an examination: sensors are classified by the quantity they measure; as passive or active; as omnidirectional or narrow-beam; by output (analogue or digital); as contact or non-contact; and by operating principle. Sensor networks are classified by their data delivery model (continuous, event-driven, observer-initiated or hybrid), their network dynamics (static, or dynamic with a mobile observer, mobile sensors or a mobile phenomenon), their communication (application and infrastructure, cooperative and non-cooperative), their structure (flat or hierarchical, single hop or multi-hop), their composition (homogeneous or heterogeneous) and their deployment (planned or random).

Why classify at all

A classification is useful because each class needs something different, so knowing the class tells you the design. A continuous-delivery network suits clustering; an event-driven one must handle a burst of reports all at once; a network with mobile sensors must keep repairing its paths. The Tilak paper was written for exactly this reason: so that a designer can recognise which kind of network they have and choose protocols that fit it.

The two halves of sensor taxonomy

Figure 10.1 Sensor taxonomy: classifying sensors, and classifying networks

Part 1: classifying sensors

By what they measure

The first and most obvious division.

GroupQuantitiesExamples of use
ThermalTemperature, heat flowCold stores, burrows, machine bearings
MechanicalPressure, force, strain, acceleration, vibration, tiltBridges, machines, earthquakes
AcousticSound, ultrasound, infrasoundChainsaw detection, volcano infrasound
OpticalLight intensity, infrared, imagesOccupancy, presence of a warm animal
MagneticMagnetic fieldDetecting vehicles by their steel
ChemicalGases, humidity, pH, pollutantsAir quality, soil, gas leaks
Biological and medicalHeart rate, blood oxygen, body temperaturePatient monitoring
Position and motionLocation, proximity, movementTracking, intrusion detection
Flow and radiationAir and liquid flow, ionising radiationVentilation, nuclear safety

Passive or active

Karl and Willig divide sensors into three kinds, and the first word is the division that matters most for energy.

  • Passive sensors only observe. They take in a physical quantity without sending anything out: a thermometer, a light sensor, a microphone, a passive infrared detector. They are cheap in energy.
  • Active sensors probe their surroundings: they emit energy and measure what comes back. An ultrasonic range finder sends a pulse of sound and times its echo; a radar sends a radio pulse. They can draw as much as the radio while they work.

The words passive and active are also used for whether a sensor needs an external power supply. Dargie and Poellabauer, the first reference book on MU's list, combine the two ideas: an active sensor requires external power and "must emit some kind of energy (e.g., microwaves, light, sound)", while a passive sensor detects energy already in the environment and derives its power from it, a passive infrared detector being their example. For an examination the safe statement is the one both books share: an active sensor emits energy and measures the response; a passive sensor only receives.

munotes.in58

Sensor Taxonomy

Omnidirectional or narrow-beam

The second division is about direction, and it applies to passive sensors.

  • Omnidirectional sensors respond equally from every direction: a thermometer, an ordinary microphone.
  • Narrow-beam sensors see only in one direction: a camera, a directional microphone, a passive infrared detector behind a lens.

So Karl and Willig's three kinds are passive omnidirectional, passive narrow-beam and active. For the network the difference matters: a narrow-beam sensor's reading depends on which way it points, so its orientation has to be known, and covering an area takes more of them.

By output, by contact and by principle

  • Analogue or digital output. An analogue sensor gives a voltage or current that the node's ADC converts; a digital sensor has its own converter and sends a number over a serial bus.
  • Contact or non-contact. A soil probe or a strain gauge must touch what it measures; an infrared thermometer or a camera measures from a distance.
  • Operating principle. How the quantity becomes an electrical signal: a change in resistance (thermistors, strain gauges), in capacitance (humidity and soil-moisture sensors), in inductance (proximity and position sensors, where a moving core changes a coil's inductance), a piezoelectric voltage from stress (vibration sensors), a photoelectric current from light, a thermoelectric voltage from a temperature difference, or a micro-electro-mechanical (MEMS) structure on a chip, as in most accelerometers.

Part 2: classifying sensor networks

The three parts of every sensing application

Tilak, Abu-Ghazaleh and Heinzelman begin by naming three participants, and the rest of their taxonomy is about how the three relate.

  • Sensor: the device that implements the physical sensing and reports its measurements; in their words it typically consists of sensing hardware, memory, battery, embedded processor, and trans-receiver.
  • Observer: the end user who wants information about the phenomenon, who may indicate interests (or queries) to the network and receive responses.
  • Phenomenon: "The entity of interest to the observer", the thing being sensed and potentially analysed; several may be observed at once.

By data delivery model

How and when readings are sent to the observer. This is the classification most often asked.

  1. Continuous. The sensors send their data continuously at a set rate: the vineyard, every 15 minutes. The paper notes that clustering is most efficient for static networks with continuous data.
  2. Event-driven. The sensors report only when an event of interest occurs: an intruder, a fire. Most of the time the network is silent; when the event happens, many nearby sensors detect it together and contend for the channel at once, which raises both the risk of losing the critical report and its delay.
  3. Observer-initiated, or request-reply. The sensors report only in response to an explicit request from the observer, sent to them directly or through other sensors: what is the temperature in zone 3 now?
  4. Hybrid. The three can coexist in one network: continuous background readings, event alarms and answers to queries all at once.
munotes.in59

Sensor Taxonomy

By network dynamics

What moves. The paper divides networks into static, where nothing moves (sensors, observer and phenomenon are all fixed, as in a group of temperature sensors), and dynamic, where something does. Three kinds of motion:

  • Mobile observer. The observer moves while the sensors stay put: a vehicle or a person carrying the collecting device drives past the field and gathers readings as it goes.
  • Mobile sensors. The sensors themselves move: sensors on animals, vehicles or drifting buoys. Paths to the observer break as sensors move and must be rebuilt.
  • Mobile phenomena. What is being watched moves: a vehicle, an animal, a fire front. The set of sensors that should report changes as the phenomenon moves; in the paper's words, sensors can hand the "responsibility of monitoring to a closer node" as the target drifts.

Motion is the commonest source of change, but the paper notes others: sensors failing and observers' interests changing.

By communication

The paper divides the network's traffic into two kinds.

  • Application communication carries the sensed data towards the observer. It is cooperative when sensors work together to meet the observer's interest (for example a cluster head combining its members' readings) and non-cooperative when each sensor reports on its own.
  • Infrastructure communication is the traffic needed to set up, maintain and optimise the network: discovering neighbours, building paths, rebuilding them when a sensor moves or dies. It is overhead, so a good protocol keeps it small, though spending some on it can reduce application traffic overall.

By structure, composition and deployment

  • Flat or hierarchical. In a flat network every node has the same role; in a hierarchical one some nodes are cluster heads. Sohraby, Minoli and Znati make the related division into two categories: networks that are mesh-based with multi-hop radio links and dynamic routing, and networks that are point-to-point or star-based with single-hop radio links and static routing to a node on the fixed network.
  • Homogeneous or heterogeneous. All nodes the same, or some with more energy, better radios or extra sensors, such as mains-powered cluster heads.
  • Planned or random deployment. Nodes placed at chosen points, or scattered, for example from the air.
munotes.in60

Sensor Taxonomy

The text book's own categorisation

Sohraby, Minoli and Znati, in their chapter 1 (Table 1.1, which they adapt from an earlier source), sort the issues of a sensor network under four headings. It is a compact taxonomy in its own right, and it overlaps with everything above.

HeadingDimensionThe two ends
SensorsSizeSmall MEMS devices to large ones such as radars and satellites
MobilityStationary (seismic sensors) or mobile (on robot vehicles)
TypePassive (acoustic, seismic, video, infrared, magnetic) or active (radar, ladar)
Operating environmentMonitoring requirementDistributed (environmental monitoring) or localised (target tracking)
Number of sitesSmall, but usually large
Spatial coverageDense or sparse; multi-hop or single-hop
DeploymentFixed and planned (factory networks) or ad hoc (air-dropped)
EnvironmentBenign (a factory floor) or adverse (a battlefield)
NatureCooperative (air traffic control) or non-cooperative (military targets)
CompositionHomogeneous or heterogeneous sensors
Energy availabilityConstrained (small sensors) or unconstrained (large ones)
CommunicationNetworkingWired on occasion, wireless more commonly
BandwidthHigh on occasion, low more typically
Processing architectureWhere the data is processedCentralised, distributed in the network, or hybrid

The same text book notes that passive sensors "tend to be low-energy devices" while active sensors such as radar and sonar "tend to be high-energy systems", which is the energy point made above in its own words.

Worked example: classifying four networks

Vineyard irrigationVolcano, ReventadorWildlife collarsCollection by a passing vehicle
Sensors usedSoil moisture (capacitive, contact), temperatureSeismometers, infrasonic microphonesGPS and accelerometersPollution sensors in fixed boxes
Passive or activePassivePassivePassive (GPS receives only)Passive
Data deliveryContinuousEvent-driven, then bulk dataContinuous, logged, sent when in rangeObserver-initiated as the vehicle passes
DynamicsStaticStaticMobile sensorsMobile observer
StructureMulti-hop meshMulti-hop tree to a gatewayOpportunistic, contact to contactSingle hop to the vehicle
CompositionHomogeneous, one mains sinkHomogeneous nodes, a GPS-equipped rootHomogeneousHomogeneous
DeploymentPlannedPlannedCarried by the animalsPlanned

Each column needs a different protocol. The vineyard wants a stable schedule and long sleeps; the volcano wants to move bursts of data reliably and time them exactly; the collars must store data until two animals, or an animal and a base station, come within range; and the roadside boxes must wake up and hand over their data in the few seconds a vehicle is in range.

Distinctions

ContinuousEvent-drivenObserver-initiatedHybrid
Who starts a reportA timerThe phenomenonThe observerAny of the three
TrafficSteadyRare bursts, many at onceOn demandMixed
Main riskWasted energy on uninteresting dataCollisions and delay at the critical momentDelay of the query and the replyServing all patterns at once
munotes.in61

Sensor Taxonomy

Passive sensorActive sensor
Emits energyNoYes, a pulse or beam
ExamplesThermometer, microphone, passive infraredUltrasonic ranger, radar
EnergyLowCan equal the radio's
Mobile observerMobile sensorsMobile phenomena
What movesThe collectorThe nodesThe thing being watched
What changesWhere the data must be deliveredThe paths between nodesWhich nodes should report

What it does not mean

"Sensor taxonomy" is not only a list of sensor types. An answer that stops at temperature, pressure and humidity has done half the job; the classification of networks is the half with the design consequences.

Passive does not mean unpowered here. In this subject a passive sensor is one that does not emit; it may still need a supply.

Event-driven does not mean light traffic. Most of the time it is silent, but at an event many sensors report at once, which is the hardest moment for the MAC protocol.

A mobile phenomenon does not make the network mobile. The sensors can all be fixed while the target moves; what changes is which of them report.

Quick revision

  • Two classifications: of sensors and of networks.
  • Sensors: what they measure; passive (only observe) or active (emit and measure the return); omnidirectional or narrow-beam; analogue or digital; contact or non-contact; operating principle. Karl and Willig: passive omnidirectional, passive narrow-beam, active.
  • Networks (Tilak and colleagues, 2002): participants sensor, observer, phenomenon; delivery continuous, event-driven, observer-initiated (request-reply), hybrid; dynamics static, or mobile observer, mobile sensors, mobile phenomena; communication application (cooperative or not) and infrastructure.
  • Also: flat or hierarchical; single or multi-hop (Sohraby and colleagues: mesh with dynamic routing, or star with static routing); homogeneous or heterogeneous; planned or random.

Test yourself

1. Explain the data delivery models of sensor networks. Continuous: sensors send at a set rate. Event-driven: sensors report only when an event of interest occurs. Observer-initiated (request-reply): sensors report only when the observer asks. Hybrid: the three coexist in one network.

2. Distinguish passive and active sensors, with examples. A passive sensor only observes and emits nothing (thermometer, microphone, passive infrared detector); an active sensor emits energy and measures what returns (ultrasonic range finder, radar) and so uses more energy.

3. What are the three kinds of mobility in the Tilak taxonomy? Mobile observer (the collector moves), mobile sensors (the nodes move, breaking paths) and mobile phenomena (the target moves, changing which nodes should report).

4. What is infrastructure communication, and why should it be kept small? The traffic needed to configure, maintain and optimise the network, such as discovering neighbours and building or repairing paths. It carries no sensed data, so it is overhead that costs energy, though a little of it can reduce application traffic overall.

munotes.in62

Sensor Taxonomy

5. Classify a network of GPS collars on elephants. Passive sensors (GPS receivers, accelerometers); continuous logging with delivery when in range; dynamic, with mobile sensors; homogeneous nodes; deployment carried by the animals; paths formed opportunistically when two collars or a collar and a base station meet.

Contents This chapter on its own page

munotes.in63

Chapter Eleven

The Operating Environment and the Design Factors

Syllabus topic Module 1, "Introduction and Overview of WSNs: WSN operating environment"

In one line

A sensor node's operating environment is wherever it is left, usually outdoors or inside machinery, unattended, exposed to weather, interference and damage, and every design decision must survive it.

In the wording a student can write in an examination: sensor nodes are deployed "very close or directly inside the phenomenon to be observed", so they usually work unattended in remote or harsh places: the interior of large machinery, the bottom of an ocean, a biologically or chemically contaminated field, a battlefield, a home or a large building, or on a human body. The survey lists the design factors such an environment imposes: fault tolerance, scalability, production costs, hardware constraints, sensor network topology, the environment itself, transmission media and power consumption.

Why the environment comes before the protocol

Because it decides what can go wrong. A network on a factory floor has power nearby, a stable temperature and a technician on call. The same network in a forest has rain, heat, frost, animals, falling branches and nobody within a day's walk. Protocols that work in the first can fail in the second, and a design that starts from the protocol instead of the environment finds this out in the field.

Where nodes are deployed

Both of MU's text book sources give the same kind of list. The survey names the interior of large machinery, the bottom of an ocean, contaminated fields, battlefields "beyond the enemy lines", and homes and large buildings. Sohraby, Minoli and Znati add open spaces, commercial buildings and "in or on a human body". What these places share is that people are not there, cannot easily go there, or should not.

The environment attacks a node in four ways, and each has a design answer.

What the environment doesExamplesDesign answer
Physical conditionsHeat, frost, rain, humidity, dust, salt spray, vibrationSealed enclosures, wide-temperature batteries (the lithium L91 works from -40 to 60 degrees C), conformal coatings
Physical damage and interferenceAnimals chewing, being trodden on, washed away, stolen or tampered withRedundant nodes, tamper-resistant casings, security that survives capture
The radio environmentObstacles, ground reflections, other radios in the same band, multipathShort hops, robust modulation, link estimation, retransmission
No attendanceNobody to reset, repair, recharge or reconfigureSelf-organisation, self-healing routes, remote reprogramming, long battery life

The radio point deserves one more sentence, because it surprises students. Nodes usually sit on or near the ground, and the survey notes that the power needed to reach a distance d grows as d to the power n, with n "closer to four for low-lying" antennas and near-ground channels. The ground itself shortens a sensor node's range, and [Radio Technology in WSNs: The Sensor Radio and Its Link Budget] shows by how much.

munotes.in64

The Operating Environment and the Design Factors

The design factors

The survey calls these the factors that "serve as a guideline to design a protocol or an algorithm for sensor networks" and that can be used to compare one scheme with another.

1. Fault tolerance

Some nodes will fail: they run out of power, suffer physical damage, or are blocked by interference. Fault tolerance is the ability to keep the network's functions running without interruption despite node failures. The survey gives a model for one node's reliability, the probability that node k has not failed within a time t:

R(t) = e^(-λt)

where λ (lambda) is that node's failure rate, the average number of failures per unit time. The model is the Poisson assumption: failures happen at random at a steady average rate. How much fault tolerance a network needs depends on its environment: a house needs little, a battlefield a great deal.

2. Scalability

A deployment may have hundreds or thousands of nodes, and in extreme cases millions. Protocols must work with that many, and must make use of the resulting density. The survey gives the density as

μ(R) = N × π × R² / A

where N nodes are scattered over an area A and R is the radio range. μ(R) is the number of nodes within radio range of each node: its expected number of neighbours. Density can range from a few nodes to a few hundred in a region less than 10 metres across.

3. Production costs

Because there are so many nodes, the cost of one decides whether the network is worth building at all. In the survey's words, if the network costs more than deploying traditional sensors, "the sensor network is not cost-justified". Cheap nodes mean small batteries, simple radios and little memory.

4. Hardware constraints

A node has four basic units (sensing, processing, transceiver, power) and possibly a location finding system, a power generator and a mobilizer, as [Inside a Sensor Node: The Five Units] explains. The survey adds that all of it may need to fit in a matchbox or even under a cubic centimetre, and that nodes must consume extremely little power, operate at high density, cost little, be dispensable and autonomous, operate unattended and adapt to their environment.

5. Sensor network topology

With hundreds to thousands of nodes deployed "within tens of feet of each other" and densities up to 20 nodes in a cubic metre, keeping the topology working needs care in three phases:

  1. Pre-deployment and deployment: nodes are thrown in as a mass or placed one by one, dropped from a plane, delivered in an artillery shell, rocket or missile, or placed by a person or a robot.
  2. Post-deployment: the topology changes as nodes move, as links are lost to jamming, noise or moving obstacles, as energy runs down, as nodes malfunction and as the task changes.
  3. Redeployment of additional nodes: new nodes are added at any time to replace failed ones or because the task has changed.
munotes.in65

The Operating Environment and the Design Factors

6. Environment

Everything in the first half of this chapter: nodes work unattended in remote places, close to or inside the phenomenon.

7. Transmission media

Links can be formed by radio, infrared or optical media. To allow operation anywhere, the medium must be available worldwide, which in practice means the licence-free industrial, scientific and medical (ISM) bands. Infrared is licence-free, robust against electrical interference and cheap, and optical links were used by the Smart Dust mote, but both need a clear line of sight. Radio needs none, which is why almost every node uses it.

8. Power consumption

The survey's example power source is small, under 0.5 ampere-hours at 1.2 volts, and in many applications it can never be replaced, so a node's lifetime depends strongly on its battery. Because each node is both a data originator and a data router, the failure of a few can force rerouting and reorganisation of the whole network. The survey divides a node's consumption into three domains: sensing, communication and data processing, and [How Long a Node Lasts: The Energy Budget Worked Out] puts real numbers on them.

Worked example: the two formulas, by hand and by program

Reliability

Take a node with a failure rate of λ = 0.2 a year: on average, one failure in five node-years (an assumption for the exercise). Its reliability after one year is

R(1) = e^(-0.2 × 1) = e^(-0.2)

which is about 0.8187. So there is roughly an 18 per cent chance a given node has failed within a year.

Now suppose every spot in the field is watched by two nodes, and that nodes fail independently of each other. The spot is lost only if both have failed, which has probability (1 - 0.8187) squared, about 0.0329. So the spot is still watched with probability about 0.9671. With three nodes it is about 0.9940. Redundancy turns an unreliable node into a reliable network.

Density

Scatter N = 200 nodes over a field 100 metres by 100 metres, so A is 10,000 square metres, with a radio range R of 20 metres:

μ(20) = 200 × π × 20² / 10,000

which is about 25.1 neighbours per node. That is plenty for routing and for redundancy, and it also means 25 radios contend for the channel around every node, which the MAC protocol must handle.

munotes.in66

The Operating Environment and the Design Factors

The program below computes both formulas over several cases, so the pattern is visible rather than one number.

import math

# Fault tolerance: the survey's R(t) = e^(-lambda t), the chance that one node
# has not failed by time t, for a failure rate lambda (failures per year).
LAMBDA = 0.2   # assumed: on average one failure per five node-years

print("years   one node   a spot watched by 2   by 3")
for t in (1, 2, 3, 5):
    r = math.exp(-LAMBDA * t)
    two = 1 - (1 - r) ** 2      # at least one of two still working
    three = 1 - (1 - r) ** 3
    print("%5d   %8.4f   %19.4f   %4.4f" % (t, r, two, three))

# Density: the survey's mu(R) = N * pi * R^2 / A, the number of nodes
# within radio range R of each node, for N nodes scattered over area A.
N, A = 200, 100 * 100          # 200 nodes in a 100 m by 100 m field
print()
print("range R (m)   neighbours mu(R)")
for R in (10, 15, 20, 30):
    mu = N * math.pi * R ** 2 / A
    print("%11d   %16.1f" % (R, mu))
years   one node   a spot watched by 2   by 3
    1     0.8187                0.9671   0.9940
    2     0.6703                0.8913   0.9642
    3     0.5488                0.7964   0.9082
    5     0.3679                0.6004   0.7474

range R (m)   neighbours mu(R)
         10                6.3
         15               14.1
         20               25.1
         30               56.5

Read the reliability rows downwards. After five years a single node has only about a 37 per cent chance of still working, but a spot watched by three nodes is still covered with about 75 per cent probability. Redundancy buys time, not immortality.

Read the density rows. μ grows with the square of the range: doubling R from 10 to 20 metres multiplies the neighbours by four, from about 6 to about 25. Turning a radio's power up to gain range therefore floods every node with neighbours to contend with, which is one reason [Energy Efficiency in Ad Hoc Networks: Where the Energy Goes] prefers to turn power down.

Distinctions

The operating environmentThe design factors
IsWhere the node is and what happens to it thereWhat the designer must take into account because of it
ExamplesA forest, a machine, the sea bed, a bodyFault tolerance, scalability, cost, hardware, topology, environment, media, power
Asked asDescribe the environment in which WSNs operateList and explain the factors that influence WSN design
Reliability R(t)Density μ(R)
MeasuresThe chance one node survives to time tThe expected number of neighbours of a node
Depends onThe failure rate λ and the time tThe number of nodes N, the range R and the area A
Grows withSmaller λ, shorter tMore nodes, longer range (as R squared)
munotes.in67

The Operating Environment and the Design Factors

What it does not mean

R(t) is not the network's reliability. It is one node's. The network's depends on redundancy, and the sum above treats failures as independent, which a flood or a fire that destroys many nodes at once does not respect.

A higher density is not always better. It gives redundancy and routes, and it also gives collisions and more overhearing. The best density is the lowest that meets coverage and connectivity, which [Deployment and Coverage: Random Against Grid] examines.

"Unattended" does not mean "untouched". Unattended nodes are exactly the ones an animal or an attacker can reach without anyone noticing.

The survey's numbers are its own time's. Its target of a node costing well under one US dollar and its 0.5 ampere-hour power source describe research goals of 2002, not the price or battery of a node you can buy.

Quick revision

  • Environment: unattended, remote or harsh: machinery, ocean floor, contaminated fields, battlefields, buildings, homes, the body. Threats: physical conditions, damage, the radio environment, no attendance.
  • Near the ground the path loss exponent is close to 4.
  • Eight design factors (Akyildiz and colleagues 2002): fault tolerance, scalability, production costs, hardware constraints, topology, environment, transmission media, power consumption.
  • R(t) = e^(-λt): one node's reliability. λ = 0.2 a year: R(1) about 0.8187; with two independent nodes a spot survives with about 0.9671, with three about 0.9940.
  • μ(R) = N × π × R² / A: neighbours per node. 200 nodes, 100 m by 100 m, R = 20 m: about 25.1; it grows as R squared.
  • Topology in three phases: deployment, post-deployment, redeployment. Media: radio, infrared, optical (the last two need line of sight). Power: sensing, communication, processing.

Test yourself

1. Describe the environment in which sensor networks operate. Nodes are placed close to or inside the phenomenon, so they usually work unattended in remote or harsh places: inside machinery, on the ocean floor, in contaminated fields, on battlefields, in buildings and homes, on the body. They face weather, damage, interference and a crowded radio band, with nobody to repair them.

2. List and explain the factors influencing sensor network design. Fault tolerance (keep working despite failures, modelled by R(t) = e^(-λt)); scalability (hundreds to millions of nodes, density μ(R) = N π R² / A); production costs (each node must be cheap); hardware constraints (four basic units in a tiny, frugal package); topology (deployment, post-deployment and redeployment phases); environment (unattended, harsh); transmission media (radio, infrared, optical, available worldwide); power consumption (sensing, communication, processing).

3. A node has a failure rate of 0.2 a year. What is its reliability after one year, and what is the chance a spot watched by two such nodes is still covered? R(1) = e^(-0.2), about 0.8187. With independent failures the spot is lost with probability (1 - 0.8187)², about 0.0329, so it is covered with about 0.9671.

munotes.in68

The Operating Environment and the Design Factors

4. 200 nodes are scattered over 100 m by 100 m with a radio range of 20 m. How many neighbours does each have on average? μ = 200 × π × 400 / 10,000, about 25.1.

5. Name the three phases of topology maintenance. Pre-deployment and deployment (nodes thrown in or placed), post-deployment (topology changes from movement, lost links, energy, malfunction, task changes), and redeployment of additional nodes.

Contents This chapter on its own page

munotes.in69

Chapter Thirteen

The Wireless Technologies a Sensor Network Can Use

Syllabus topic Module 1, "Introduction and Overview of WSNs: Radio technology in WSNs"

In one line

A sensor network can be built on short-range personal-area radios (IEEE 802.15.4 and what runs on it, Bluetooth Low Energy), on Wi-Fi, on the new low-power wide-area networks (LoRaWAN, Sigfox, NB-IoT, Wi-SUN) or on the cellular network, and the choice trades range and data rate against energy and infrastructure.

In the wording a student can write in an examination: the wireless technologies available to a WSN include IEEE 802.15.4 (low-rate personal area networks, 20 to 250 kb/s, the basis of Zigbee and 6LoWPAN); Bluetooth Low Energy (coin-cell devices, star topology around a central device); IEEE 802.11 Wi-Fi (high rate, high power); low-power wide-area networks such as LoRaWAN, Sigfox and NB-IoT (kilometres of range at very low data rates and duty cycles); and cellular data such as GPRS. The designer chooses by range, data rate, energy, topology, infrastructure needed and cost.

Why there is more than one answer

Because the four things a designer wants pull against each other. Long range needs either high power or a very low data rate. A high data rate needs more power or a short range. Needing no infrastructure means the nodes must relay for each other, which costs energy. So each technology picks a corner of the trade-off, and the application decides which corner it can live in.

The text book's comparison

MU's first text book compares four candidate lower-layer technologies in one table. It is worth knowing as printed, because it shows where 802.15.4 sits.

GPRS/GSM, 1xRTT/CDMAIEEE 802.11b/gIEEE 802.15.1IEEE 802.15.4
Market name2.5G/3GWi-FiBluetoothZigBee
Network targetWAN/MANWLAN and hotspotPAN and DAN (desk area network)WSN
Application focusWide area voice and dataEnterprise applications (data and VoIP)Cable replacementMonitoring and control
Bandwidth (Mbps)0.064 to 0.128+11 to 540.70.020 to 0.25
Transmission range (ft)3000+1 to 300+1 to 30+1 to 300+
Design factorsReach and transmission qualityEnterprise support, scalability and costCost, ease of useReliability, power and cost

Read the last column against the others. 802.15.4 has the lowest data rate of the four and a range like Wi-Fi's, and it is the only one designed with power as a design factor. That is what "designed for sensor networks" means.

IEEE 802.15.4, and what is built on it

IEEE 802.15.4 defines only the two lowest layers, the physical layer and the MAC, for low-rate wireless personal area networks (LR-WPANs). It operates in three unlicensed bands (868 to 868.6 MHz, 902 to 928 MHz and 2400 to 2483.5 MHz, the last worldwide) at 20, 40 and 250 kb/s in its original form, and its frames carry at most 127 bytes. MU names it as the routing module's case study, and five chapters from [IEEE 802.15.4: The Standard, Its Devices and Its Topologies] take it apart.

munotes.in76

The Wireless Technologies a Sensor Network Can Use

Because it defines only two layers, other standards supply the rest:

  • Zigbee adds a network layer and an application framework. Its network layer supports star, tree and mesh topologies; the device that starts and controls the network is the Zigbee coordinator, which the specification defines as "an IEEE 802.15.4 PAN coordinator"; a Zigbee router is "capable of routing messages between devices"; a Zigbee end device is neither, and only sends and receives its own traffic. Zigbee's routing is taught in [Built on 802.15.4: Zigbee Routing, Security and the Later Amendments].
  • 6LoWPAN carries IPv6 over 802.15.4 frames, so that sensor nodes can be ordinary Internet hosts. It is the subject of [Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP].

Strengths: designed for battery-powered nodes; supports multi-hop mesh networks through the layers above it; worldwide 2.4 GHz band; cheap radios. Limits: tens of metres per hop, so large areas need many hops; the 2.4 GHz band is shared with Wi-Fi and Bluetooth.

Bluetooth Low Energy

Bluetooth Low Energy (Bluetooth LE, also branded Bluetooth Smart) was introduced in Bluetooth 4.0 and extended in 4.1 and later versions. RFC 7668 describes its target plainly: devices "that operate with very low-capacity (e.g., coin cell) batteries or minimalistic power sources". A device in the central role (typically a phone or a router) manages connections to several peripherals (the sensors), giving a star topology. Its packets are small: the default MTU of its link-layer channels is 27 octets, 23 of them available to the upper layers.

Strengths: in almost every phone, so a person's phone can be the gateway; very low energy per connection; good for wearable and personal sensors, such as the heart-rate monitor RFC 7668 gives as its example. Limits: short range; the star topology means a node must be within reach of a central device; not designed for large multi-hop fields.

Wi-Fi

IEEE 802.11 (Wi-Fi) gives tens of megabits per second, which the text book's table puts at 11 to 54 for the 802.11b and g generations. It is designed for laptops and phones with large batteries or mains power, and its receivers are built to be on whenever the device is in use. Strengths: existing infrastructure everywhere, high data rates, IP built in. Limits: far more power than a sensor node can sustain on small batteries; access points needed. It suits sensors with mains power (cameras, gateways) and the gateway's own link to the Internet.

The low-power wide-area networks

A newer family aims at a different corner of the trade-off. RFC 8376, the IETF's overview, puts their shared goal as "supporting large numbers of very low-cost, low-throughput devices with very low power consumption, so that even battery-powered devices can be deployed for years", and says they accept "severe bandwidth and duty cycle constraints" in return for "multiple-kilometer radio links". LoRaWAN, Sigfox and NB-IoT are star networks: every device talks directly to a base station or gateway, with no relaying.

munotes.in77

The Wireless Technologies a Sensor Network Can Use

LoRaWAN uses ISM bands, for example 433 MHz and 868 MHz in the European Union and 915 MHz in the Americas. In the EU 868 MHz band its three default channels are at 868.1, 868.2 and 868.3 MHz, and its data rate runs from 250 bit/s with frames of 59 octets up to 50,000 bit/s with frames of 250 octets. A device's transmission may be received by several gateways at once, and the network adjusts each device's data rate and power by an adaptive data rate scheme. After each transmission a basic (Class A) device opens two short receive windows, and only then may it transmit again.

Sigfox uses ultra narrow band transmission: channels of 100 or 600 hertz, 100 or 600 baud, modulated by DBPSK. Its devices send "only a few bytes per day, week, or month", which RFC 8376 says allows them to last "up to 10-15 years" on one battery.

NB-IoT (Narrowband IoT) runs in licensed mobile-operator spectrum, using a narrow band of 180 kHz inside, beside or outside an LTE carrier. Its peak rates are 60 kbit/s up and 30 kbit/s down, and its design targets include a module cost of less than 5 US dollars, coverage of 164 dB maximum coupling loss, battery life of over 10 years and about 55,000 devices per cell.

Wi-SUN FAN (Field Area Network) is the exception to the star pattern: RFC 8376 describes it as "an IPv6 wireless mesh network" using frequency hopping, typically 2 to 3 km in line of sight and up to 300 kbit/s.

Worked example: what a 1 per cent duty cycle allows

Some ISM bands limit how much of the time a device may transmit, a limit RFC 8376 says is "usually expressed as a percentage of time per hour", and its LoRaWAN table gives 1 per cent as the tightest such limit. Take a LoRaWAN device in the EU 868 MHz band at its slowest and longest-reaching setting, 250 bit/s, sending a full 59-octet frame.

  • Time on air for one frame: 59 × 8 / 250 = 1.888 seconds.
  • Transmitting time allowed in an hour: 3,600 × 0.01 = 36 seconds.
  • Frames allowed in an hour: 36 / 1.888, about 19.

So even this long-range network can carry about nineteen small messages an hour from each device at its longest range, which is plenty for a meter and hopeless for anything that streams.

munotes.in78

The Wireless Technologies a Sensor Network Can Use

Cellular data

A node can also use the mobile phone network directly, through GPRS or later data services, taught in [New Data Services: HSCSD, GPRS and EDGE]. It needs a SIM and a subscription for each node and draws considerable power, so it is usually used by the gateway rather than by every node: the sensor field talks to the gateway over 802.15.4, and the gateway talks to the task manager over the cellular network.

Worked example: choosing for three applications

ApplicationNeedsChoice, and why
Vineyard irrigation, 40 nodes over a few hectaresLow rate, a season on batteries, no mains in the field802.15.4 mesh with a gateway at the pump house, or LoRaWAN if one gateway can reach every node
Water meters across a cityA few bytes a day, ten years on a battery, no maintenance, operator-run networkNB-IoT or another LPWAN: no gateway to install, and one reading a day fits easily in the duty cycle
A patient's heart-rate sensorContinuous readings to a nearby phone, tiny batteryBluetooth Low Energy, with the phone as gateway

The same reasoning chooses for any application: first the range and topology (can every node reach a gateway directly, or must nodes relay?), then the data rate and duty cycle (does the traffic fit?), then energy and cost.

A last classification: how constrained is the node?

RFC 7228 names three classes of constrained device, a useful vocabulary when choosing a stack.

ClassData size (RAM)Code size (flash)
Class 0Much less than 10 KiBMuch less than 100 KiB
Class 1About 10 KiBAbout 100 KiB
Class 2About 50 KiBAbout 250 KiB

The MSP430F1611 of [Inside a Sensor Node: The Five Units], with 10 kB of RAM and 48 kB of flash, sits between classes 0 and 1: too small for a full Internet protocol stack, which is why sensor networks needed stacks of their own.

Distinctions

802.15.4-basedBluetooth LEWi-FiLPWAN (LoRaWAN, Sigfox, NB-IoT)
Range per linkTens of metresShortTens of metresKilometres
Data rate20 to 250 kb/sLow, small packetsMegabits per secondHundreds of bits to tens of kilobits a second
TopologyStar, tree or meshStar around a centralStar around an access pointStar around a gateway or base station (Wi-SUN: mesh)
Relaying by nodesYes, in meshNoNoNo
EnergyLowVery lowHighVery low, at low duty cycle
SpectrumUnlicensedUnlicensedUnlicensedUnlicensed, except NB-IoT (licensed)

What it does not mean

Zigbee and 802.15.4 are not the same thing. 802.15.4 is the radio and MAC; Zigbee is one of several stacks that run on top of it.

munotes.in79

The Wireless Technologies a Sensor Network Can Use

A long-range LPWAN is not a replacement for a mesh in every case. It trades data rate and duty cycle for range, and a node that must report every few seconds does not fit.

Wi-Fi is not wrong for sensors; it is wrong for battery sensors. A mains-powered camera or a gateway uses it well.

"Low power" is not a property of a radio alone. Bluetooth LE or 802.15.4 left listening all the time will empty a coin cell quickly. The MAC protocol's duty cycle decides the energy as much as the radio does.

Quick revision

  • Text book Table 1.3: GPRS 0.064 to 0.128+ Mbps, 3000+ ft; Wi-Fi 11 to 54 Mbps, 1 to 300+ ft; Bluetooth 0.7 Mbps, 1 to 30+ ft; 802.15.4/ZigBee 0.020 to 0.25 Mbps, 1 to 300+ ft, designed for reliability, power and cost.
  • 802.15.4: PHY and MAC only, three bands, 20/40/250 kb/s, 127-byte frames; Zigbee (star, tree, mesh; coordinator = 802.15.4 PAN coordinator) and 6LoWPAN above it.
  • Bluetooth LE: from Bluetooth 4.0, coin cells, central and peripherals, star, 27-octet MTU.
  • LPWAN: star (Wi-SUN mesh), kilometres, very low rate and duty cycle. LoRaWAN EU 868: channels 868.1/868.2/868.3 MHz, 250 bit/s to 50,000 bit/s; Sigfox ultra narrow band, 100 or 600 baud, DBPSK; NB-IoT licensed, 180 kHz, 60/30 kbit/s peak, 164 dB, over 10 years, about 55,000 devices a cell; Wi-SUN 2 to 3 km, 300 kbit/s.
  • 1 per cent duty cycle at 250 bit/s: 1.888 s a frame, 36 s an hour, about 19 frames an hour.
  • RFC 7228: Class 0 (much less than 10 KiB RAM), Class 1 (about 10 KiB), Class 2 (about 50 KiB).

Test yourself

1. Compare 802.15.4, Bluetooth, Wi-Fi and GPRS as technologies for a sensor network. From the text book's table: GPRS reaches kilometres at 0.064 to 0.128+ Mbps for wide-area voice and data; Wi-Fi gives 11 to 54 Mbps over up to 300+ feet for enterprise data; Bluetooth gives 0.7 Mbps over up to 30+ feet for cable replacement; 802.15.4 gives 0.020 to 0.25 Mbps over up to 300+ feet and is the only one designed around power, reliability and cost for monitoring and control.

2. What is the relationship between IEEE 802.15.4 and Zigbee? 802.15.4 defines the physical and MAC layers of a low-rate personal area network. Zigbee defines a network layer and application framework on top of it, supporting star, tree and mesh topologies; the Zigbee coordinator is an 802.15.4 PAN coordinator.

3. What do low-power wide-area networks trade, and for what? They accept very low data rates and strict duty cycles in exchange for multi-kilometre links and years of battery life, and they use a star topology with no relaying by the nodes.

munotes.in80

The Wireless Technologies a Sensor Network Can Use

4. A LoRaWAN device sends 59-octet frames at 250 bit/s under a 1 per cent duty cycle. How many frames can it send an hour? One frame takes 59 × 8 / 250 = 1.888 s; the device may transmit for 3,600 × 0.01 = 36 s an hour; 36 / 1.888 gives about 19 frames.

5. Which technology would you choose for a wearable heart-rate sensor, and why? Bluetooth Low Energy: it is designed for coin-cell devices, it is in every phone, so the phone serves as the gateway, and a star topology around the phone is all a single wearer needs.

Contents This chapter on its own page

munotes.in81

Chapter Fourteen

Network Architecture: Sources, Sinks, Hops and Mobility

Syllabus topic Module 1, "Introduction and Overview of WSNs: Network architecture and optimization"

In one line

A sensor network's architecture is described by where data comes from (its sources), where it must go (its sinks), how many hops lie between them, and what moves: the nodes, the sink or the thing being observed.

In the wording a student can write in an examination: the basic sensor network scenarios are defined by (1) the types of sources and sinks: a source is any node that provides data, and a sink is where the data is required, which may be a node inside the network, a device outside it such as a user's handheld, or a gateway to a larger network; (2) single-hop or multi-hop communication between them; (3) multiple sources and multiple sinks; and (4) three types of mobility: node mobility, sink mobility and event mobility.

Why describe a network this way

Because the same hardware behaves as completely different networks depending on these four answers. A protocol that is efficient with one fixed sink and fixed nodes can collapse when the sink is a firefighter walking through a building. Before choosing any protocol, the designer writes down the scenario, and this is the vocabulary for doing it.

Sources and sinks

A source is any entity in the network that can provide data: usually a sensor node, sometimes an actuator reporting that it has acted. A sink is the entity where the information is required. Karl and Willig note that in most applications "there is a clear difference between sources of data" and sinks, that there are usually, though not always, more sources than sinks, and that "the sink is oblivious or not interested in the identity of the sources; the data itself is much more important", which is the data-centric idea of [What a Wireless Sensor Network Is] in the book's own words.

A sink can be one of three things.

  1. Part of the sensor network itself: another sensor node, or an actuator that must react (the valve that opens when the soil is dry).
  2. A device outside the network that talks to it directly, in Karl and Willig's example "the firefighter's PDA communicating with a WSN". Today it would be a phone or a tablet.
  3. A gateway to a larger network, through which the data reaches the Internet and the task manager, as in [The Architectural Elements of a Sensor Network].

Single hop or multiple hops

If every source can reach its sink directly, the network is single-hop: a star around the sink. It is simple, and nodes never relay. But radio range is limited, and near the ground it is tens of metres ([Radio Technology in WSNs: The Sensor Radio and Its Link Budget]), so in most fields many sources cannot reach the sink at all.

munotes.in82

Network Architecture: Sources, Sinks, Hops and Mobility

The answer is multi-hop communication: nodes relay for each other. Karl and Willig give the reason exactly: communication over long distances "is only possible using prohibitively high transmission power", and "the use of intermediate nodes as relays can reduce the total required power". Multi-hop also works around obstacles: two nodes on either side of a wall may not reach each other directly but may both reach a third.

Note the careful wording: relays can reduce the power. Whether they do depends on the distances and on the radio, and the next chapter, [Single Hop or Multiple Hops: The Energy Argument Worked Out], works out when they do not. A multi-hop network also pays in delay (every hop adds some) and in load on the relays (a node near the sink forwards everybody's data).

Multiple sources and multiple sinks

Real networks have many of each, and the pattern of traffic between them decides what the routing must do.

PatternWho talks to whomExample
Many sources, one sinkAll readings converge on one collection point: convergecastThe vineyard reporting to the pump house
One source, many sinksOne node's data is wanted in several placesA fire alarm sent to every firefighter's handheld
Many sources, many sinksEach sink needs data from some sourcesSeveral users querying different parts of a building

With several sinks, a source may report to any one of them (the nearest, to save energy) or to all of them (when every sink needs the data). Several sinks also remove the single point of failure that one sink is, and share out the relaying load that concentrates around it.

The three kinds of mobility

A network's structure changes when something moves, and three different things can move.

1. Node mobility. The sensor nodes themselves move: collars on animals, nodes on vehicles or robots, sensors floating in a river. Karl and Willig add that nodes may also move deliberately, after deployment, to positions where they can do their sensing task better. Routes break as nodes move and must be rebuilt, which is the same problem a mobile ad hoc network faces ([Ad Hoc Networks: MANETs, and How a Sensor Network Differs]).

2. Sink mobility. The sink moves while the nodes stay put: a firefighter walking through a building, a vehicle driving past roadside nodes. Karl and Willig note that in principle this is ordinary node mobility, but it can cause difficulties for protocols that are efficient only in fully static scenarios, because the routes all point towards where the sink was.

3. Event mobility. The phenomenon moves: an intruder, a vehicle, a fire front. The nodes and the sink can all be fixed while the set of nodes that should be reporting changes. In Karl and Willig's words, in tracking applications it is the network's explicit task "to ensure that some form of activity happens in nodes that surround the phenomenon under observation", and ideally what was learned at one place should be available at the next.

munotes.in83

Network Architecture: Sources, Sinks, Hops and Mobility

Correlated mobility. In any of these, a group can move together: nodes carried by a storm, a river or some other fluid, or a group of people walking together.

Worked example: a bus as a mobile sink

A town council places pollution nodes on 50 lamp-posts along a bus route and lets the bus collect their readings as it passes, instead of building a multi-hop network or a gateway on every post. This is sink mobility: the nodes are fixed, the sink drives by.

The readings to collect. Each node reads every 15 minutes, 96 readings a day, and each reading is 12 bytes, so a node holds 96 × 12 = 1,152 bytes a day.

The time the bus is in range. The node's radio reaches 30 metres, so the bus is in range over 60 metres of road. At 36 km/h the bus travels 36,000 / 3,600 = 10 metres a second, so it is in range for 60 / 10 = 6 seconds.

What 6 seconds can carry. At 250 kbps the radio could move 6 × 250,000 / 8 = 187,500 bytes in that time, if the whole contact were usable. Even allowing for the time to find the bus, for packet overhead and for retransmissions, a day's 1,152 bytes fit comfortably.

What the scenario demands of the protocol. The nodes must be awake, or wake quickly, when the bus comes, which is a problem for a node that sleeps 99 per cent of the time: either the bus announces itself with a signal the node's short listening windows can catch, or the node learns the timetable. And if the bus breaks down, the data waits: a mobile sink trades a network's infrastructure for delay.

Distinctions

Node mobilitySink mobilityEvent mobility
What movesThe sensor nodesThe collectorThe phenomenon
What breaks or changesRoutes between nodesRoutes towards the sinkWhich nodes should report
ExampleAnimal collarsA bus or a firefighterA vehicle being tracked
Single-hopMulti-hop
ReachOne radio rangeMany
RelayingNoneEvery node may relay
DelayOne hopAdds up over hops
EnergyCan be high for distant nodesCan be lower, not always (next chapter)
Weak pointNodes out of rangeRelays near the sink, which carry the most traffic
Sink inside the networkSink outside, talking directlyGateway sink
ExampleAn actuator nodeA firefighter's handheldA gateway to the Internet
ResourcesLike a nodeUsually moreMost, often mains powered
munotes.in84

Network Architecture: Sources, Sinks, Hops and Mobility

What it does not mean

A sink is not always a gateway. It may be an actuator in the field or a handheld that never connects to the Internet.

Multi-hop does not always save energy. It reduces the transmit power per hop; it adds receptions and more transmissions. When the radio's electronics dominate its energy, one long hop can be cheaper, as the next chapter computes.

Event mobility does not require any node to move. The fire moves; the sensors stay where they are.

A mobile sink is not free. It removes relaying and infrastructure and adds delay, plus the problem of waking sleeping nodes when it passes.

Quick revision

  • Scenarios (Karl and Willig, chapter 3): sources and sinks; single-hop or multi-hop; multiple sources and sinks; three kinds of mobility.
  • Source: provides data. Sink: where data is required: inside the network, an outside device (the firefighter's PDA), or a gateway.
  • The sink "is oblivious or not interested in the identity of the sources".
  • Multi-hop: needed because long hops need "prohibitively high transmission power"; relays can reduce total power, not always; costs delay and relay load.
  • Traffic patterns: many to one (convergecast), one to many, many to many; several sinks remove the single point of failure.
  • Mobility: node, sink, event; mobility can be correlated.
  • Bus sink: 1,152 bytes a day per node; 6 seconds in range at 36 km/h over 60 m; the radio could carry 187,500 bytes in that time.

Test yourself

1. Explain the three types of mobility in a sensor network, with an example of each. Node mobility: the sensors move, such as collars on animals, and routes between them break. Sink mobility: the collector moves, such as a bus passing roadside nodes, and routes pointing at the sink go stale. Event mobility: the phenomenon moves, such as a tracked vehicle, and the set of nodes that should report changes, even though no node moves.

2. What can a sink be? A node inside the network (another sensor or an actuator), a device outside that communicates with the network directly (a user's handheld), or a gateway to a larger network such as the Internet.

3. Why do sensor networks use multi-hop communication, and what does it cost? Because radio range is limited and reaching far in one hop needs prohibitively high power; relays can reduce the total power and get round obstacles. It costs delay at every hop and a heavy load on the relays near the sink, and it does not always save energy.

munotes.in85

Network Architecture: Sources, Sinks, Hops and Mobility

4. A bus at 36 km/h passes a node whose radio reaches 30 m. How long is the node in range? Over 60 m of road at 10 m/s: 6 seconds.

5. What is convergecast? The many-to-one traffic pattern in which readings from many sources converge on one sink, typical of data collection networks.

Contents This chapter on its own page

munotes.in86

Chapter Fifteen

Single Hop or Multiple Hops: The Energy Argument Worked Out

Syllabus topic Module 1, "Introduction and Overview of WSNs: Network architecture and optimization"

In one line

Several short hops need less transmit power than one long hop, but every extra hop adds a reception and the cost of the radio's electronics, so multi-hop saves energy only when the distances are long enough for the amplifier's energy to outweigh the electronics.

In the wording a student can write in an examination: with the first-order radio model, sending k bits a distance d costs E(elec) × k + ε(amp) × k × d², and receiving them costs E(elec) × k. For a node n hops of r metres from the base station, direct transmission costs k(E(elec) + ε(amp) n² r²), while minimum transmission energy (MTE) multi-hop routing costs k((2n - 1)E(elec) + ε(amp) n r²). Direct transmission is cheaper whenever n r² < 2 E(elec) / ε(amp). Multi-hop wins at long distances, especially where loss grows as d⁴, and even then there is an optimal number of hops; and in multi-hop the node nearest the base station spends by far the most energy.

The claim, and why it sounds right

Karl and Willig state the reason for multi-hop: long-distance communication needs "prohibitively high transmission power", and relays "can reduce the total required power". The arithmetic behind it is the path loss of [Radio Technology in WSNs: The Sensor Radio and Its Link Budget]. If the power needed grows as d², one hop of 100 metres needs power proportional to 100², which is 10,000, while two hops of 50 metres need 2 × 50², which is 5,000: half. With d⁴ the saving is bigger still.

What that argument leaves out is everything the radio spends that does not depend on distance. Every hop means one more transmission and one more reception, and each of those runs the radio's electronics (oscillator, mixers, filters, the digital processing) for the length of the message. [The Radio, the Sensors and the Power Supply of a Node] showed that on a real radio these electronics, not the power amplifier, take most of the current. So the question is not whether short hops need less transmit power. It is whether the power saved is more than the electronics spent.

The first-order radio model

The LEACH paper of 2000 wrote the radio down in two equations.

E(transmit, k, d) = E(elec) × k + ε(amp) × k × d²

E(receive, k) = E(elec) × k

Here k is the number of bits, d the distance, E(elec) the energy per bit to run the transmitter or receiver electronics, and ε(amp) the energy per bit per square metre in the transmit amplifier. The paper used E(elec) = 50 nJ/bit and ε(amp) = 100 pJ/bit/m², and noted that with these values "receiving a message is not a low cost operation".

munotes.in87

Single Hop or Multiple Hops: The Energy Argument Worked Out

The linear network

A line of n nodes r metres apart, sending to the base station directly or hop by hop

Figure 15.1 One long hop, or n short hops with n - 1 receptions

Put n nodes on a line, r metres apart, with the base station just beyond node 1. Take the node farthest out, n × r metres from the base station, and one message of k bits.

Direct: one transmission over n × r metres, and no reception except at the base station:

E(direct) = E(elec) × k + ε(amp) × k × (n × r)²

= k × (E(elec) + ε(amp) × n² × r²)

Multi-hop (MTE): n transmissions of r metres each, and n - 1 receptions by the relays on the way:

E(MTE) = n × (E(elec) × k + ε(amp) × k × r²) + (n - 1) × E(elec) × k

= k × ((2n - 1) × E(elec) + ε(amp) × n × r²)

When is direct cheaper? When E(direct) < E(MTE). Divide both sides by k and subtract: ε(amp) n² r² - ε(amp) n r² < (2n - 1)E(elec) - E(elec), that is ε(amp) n r² (n - 1) < 2(n - 1) E(elec). For n greater than 1 the factor n - 1 cancels, leaving

n × r² < 2 × E(elec) / ε(amp)

With the paper's values, 2 × 50 nJ / 100 pJ is 2 × 500 = 1,000 square metres. So for nodes 10 metres apart, direct transmission is cheaper for any node up to 10 hops out, which is 100 metres from the base station. That is the paper's own finding: when transmission distances are short or the electronics energy is high, "direct transmission is more energy-efficient on a global scale than MTE routing".

The program, across distances

# One long hop or many short ones? The first-order radio model of the LEACH
# paper (HICSS 2000): sending k bits over d metres costs E_ELEC*k + EPS_AMP*k*d^2,
# receiving them costs E_ELEC*k.
E_ELEC = 50e-9     # joules per bit, transmitter or receiver electronics
EPS_AMP = 100e-12  # joules per bit per square metre, transmit amplifier
K = 2000           # bits in one message, as in the paper

def tx(d):
    return E_ELEC * K + EPS_AMP * K * d * d

def rx():
    return E_ELEC * K

def direct(n, r):
    """The node n hops out sends straight to the base station, n*r metres away."""
    return tx(n * r)

def multihop(n, r):
    """Minimum transmission energy routing: n hops of r metres, n - 1 receptions."""
    return n * tx(r) + (n - 1) * rx()

print("Direct is cheaper exactly when n * r^2 < 2 * E_ELEC / EPS_AMP = %.0f m^2"
      % (2 * E_ELEC / EPS_AMP))
print()
print("  r (m)   n   n*r^2   direct (uJ)   multi-hop (uJ)   cheaper")
for r in (5, 10, 20, 40):
    for n in (2, 5, 10):
        d, m = round(direct(n, r) * 1e6, 1), round(multihop(n, r) * 1e6, 1)
        verdict = "equal" if d == m else ("direct" if d < m else "multi-hop")
        print("%7d %3d %7d %13.1f %16.1f   %s" % (r, n, n * r * r, d, m, verdict))

# Who pays? With multi-hop on a line of 5 nodes 20 m apart, node 1 (next to the
# base station) sends its own message and relays the other four.
print()
print("energy per node for one message from every node, n = 5, r = 20 m:")
n, r = 5, 20
for i in range(1, n + 1):
    relays = n - i                      # messages from nodes farther out
    energy = (relays + 1) * tx(r) + relays * rx()
    print("  node %d: %6.1f uJ (sends %d, receives %d)"
          % (i, energy * 1e6, relays + 1, relays))
munotes.in88

Single Hop or Multiple Hops: The Energy Argument Worked Out

Direct is cheaper exactly when n * r^2 < 2 * E_ELEC / EPS_AMP = 1000 m^2

  r (m)   n   n*r^2   direct (uJ)   multi-hop (uJ)   cheaper
      5   2      50         120.0            310.0   direct
      5   5     125         225.0            925.0   direct
      5  10     250         600.0           1950.0   direct
     10   2     200         180.0            340.0   direct
     10   5     500         600.0           1000.0   direct
     10  10    1000        2100.0           2100.0   equal
     20   2     800         420.0            460.0   direct
     20   5    2000        2100.0           1300.0   multi-hop
     20  10    4000        8100.0           2700.0   multi-hop
     40   2    3200        1380.0            940.0   multi-hop
     40   5    8000        8100.0           2500.0   multi-hop
     40  10   16000       32100.0           5100.0   multi-hop

energy per node for one message from every node, n = 5, r = 20 m:
  node 1: 1300.0 uJ (sends 5, receives 4)
  node 2: 1020.0 uJ (sends 4, receives 3)
  node 3:  740.0 uJ (sends 3, receives 2)
  node 4:  460.0 uJ (sends 2, receives 1)
  node 5:  180.0 uJ (sends 1, receives 0)

Read the verdict column. Every row with n × r² below 1,000 says direct; every row above it says multi-hop; the row at exactly 1,000 says equal. The algebra and the program agree, which is the point of running it. With nodes 5 metres apart, a node 10 hops out spends 600 microjoules sending directly against 1,950 by multi-hop: over three times as much for the "energy-saving" route.

Read the last block. On a line of 5 nodes 20 metres apart, one message from every node costs node 1, next to the base station, 1,300 microjoules and node 5 only 180: more than seven times as much. Node 1 dies first, then node 2, and when they are gone the rest of the line is cut off. This is the energy hole around a sink, and the LEACH paper observed it in its simulations: with MTE routing the nodes closest to the base station die first, while with direct transmission it is the farthest ones.

munotes.in89

Single Hop or Multiple Hops: The Energy Argument Worked Out

A better model, and the best number of hops

Near the ground, and at longer distances, loss grows faster than d². The same authors' journal paper of 2002 uses two regimes: free space, ε(fs) × d², up to a crossover distance, and multipath, ε(mp) × d⁴, beyond it, with ε(fs) = 10 pJ/bit/m² and ε(mp) = 0.0013 pJ/bit/m⁴. The crossover is where the two are equal, the square root of ε(fs) / ε(mp), about 87.7 metres.

# The LEACH journal model (2002): free-space loss (d^2) below a crossover
# distance d0, multipath loss (d^4) beyond it.
E_ELEC = 50e-9       # joules per bit
EPS_FS = 10e-12      # joules per bit per square metre
EPS_MP = 0.0013e-12  # joules per bit per metre to the fourth
K = 2000             # bits in one message
D0 = (EPS_FS / EPS_MP) ** 0.5

def tx(d):
    amp = EPS_FS * d ** 2 if d < D0 else EPS_MP * d ** 4
    return E_ELEC * K + amp * K

def chain(total, n):
    """Energy to carry one message over `total` metres in n equal hops."""
    return n * tx(total / n) + (n - 1) * E_ELEC * K

print("crossover distance d0 = %.1f m" % D0)
print()
print("distance   1 hop (uJ)   2 hops   4 hops   8 hops")
for total in (50, 100, 200, 400):
    row = [chain(total, n) * 1e6 for n in (1, 2, 4, 8)]
    print("%6d m %12.1f %8.1f %8.1f %8.1f" % (total, *row))
crossover distance d0 = 87.7 m

distance   1 hop (uJ)   2 hops   4 hops   8 hops
    50 m        150.0    325.0    712.5   1506.2
   100 m        360.0    400.0    750.0   1525.0
   200 m       4260.0    820.0    900.0   1600.0
   400 m      66660.0   8620.0   1740.0   1900.0

Each row has a best entry, and it moves right as the distance grows. Over 50 and 100 metres, one hop is cheapest. Over 200 metres, two hops (820 microjoules) beat one (4,260), because the single hop has crossed into the d⁴ regime. Over 400 metres, four hops (1,740) are best, about 38 times cheaper than one, and eight hops are worse than four (1,900), because the extra receptions and electronics outweigh the little amplifier energy left to save.

So the full answer is neither "single hop" nor "multi-hop": there is an optimal hop length, set by the ratio of the electronics energy to the amplifier energy and by how fast the loss grows with distance.

Distinctions

Direct transmissionMulti-hop (MTE) routing
Transmissions per message1, longn, short
Receptions per message0 (except the base station)n - 1
Energy (first-order model)k(E(elec) + ε(amp) n² r²)k((2n - 1)E(elec) + ε(amp) n r²)
Cheaper whenn r² < 2 E(elec) / ε(amp)n r² > 2 E(elec) / ε(amp)
Who dies firstThe farthest nodesThe nodes nearest the base station
munotes.in90

Single Hop or Multiple Hops: The Energy Argument Worked Out

Free-space modelTwo-regime model
Amplifier energyε(amp) d²ε(fs) d² below about 87.7 m, ε(mp) d⁴ beyond
Constants100 pJ/bit/m²10 pJ/bit/m² and 0.0013 pJ/bit/m⁴
PaperLEACH, HICSS 2000LEACH, IEEE Transactions on Wireless Communications 2002

What it does not mean

Multi-hop is not always the energy-saving choice. With the LEACH constants and hops of 10 metres, sending directly is cheaper for every node within 100 metres of the base station. The claim needs its condition.

More hops are not always better, even when multi-hop wins. The two-regime run found four hops best over 400 metres and eight hops worse.

Saving total energy is not the same as a long network lifetime. Multi-hop can use less energy in total and still kill the nodes next to the sink early, because it concentrates the relaying on them. Lifetime is about the worst node, which [Optimization Goals: Quality of Service, Energy Efficiency and Lifetime] takes up.

The model is not a real radio. It ignores listening, collisions, retransmissions and the MAC's overhead, and real radios often cannot vary their power continuously. It is a tool for reasoning about the trade-off, not a prediction to the microjoule.

Quick revision

  • First-order model (LEACH 2000): transmit E(elec)k + ε(amp)k d², receive E(elec)k; E(elec) = 50 nJ/bit, ε(amp) = 100 pJ/bit/m².
  • Line of n nodes r apart: E(direct) = k(E(elec) + ε(amp) n² r²); E(MTE) = k((2n - 1)E(elec) + ε(amp) n r²).
  • Direct is cheaper when n r² < 2 E(elec) / ε(amp), which is 1,000 m² with these constants.
  • Program: r = 5 m, n = 10: 600 against 1,950 microjoules, direct wins.
  • Energy hole: n = 5, r = 20 m, node 1 spends 1,300 microjoules, node 5 180; MTE kills the nodes near the base station first.
  • Two-regime model (2002): ε(fs) = 10 pJ/bit/m², ε(mp) = 0.0013 pJ/bit/m⁴, crossover about 87.7 m; best hops: 1 up to 100 m, 2 at 200 m, 4 at 400 m (8 is worse).

Test yourself

1. Write the first-order radio model and explain its terms. Transmitting k bits over d metres costs E(elec) × k + ε(amp) × k × d²; receiving costs E(elec) × k. E(elec) is the energy per bit of the transmitter or receiver electronics and ε(amp) the energy per bit per square metre of the transmit amplifier. LEACH used 50 nJ/bit and 100 pJ/bit/m².

munotes.in91

Single Hop or Multiple Hops: The Energy Argument Worked Out

2. Derive the condition under which direct transmission uses less energy than multi-hop routing on a line of n nodes spaced r apart. E(direct) = k(E(elec) + ε(amp) n² r²) and E(MTE) = k((2n - 1)E(elec) + ε(amp) n r²). Direct < MTE gives ε(amp) n r² (n - 1) < 2(n - 1)E(elec), so n r² < 2 E(elec) / ε(amp).

3. With E(elec) = 50 nJ/bit and ε(amp) = 100 pJ/bit/m², up to how many hops of 10 metres is direct transmission cheaper? 2 × 50 / 0.1 = 1,000 m², so n × 100 < 1,000, n < 10: for any node fewer than 10 hops out; at exactly 10 hops the two are equal.

4. Why do the nodes nearest the base station die first under multi-hop routing? Because they relay every message from the nodes farther out. On a line of five nodes, node 1 sends five messages and receives four for every one message node 5 sends.

5. With loss growing as d⁴ beyond about 88 metres, how many hops are best for 400 metres, and why not more? Four hops, at about 1,740 microjoules. Eight hops cost about 1,900, because each extra hop adds a transmission's and a reception's electronics energy while the amplifier energy left to save is small.

Contents This chapter on its own page

munotes.in92

Chapter Sixteen

Optimization Goals: Quality of Service, Energy Efficiency and Lifetime

Syllabus topic Module 1, "Introduction and Overview of WSNs: Optimization goals, figures of merit"

In one line

A sensor network is optimized not for throughput but for the quality of the information it delivers, for the energy each piece of information costs, and above all for how long it keeps working, which has to be defined before it can be measured.

In the wording a student can write in an examination: the optimization goals of a WSN are (1) quality of service, measured by the network's usefulness to the application: event detection and reporting probability, event classification error, event detection delay, missing reports, approximation accuracy and tracking accuracy; (2) energy efficiency, measured as energy per correctly received bit or energy per reported (unique) event, with a trade-off against delay; (3) network lifetime, defined as the time to first node death, the network half-life, the time to partition, the time to loss of coverage or the time to failure of the first event notification; (4) scalability; and (5) robustness. The last two are the next chapter.

Why a sensor network needs its own goals

The Internet is judged by throughput, delay and loss, because it exists to move bits. A sensor network exists to tell its user something about the world, and Karl and Willig say bluntly that the familiar measures miss the point: "The packet delivery ratio is an insufficient metric; what is relevant is the amount and quality of information that can be extracted at given sinks". A network that delivers 60 per cent of its packets may be perfect if every fire is still reported; a network that delivers 99 per cent may be useless if the 1 per cent lost are the alarms.

So the designer has to decide what "good" means for this application before tuning anything, and the goals below are the menu.

Goal 1: quality of service, measured on the information

Each measure goes with one of the four sensing tasks of [Applications of Wireless Sensor Networks].

MeasureQuestion it answersTask it belongs to
Event detection or reporting probabilityOf the events that happened, what fraction did the sink hear about?Event detection
Event classification errorOf the events reported, how many were called the wrong kind?Event detection
Event detection delayHow long from the event happening to the sink knowing?Event detection
Missing reportsIn a periodic application, how many scheduled readings never arrived?Periodic measurement
Approximation accuracyHow close is the map the sink builds to the real field?Function approximation
Tracking accuracyHow far is the reported position from the target's true position?Tracking

Notice that none of them is about bits. Event detection probability is the one to remember: a fire network is judged by the chance it reports a fire, not by how many packets arrive.

munotes.in93

Optimization Goals: Quality of Service, Energy Efficiency and Lifetime

Goal 2: energy efficiency

Energy is spent to get information, so the measures relate the two.

  • Energy per correctly received bit: the total energy spent in the network, divided by the number of bits that reached their destination intact. It counts every retransmission, every relay and every idle listen. It suits periodic data collection.
  • Energy per reported (unique) event: the energy spent divided by the number of distinct events reported. It suits event detection, where many nodes may report the same fire and only the first report counts as information.
  • The delay and energy trade-off: a network can nearly always save energy by waiting (sleeping longer, collecting several readings into one packet), so energy efficiency is usually stated at a given delay. An alarm may be allowed a second; a weekly report a week.

Goal 3: network lifetime, and why it must be defined

Karl and Willig call lifetime "a very important figure of merit" and note that "the precise definition of lifetime depends on the application at hand". The standard definitions:

  1. Time to first node death. The network is considered dead when the first node runs out of energy. The simplest, and the most pessimistic.
  2. Network half-life. The time until 50 per cent (or another fixed fraction) of the nodes have failed.
  3. Time to partition. The time until the network splits into two or more parts that cannot reach each other, so that some live nodes can no longer reach the sink.
  4. Time to loss of coverage. The time when, for the first time, some point of the observed region is no longer covered by any sensor node.
  5. Time to failure of the first event notification. The time when, for the first time, an event somewhere in the region could not be reported to the sink, because no node covering it is alive or because the ones covering it cannot reach the sink.

Each suits a different application. A network that must watch every point (a perimeter intrusion system) is dead at loss of coverage; one that only needs statistics (an average soil moisture) can carry on to its half-life; one whose nodes are all equally important is dead at the first node death.

Worked example: five lifetimes for one network

A 5 by 5 grid of 25 nodes, 10 metres apart, reports to a sink beside the middle of one edge. Nodes within 15 metres are neighbours. Every round each node sends one reading to the sink by the fewest hops, relays pay to receive and resend, and a node that has been cut off from the sink still wastes one transmission trying. Each node starts with 0.5 J, as in the LEACH simulations of [Single Hop or Multiple Hops: The Energy Argument Worked Out], and the radio follows the same first-order model. The program measures when each definition says the network died.

munotes.in94

Optimization Goals: Quality of Service, Energy Efficiency and Lifetime

# Five definitions of network lifetime, measured on one simulated network.
# 25 nodes on a 5 x 5 grid 10 m apart; the sink sits beside the middle of one
# edge. Nodes within 15 m of each other (diagonals included) are neighbours.
# Every round each node sends one 2000-bit reading along a fewest-hops path to
# the sink, and relays pay to receive and resend. A node cut off from the sink
# still transmits its reading once, uselessly. Radio: the LEACH first-order model.
import math
from collections import deque

E_ELEC, EPS, K, E0 = 50e-9, 100e-12, 2000, 0.5   # J/bit, J/bit/m^2, bits, J
NODES = [(10 * x, 10 * y) for x in range(5) for y in range(5)]
SINK = (-10, 20)
RANGE = 15
energy = {n: E0 for n in NODES}

def tx(a, b):
    return E_ELEC * K + EPS * K * math.dist(a, b) ** 2

def next_hops(alive):
    """Breadth-first search from the sink: each connected node's next hop."""
    parent, queue = {}, deque([SINK])
    while queue:
        p = queue.popleft()
        for q in sorted(alive):
            if q not in parent and math.dist(p, q) <= RANGE:
                parent[q] = p
                queue.append(q)
    return parent

def covered(point, sensors, rs=15):
    return any(math.dist(point, s) <= rs for s in sensors)

found, rnd = {}, 0
while len(found) < 5:
    rnd += 1
    alive = {n for n in NODES if energy[n] > 0}
    parent = next_hops(alive)
    checks = {
        "time to first node death": len(alive) < len(NODES),
        "time to partition": len(parent) < len(alive),
        "time to failure of first event notification":
            not all(covered(p, parent) for p in NODES),
        "time to loss of coverage": not all(covered(p, alive) for p in NODES),
        "network half-life": len(alive) <= len(NODES) // 2,
    }
    for name, happened in checks.items():
        if happened and name not in found:
            found[name] = rnd
    for n in alive:
        if n in parent:                          # carry the reading to the sink
            hop = n
            while hop != SINK:
                energy[hop] -= tx(hop, parent[hop])
                if parent[hop] != SINK:
                    energy[parent[hop]] -= E_ELEC * K
                hop = parent[hop]
        else:                                    # cut off: one wasted attempt
            energy[n] -= E_ELEC * K + EPS * K * RANGE ** 2

for name, r in sorted(found.items(), key=lambda item: item[1]):
    print("%-45s round %d" % (name, r))
time to first node death                      round 114
time to partition                             round 280
time to failure of first event notification   round 280
network half-life                             round 2998
time to loss of coverage                      round 3017

The same network, the same energy, and its "lifetime" is anything from 114 rounds to 3,017, a factor of more than twenty-six. Four things in the output repay attention.

munotes.in95

Optimization Goals: Quality of Service, Energy Efficiency and Lifetime

The first death comes early because the fewest-hops tree is lopsided: of the three nodes beside the sink, one carries 19 of the 25 readings every round and the other two carry 3 each (a separate count of the tree confirms it). But the network keeps working: the other two carry on, and the routes are rebuilt around the dead node.

Partition comes at round 280, when all three nodes beside the sink are dead. From then on, however much energy the other nodes have, nothing reaches the sink. For a data collection network this, not the half-life, is the real end.

Failure of the first event notification comes at the same moment as partition. Here the sensing range equals the radio range, so the cut-off nodes' area is covered by no connected node from the instant they are cut off. With a larger sensing range the two could come apart.

Half-life and loss of coverage come thousands of rounds later, because cut-off nodes spend almost nothing and live on, uselessly. A designer who quoted the half-life would be describing a network that had delivered nothing for over 2,700 rounds.

So an answer that says the network lifetime is 2,998 rounds without saying which definition is not an answer.

Distinctions

DefinitionNetwork is dead whenPessimismSuits
First node deathOne node is out of energyMost pessimisticNetworks where every node matters
Half-lifeHalf the nodes are deadOptimisticStatistical applications
PartitionSome live nodes cannot reach the sinkMiddleData collection to one sink
Loss of coverageSome point is not coveredDepends on redundancySurveillance of every point
First failed event notificationAn event cannot be reportedCombines coverage and connectivityEvent detection
Energy per correctly received bitEnergy per reported unique event
CountsAll energy, per useful bit deliveredAll energy, per distinct event reported
SuitsPeriodic collectionEvent detection
PunishesRetransmissions, idle listeningDuplicate reports of one event

What it does not mean

Quality of service here is not bandwidth and jitter. It is detection probability, delay of an event report, accuracy of a map or a track.

Energy efficiency is not the same as low power. A node that draws little but delivers nothing is not efficient; the measure is energy per useful result.

Network lifetime is not one number until a definition is chosen. The simulation gave five numbers from one network.

A long half-life does not mean a working network. Half the nodes can be alive and cut off from the sink.

Quick revision

  • Goals (Karl and Willig 3.2): quality of service, energy efficiency, lifetime, scalability, robustness.
  • QoS measures: event detection or reporting probability, event classification error, event detection delay, missing reports, approximation accuracy, tracking accuracy. "The packet delivery ratio is an insufficient metric."
  • Energy measures: energy per correctly received bit; energy per reported unique event; the delay and energy trade-off.
  • Lifetime: first node death, half-life, partition, loss of coverage, failure of the first event notification.
  • Simulation, 25 nodes, 0.5 J each: first death 114, partition 280, first failed notification 280, half-life 2,998, loss of coverage 3,017.
munotes.in96

Optimization Goals: Quality of Service, Energy Efficiency and Lifetime

Test yourself

1. What are the optimization goals of a wireless sensor network? Quality of service (measured on the information: detection probability, classification error, detection delay, missing reports, approximation and tracking accuracy), energy efficiency (energy per correctly received bit or per unique event, traded against delay), network lifetime, scalability and robustness.

2. Give five definitions of network lifetime. Time to first node death; network half-life (half the nodes failed); time to partition (some live nodes can no longer reach the sink); time to loss of coverage (some point no longer covered); time to failure of the first event notification (some event cannot be reported).

3. Why is the packet delivery ratio insufficient as a quality measure for a sensor network? Because the user wants information, not packets. What matters is how much and how good the information at the sink is: whether events are detected and reported, how accurate the map or track is. A network can lose many packets and still report every event, or deliver most packets and miss the important ones.

4. In the simulation, why did the network's half-life come about 2,700 rounds after it was partitioned? Because once the nodes beside the sink died, the other nodes were cut off and spent almost no energy, so they stayed alive for thousands of rounds while delivering nothing. The half-life counts them as alive.

5. Which lifetime definition would you use for a perimeter intrusion detection network, and why? Time to loss of coverage, or time to failure of the first event notification: the network has failed as soon as an intruder could cross some point without being detected and reported.

Contents This chapter on its own page

munotes.in97

Chapter Seventeen

Figures of Merit: Scalability, Robustness and Measuring a Network

Syllabus topic Module 1, "Introduction and Overview of WSNs: Optimization goals, figures of merit"

In one line

A figure of merit is a number that says how good a network is at something; for a sensor network the most important are how it copes with growth (scalability), how it copes with failure (robustness), and the measured delivery, delay, overhead and energy of its traffic.

In the wording a student can write in an examination: besides quality of service, energy efficiency and lifetime, a WSN is optimized for scalability, the ability to keep its performance as the number of nodes, the density or the area grows, and robustness, the ability to keep performing when nodes fail, links fluctuate or the topology changes. Protocols are compared by figures of merit such as throughput, packet delivery ratio, end-to-end delay, jitter, control overhead and energy per delivered packet.

Why define them once

Because every later chapter compares protocols, and a comparison is only fair when both sides are measured the same way. AODV is better than DSR means nothing until you say better at what, measured how, on what traffic. This chapter is the vocabulary.

Goal 4: scalability

A WSN may have tens of nodes or tens of thousands, and Karl and Willig require that "the employed architectures and protocols must be able scale to these numbers". Scale has three dimensions:

  • Number of nodes. A protocol whose cost grows faster than the number of nodes will fail at some size.
  • Density. Karl and Willig list a "wide range of densities" as a requirement in itself: the number of nodes per unit area differs between applications and changes over time within one network as nodes fail or move.
  • Area. More hops between the farthest node and the sink mean more delay and more relaying.

What makes a protocol scalable. The guideline Karl and Willig give is locality: a node "should attempt to limit the state that they accumulate during protocol processing to only information about their direct neighbors". A node that keeps a table with an entry for every other node needs memory that grows with the network; a node that knows only its neighbours needs the same memory in a network of fifty or fifty thousand.

Worked: the cost of asking a question. A sink wants the temperature in one room of a large building. If it floods the query to every node, every node transmits it once, so the cost grows with the whole network: 100 transmissions in a 100-node network, 10,000 in a 10,000-node one. If the query is scoped to the room (sent only towards nodes whose location is in the room, as the design principles of [Design Principles: Data Centricity, Location, Activity and Heterogeneity] allow), its cost grows with the size of the room and the distance to it, not with the size of the building. The first design does not scale; the second does.

munotes.in98

Figures of Merit: Scalability, Robustness and Measuring a Network

Goal 5: robustness

A network is robust if its performance degrades gracefully, not suddenly, when things go wrong: nodes fail or run flat, links fade in and out, obstacles appear, interference rises, or nodes move. Karl and Willig's requirements of fault tolerance (tolerating node failures through redundant deployment) and maintainability (monitoring its own health and adapting, for example by giving lower quality when energy runs short) are two sides of it.

Graceful degradation is the phrase to use. When 10 per cent of the nodes fail, a robust network loses a little coverage or accuracy; a brittle one stops reporting altogether. The simulation in the previous chapter showed a brittle point: when the three nodes beside the sink died, the whole network was cut off at once. A more robust design gives the sink more neighbours, or uses several sinks, or rotates the heavy relaying role.

The measures used to compare protocols

These are the figures of merit that simulations and experiments report. Each has a definition that is worth writing exactly.

  1. Throughput: the amount of data delivered to its destinations per unit time, in bits per second. Only data that arrives counts.
  2. Goodput: throughput counting only the useful payload, not headers, retransmitted duplicates or control packets.
  3. Packet delivery ratio (PDR): packets received at their destinations divided by packets sent by their sources, usually as a percentage.
  4. End-to-end delay: for each delivered packet, the time it was received less the time it was sent; reported as an average (and often also a maximum or a percentile), over delivered packets only.
  5. Jitter: the variation of the delay from packet to packet, for example as the spread between the smallest and largest delay, or their standard deviation.
  6. Control overhead, or routing overhead: the control packets (route requests, replies, errors, beacons) a protocol sends, often normalised as control packets per data packet delivered.
  7. Energy per delivered packet (or per correctly received bit): total energy spent divided by what was delivered, the energy measure of [Optimization Goals: Quality of Service, Energy Efficiency and Lifetime].

Worked example: measuring a small trace

A source sends ten readings of 64 bytes, one a second from time 0 to time 9 seconds, to a sink four hops away. The routing protocol sent 30 control packets during the run. The trace records what arrived.

PacketSent at (s)ArrivedDelay (ms)
10yes42
21yes38
32lost
43yes55
54yes41
65lost
76yes120
87yes44
98yes39
109yes47
Total8 arrived426
munotes.in99

Figures of Merit: Scalability, Robustness and Measuring a Network

Packet delivery ratio. 8 of 10 arrived: 8 / 10 = 0.8, so the PDR is 80 per cent.

Average end-to-end delay. Over the 8 delivered packets only: 426 / 8 = 53.25 milliseconds. (Averaging over all ten, with the lost ones counted as zero, would be a mistake: it would make losses look like good performance.)

Jitter. The delays run from 38 to 120 milliseconds, a spread of 120 - 38 = 82 milliseconds, almost all of it from packet 7. One slow packet (perhaps retransmitted, or held in a queue) dominates the jitter, and it also pulls the average up: without it the other seven average 306 / 7, about 43.7 milliseconds. That is why reports give a percentile or the median beside the average.

Throughput. 8 packets of 64 bytes arrived, which is 8 × 64 × 8 = 4,096 bits, over the 10-second run: 4,096 / 10 = 409.6 bits per second.

Control overhead. 30 control packets for 8 delivered data packets: 30 / 8 = 3.75 control packets per delivered packet. A protocol that delivered the same 8 with 12 control packets would have an overhead of 1.5, and if the two were otherwise equal it would be the better one for a battery-powered network, because every control packet costs energy.

Distinctions

ThroughputPacket delivery ratio
IsData delivered per secondFraction of packets delivered
Unitbits per secondper cent
Depends on the offered loadYes, stronglyLess so
In the trace409.6 bit/s80 per cent
End-to-end delayJitter
IsHow long a packet takesHow much that time varies
In the trace53.25 ms average82 ms spread
Matters most forAlarms, controlStreams, voice, synchronised sampling
ScalabilityRobustness
QuestionDoes it still work when the network grows?Does it still work when things fail?
Design answerLocality, scoping, hierarchyRedundancy, several sinks, adaptive routes
Failure looks likeOverhead or state growing faster than the networkA sudden collapse after a few failures

What it does not mean

A high packet delivery ratio does not mean a good sensor network. It is the measure Karl and Willig call insufficient; it says nothing about whether the right packets, the alarms, arrived.

Delay is not averaged over lost packets. Only delivered packets have a delay.

Throughput is not the radio's data rate. A 250 kbps radio in the trace delivered 409.6 bits per second, because the application sent little and some of it was lost.

Scalability is not only about the number of nodes. Density and area matter as much, and a protocol can scale in one and fail in another.

munotes.in100

Figures of Merit: Scalability, Robustness and Measuring a Network

Low overhead is not always better. A protocol that sends almost no control packets may deliver less, or react slowly when routes break. Overhead must be read beside delivery and delay.

Quick revision

  • Scalability: keep performance as nodes, density and area grow; the guideline is locality (state only about direct neighbours) and scoping (a query's cost should grow with the region asked about, not the network).
  • Robustness: graceful degradation under failures, fading, interference and movement; fault tolerance by redundancy; maintainability.
  • Measures: throughput, goodput, packet delivery ratio, end-to-end delay (over delivered packets), jitter, control overhead, energy per delivered packet.
  • Trace: PDR 80 per cent; average delay 53.25 ms; jitter 82 ms spread; throughput 409.6 bit/s; overhead 3.75 control packets per delivered packet.

Test yourself

1. Define scalability and robustness as optimization goals of a WSN. Scalability: the network keeps its performance as the number of nodes, the density or the area grows, which requires protocols whose state and overhead grow slowly, typically by locality. Robustness: performance degrades gracefully when nodes fail, links fluctuate or the topology changes.

2. Define packet delivery ratio, end-to-end delay and throughput. PDR: packets received at their destinations divided by packets sent. End-to-end delay: receive time minus send time for each delivered packet, averaged over delivered packets. Throughput: data delivered per unit time, in bits per second.

3. Ten packets are sent and eight arrive with delays totalling 426 ms. Give the PDR and the average delay. PDR 8 / 10 = 0.8, which is 80 per cent; average delay 426 / 8 = 53.25 ms.

4. Why is flooding a query to the whole network not scalable, and what is the alternative? Its cost grows with the number of nodes: every node transmits the query once. Scoping the query to the region of interest, by location or by attribute, makes the cost grow with the region instead.

5. What is control overhead, and why does it matter more in a sensor network than on the Internet? The control packets a protocol sends (route requests, replies, errors, beacons), often measured per data packet delivered. In a sensor network every transmission costs battery energy, so overhead directly shortens the network's lifetime.

Contents This chapter on its own page

munotes.in101

Chapter Eighteen

Deployment and Coverage: Random Against Grid

Syllabus topic Module 1, "Introduction and Overview of WSNs: Network architecture and optimization" (and the paired practical, "Wireless Sensor Network Deployment and Coverage Analysis")

In one line

Coverage asks whether every point of the field is watched by at least one sensor; connectivity asks whether every sensor can get its data to the sink; a planned grid achieves both with far fewer nodes than scattering them at random, but random deployment is often the only kind possible.

In the wording a student can write in an examination: a point is covered if it lies within the sensing range r(s) of at least one working node, and the coverage of a deployment is the fraction of the field that is covered. A network is connected if every node has a multi-hop path to the sink over links no longer than the communication range r(c). In deterministic (grid) deployment nodes are placed at planned points; in random deployment they are scattered, for example from the air. For uniform random deployment of N nodes over an area A, the expected coverage is approximately 1 - e^(-N π r(s)² / A). Zhang and Hou proved that if r(c) is at least 2 r(s), complete coverage of a convex area implies connectivity.

Why coverage and connectivity are separate questions

A node can be watching its patch of ground perfectly and have no way to tell anyone: that is coverage without connectivity. And a network can be perfectly connected with a gap in the middle nobody watches: connectivity without coverage. A monitoring network needs both, and they depend on two different ranges: what a node can sense and how far it can talk.

The two ways to deploy

Deterministic or grid deployment. Nodes are placed at chosen points, usually on a regular pattern: a square grid or a triangular (hexagonal) one. It needs someone to walk the field or a structure to mount the nodes on, and it gives predictable coverage with the fewest nodes.

Random deployment. Nodes are scattered, "dropping from a plane" in the survey's list, or thrown by hand. It is the only choice for a forest, a battlefield or a disaster area, and it gives uneven density: clumps where nodes landed together and holes where none did. The answer to holes is more nodes, and the question is how many more.

The square grid's spacing rule

On a square grid with spacing s, the point farthest from every node is the centre of a cell, half a diagonal from the four nodes at its corners. By Pythagoras, the square of that distance is

(s / 2)² + (s / 2)² = s² / 2

so the distance is s divided by the square root of 2, about s / 1.414. The whole field is covered exactly when that centre is within sensing range, that is when s is at most the square root of 2 times r(s), about 1.414 r(s).

munotes.in102

Deployment and Coverage: Random Against Grid

With r(s) = 10 metres, the spacing must be at most about 14.14 metres. The grid is connected if neighbouring nodes are within radio range, s at most r(c), which is easily met when r(c) is 20 metres.

Zhang and Hou's theorem

Zhang and Hou proved the result that links the two questions: "if the radio range is at least twice of the sensing range, a complete coverage of a convex area implies connectivity among the working set of nodes". They also showed that the condition is necessary: with r(c) below 2 r(s) there are coverings that are not connected.

The reason is short enough to write in an answer. If every point is covered, take any two nodes whose sensing discs touch or overlap; their centres are at most 2 r(s) apart, so with r(c) at least 2 r(s) they can talk directly. In a convex field fully covered by discs, the discs form an unbroken chain from any node to any other, so every node can reach every other, and the sink in particular.

Its practical meaning is large: with r(c) at least 2 r(s), a designer need only ensure coverage, and connectivity comes free. The program below uses exactly that ratio, r(s) = 10 metres and r(c) = 20 metres.

The practical's experiment

The program places 49 and 64 nodes on grids, and 49, 64, 100 and 150 nodes at random in twenty different layouts each, in a 100 metre square field. Coverage is sampled at 2,500 points, one every 2 metres; connectivity is tested by searching the graph of links no longer than 20 metres.

49 nodes on a grid and 49 at random, each with its 10 m sensing disc

Figure 18.1 The same number of nodes, planned and scattered: the program's first random layout

# Grid against random deployment: coverage and connectivity.
# A 100 m by 100 m field, sensing radius 10 m, radio range 20 m (twice it).
import math
import random

SIDE, RS, RC = 100.0, 10.0, 20.0
POINTS = [(x + 1.0, y + 1.0) for x in range(0, 100, 2) for y in range(0, 100, 2)]

def coverage(nodes):
    """Fraction of 2,500 sample points within RS of at least one node."""
    hit = sum(1 for p in POINTS if any(math.dist(p, n) <= RS for n in nodes))
    return hit / len(POINTS)

def connected(nodes):
    """Is every node reachable from the first, over links of at most RC?"""
    seen, stack = {0}, [0]
    while stack:
        i = stack.pop()
        for j, n in enumerate(nodes):
            if j not in seen and math.dist(nodes[i], n) <= RC:
                seen.add(j)
                stack.append(j)
    return len(seen) == len(nodes)

def grid(k):
    s = SIDE / k                                   # spacing between rows
    return [(s / 2 + s * i, s / 2 + s * j) for i in range(k) for j in range(k)]

for k in (7, 8):
    g, s = grid(k), SIDE / k
    worst = s / math.sqrt(2)                       # a cell's centre to its corners
    print("grid %d x %d = %d nodes, spacing %.2f m: sampled coverage %.4f, connected %s"
          % (k, k, len(g), s, coverage(g), connected(g)))
    print("   farthest point from any node: %.2f m, so every point is covered: %s"
          % (worst, worst <= RS))

print()
print("random, 20 layouts each:")
for n in (49, 64, 100, 150):
    covs, links = [], 0
    for seed in range(20):
        rnd = random.Random(seed)
        nodes = [(rnd.uniform(0, SIDE), rnd.uniform(0, SIDE)) for _ in range(n)]
        covs.append(coverage(nodes))
        links += connected(nodes)
    ideal = 1 - math.exp(-n * math.pi * RS ** 2 / SIDE ** 2)
    print("  %3d nodes: mean coverage %.4f (formula %.4f), worst %.4f, connected in %d of 20"
          % (n, sum(covs) / 20, ideal, min(covs), links))
munotes.in103

Deployment and Coverage: Random Against Grid

grid 7 x 7 = 49 nodes, spacing 14.29 m: sampled coverage 1.0000, connected True
   farthest point from any node: 10.10 m, so every point is covered: False
grid 8 x 8 = 64 nodes, spacing 12.50 m: sampled coverage 1.0000, connected True
   farthest point from any node: 8.84 m, so every point is covered: True

random, 20 layouts each:
   49 nodes: mean coverage 0.7502 (formula 0.7855), worst 0.6940, connected in 3 of 20
   64 nodes: mean coverage 0.8369 (formula 0.8661), worst 0.7980, connected in 12 of 20
  100 nodes: mean coverage 0.9365 (formula 0.9568), worst 0.9024, connected in 18 of 20
  150 nodes: mean coverage 0.9841 (formula 0.9910), worst 0.9704, connected in 20 of 20

Reading the output

The sampled 7 by 7 grid looks perfect and is not. Its spacing is 100 / 7, about 14.29 metres, just over the 14.14-metre limit, so the centre of every cell is 10.10 metres from the nearest node and a tiny patch around it is uncovered. The 2-metre sampling never lands in those patches and reports 1.0000. The exact check on the next line says what the sampling cannot. A coverage figure is only as good as the points it was measured at, which matters in the practical: a simulation that samples coarsely can pass a layout with holes. The 8 by 8 grid, at 12.5-metre spacing, covers every point.

Random deployment needs well over twice as many nodes, and still leaves holes. With 49 random nodes, the mean coverage is about 75 per cent and the worst layout under 70 per cent; the grid of the same 49 nodes misses almost nothing. Random layouts reach an average of about 98 per cent only at 150 nodes, and even then the worst of twenty layouts leaves about 3 per cent uncovered.

munotes.in104

Deployment and Coverage: Random Against Grid

The formula is optimistic, and the program shows why. 1 - e^(-N π r(s)² / A) assumes nodes scattered over an unbounded plane. In a real field, a node near the edge wastes part of its disc outside the field, so the measured coverage is always a little below the formula: 0.7502 against 0.7855 at 49 nodes. A student who quoted the formula for a small field would overstate the coverage.

Connectivity improves with density as well. With 49 random nodes, only 3 of the 20 layouts are fully connected: some node is always stranded more than 20 metres from all the others. At 150 nodes all 20 are connected. The grids, with neighbours 12.5 to 14.3 metres apart, are always connected, as the theorem promises for a fully covered field.

Distinctions

Grid deploymentRandom deployment
HowNodes placed at planned pointsNodes scattered, by hand or from the air
Nodes for full coverage here64 (8 by 8 grid)Over 150 for about 98 per cent on average
CoveragePredictable; full if s is at most 1.414 r(s)Uneven: clumps and holes
NeedsAccess to the fieldNothing but a way to scatter
SuitsBuildings, farms, factoriesForests, battlefields, disaster areas
CoverageConnectivity
QuestionIs every point sensed?Can every node reach the sink?
Depends onSensing range r(s) and placementCommunication range r(c) and placement
Linked byZhang and Hou: r(c) at least 2 r(s) makes full coverage imply connectivity

What it does not mean

A sampled coverage of 100 per cent is not proof of full coverage. The 7 by 7 grid had holes the samples missed.

Random deployment is not inferior design. It is often the only possible design; the answer is to deploy more nodes and to let redundant ones sleep, which [Energy Efficiency in Ad Hoc Networks: Where the Energy Goes] discusses.

Coverage does not imply connectivity in general. It does when r(c) is at least 2 r(s) and the area is convex. With a shorter radio range, a fully covered field can be disconnected.

The formula is not the answer for a real field. It ignores the edges, and it describes an average: any one layout can be much worse, as the worst-case column showed.

Quick revision

  • Coverage: fraction of the field within r(s) of some working node. Connectivity: every node has a path to the sink over links of at most r(c).
  • Grid: planned; square grid covers fully when s is at most 1.414 r(s) (14.14 m for r(s) = 10 m). Random: scattered; uneven; needs many more nodes.
  • Random coverage formula, ignoring edges: 1 - e^(-N π r(s)² / A).
  • Zhang and Hou: if r(c) is at least 2 r(s), full coverage of a convex area implies connectivity, and the condition is necessary.
  • Program (100 m field, r(s) = 10 m, r(c) = 20 m): 7 by 7 grid has tiny holes (cell centres 10.10 m from nodes), which 2 m sampling missed; 8 by 8 covers fully; random 49 nodes about 75 per cent, 150 nodes about 98 per cent; connected in 3 of 20 layouts at 49 nodes, 20 of 20 at 150.
munotes.in105

Deployment and Coverage: Random Against Grid

Test yourself

1. Define coverage and connectivity. Coverage: the fraction of the field lying within the sensing range of at least one working node. Connectivity: every node can reach the sink by a multi-hop path over links no longer than the communication range.

2. On a square grid, what spacing covers the whole field for a sensing range r(s)? At most the square root of 2 times r(s), about 1.414 r(s), because the farthest point from every node is a cell's centre, whose squared distance from its corners is (s / 2)² + (s / 2)², that is s² / 2.

3. State Zhang and Hou's result on coverage and connectivity, and give its practical meaning. If the radio range is at least twice the sensing range, complete coverage of a convex area implies connectivity; the condition is also necessary. A designer can then concentrate on coverage and get connectivity free.

4. Why is measured coverage in a real field lower than 1 - e^(-N π r² / A)? The formula assumes an unbounded plane. In a bounded field, nodes near the edges waste part of their sensing discs outside the field, so less of the field is covered than the formula predicts.

5. Compare grid and random deployment of 49 nodes in the simulation. The grid covered almost all the field (with tiny holes at cell centres because its spacing slightly exceeded 1.414 r(s)) and was always connected. Random layouts covered about 75 per cent on average, under 70 per cent at worst, and were fully connected in only 3 of 20 layouts.

Contents This chapter on its own page

munotes.in106

Chapter Nineteen

Design Principles: Distributed Organisation and In-network Processing

Syllabus topic Module 1, "Introduction and Overview of WSNs: Design principles for WSNs" (and the paired practical, "Sensor Mote and Base Station Communication ... data aggregation and sink node operation")

In one line

A sensor network is designed to organise itself without a central controller, and to process data inside the network, combining readings as they travel, so that it sends answers instead of raw numbers.

In the wording a student can write in an examination: the design principles of WSNs include (1) distributed organisation: nodes cooperate and organise the network themselves, without a central controller; (2) in-network processing: data is processed inside the network rather than only at the edge, by aggregation (combining values on their way to the sink), distributed source coding and compression, distributed and collaborative signal processing, and mobile code or agents; (3) adaptive fidelity and accuracy: trading the quality of the result against energy; plus the principles of the next chapter: data centricity, exploiting location, activity patterns and heterogeneity, and component-based protocol stacks with cross-layer optimisation.

Why sensor networks need design principles of their own

Every principle in this chapter and the next answers the same problem from [The Challenges of Wireless Sensor Networks]: the radio is expensive and the nodes are many, small and unreliable. A central controller would be a single point of failure and a traffic bottleneck; raw data sent to the edge would cost the most expensive thing a node does, transmitting. So the principles push decisions and computation out into the network.

Principle 1: distributed organisation

What it is. There is no central controller that knows every node and gives it instructions. Nodes discover their neighbours, form routes, elect cluster heads and schedule their sleeping by exchanging messages with each other, locally.

Why. A central controller needs to know the whole network, which does not scale ([Figures of Merit: Scalability, Robustness and Measuring a Network]); it needs every node to talk to it, which concentrates traffic around it; and when it fails, everything fails. Distributed schemes degrade gracefully instead.

The cost. A distributed decision is made with local knowledge, so it is usually not the best possible one. The network trades optimality for robustness and scale. Where a central entity exists anyway (the sink), some decisions can be centralised to improve them; LEACH's later version does exactly this for choosing clusters ([LEACH: Clusters That Take Turns]).

Principle 2: in-network processing

Karl and Willig explain the idea with an example: to find the highest or the average temperature in an area, "readings from individual sensors can be aggregated as they propagate through the network, reducing the amount of data to be transmitted and hence improving the energy efficiency". There are four forms.

Aggregation

Nodes combine the data they receive with their own before forwarding it. The simplest case is a collection tree: every node forwards towards the sink through its parent, and each node combines its own reading with its children's into one message. Typical aggregates are SUM, COUNT, AVERAGE, MIN and MAX.

munotes.in107

Design Principles: Distributed Organisation and In-network Processing

How to aggregate correctly. TinyDB, the query system for sensor networks described in [Middleware Approaches, and TinyDB], implements every aggregate through three functions:

  1. an initializer i, which turns one reading into a partial state record;
  2. a merging function f, which combines two partial state records into one; it must be commutative and associative, so that the order in which records meet does not matter;
  3. an evaluator e, which turns the final record at the sink into the answer.

For AVERAGE the partial state record is the pair (SUM, COUNT): i(x) = (x, 1), f adds the pairs, and e(S, C) = S / C. MAX needs only the largest value so far; COUNT only the count.

Not everything aggregates. An aggregate whose partial state stays small (SUM, COUNT, MAX, AVERAGE) saves a lot. The median does not: to merge two medians you need all the values behind them, so its "partial state" is the whole list and nothing is saved.

A caution the survey adds. Akyildiz and colleagues give it as a design principle of the network layer that data aggregation is useful only when it does not hinder the collaborative effort of the nodes. Aggregation delays data (a node must wait for its children), and it destroys detail the sink might have needed. An alarm must never wait to be averaged.

Distributed source coding and compression

Neighbouring nodes see similar values: two temperature sensors a metre apart read almost the same. Their data is therefore correlated, and a network can send less than the sum of what they measure. The simplest version: send the first node's reading in full and the second as a small difference from it. More advanced schemes let each node compress its own readings knowing only that its neighbours' are correlated, without exchanging them first.

Distributed and collaborative signal processing

Some questions need several nodes' raw signals at once. The direction of a sound, the position of a vehicle, whether a seismic tremor is an earthquake: each single node sees too little. Nodes near the event exchange their signals (or features extracted from them), compute the answer together, and send only the answer. This is the "collaboration" in Karl and Willig's required mechanism: "several sensors have to collaborate to detect an event".

Mobile code and agent-based networking

Instead of moving data to where the program is, move the program to where the data is. A small piece of code, an agent, travels from node to node, processes the data it finds, and carries a result on. It suits tasks whose processing is needed only in one area, such as following a target.

munotes.in108

Design Principles: Distributed Organisation and In-network Processing

Principle 3: adaptive fidelity and accuracy

Fidelity is how accurate or detailed a result is. A network can usually produce a better result by spending more energy: more nodes sampling, more often, with more retransmissions. Adaptive fidelity means choosing that trade-off deliberately and changing it as conditions change: sample every ten minutes when nothing is happening and every ten seconds when a fire is suspected; let most nodes sleep until an event is detected, then wake more to track it; give lower quality when batteries run low. It is Karl and Willig's "exploit trade-offs" mechanism made into a design rule.

The sink's part in aggregation

The practical asks for a scenario with "multiple sensor motes and a base station to study data aggregation and sink node operation". The sink is where aggregation ends: it receives the partial state records from its children, merges them with f, and applies the evaluator e to produce the answer, which it passes to the gateway and the user. In a TinyDB-style system it also starts each round: it sends the query out, and each epoch (the period of one round of sampling) every node samples, merges its children's records with its own, and forwards one record.

Worked example: aggregating an average on a tree

Fifteen nodes report temperatures to a sink along the tree in the figure. Each node's reading is printed beneath it.

A collection tree of 15 nodes with each node's temperature

Figure 19.1 The collection tree of the worked example

The program counts the transmissions with and without aggregation, computes the average the TinyDB way, and then makes the mistake a first program usually makes.

# In-network aggregation on a collection tree, TinyDB style.
# Node -> parent (0 is the sink), and each node's temperature reading.
PARENT = {1: 0, 2: 0, 3: 0, 4: 1, 5: 1, 6: 2, 7: 2, 8: 3, 9: 3,
          10: 4, 11: 4, 12: 5, 13: 6, 14: 8, 15: 8}
READING = {1: 31, 2: 29, 3: 33, 4: 30, 5: 28, 6: 35, 7: 32, 8: 27,
           9: 30, 10: 34, 11: 29, 12: 31, 13: 36, 14: 28, 15: 30}

def depth(n):
    return 0 if n == 0 else 1 + depth(PARENT[n])

def children(n):
    return sorted(c for c, p in PARENT.items() if p == n)

# Without aggregation every reading is forwarded hop by hop to the sink.
raw = sum(depth(n) for n in PARENT)
print("transmissions without aggregation:", raw)
print("transmissions with aggregation:   ", len(PARENT), "(one record per node)")

# TinyDB's AVERAGE: initializer i(x) = <x, 1>, merge f adds the pairs,
# evaluator e(<S, C>) = S / C.
def record(n):
    s, c = READING[n], 1
    for k in children(n):
        ks, kc = record(k)
        s, c = s + ks, c + kc
    return s, c

s, c = 0, 0
for k in children(0):
    ks, kc = record(k)
    s, c = s + ks, c + kc
print()
print("partial state at the sink: SUM = %d, COUNT = %d, AVERAGE = %.4f" % (s, c, s / c))

# The mistake: each node forwards the plain average of itself and its children.
def naive(n):
    values = [READING[n]] + [naive(k) for k in children(n)]
    return sum(values) / len(values)

wrong = sum(naive(k) for k in children(0)) / len(children(0))
print("averaging averages instead gives:      %.4f" % wrong)
print("MAX needs no count: %d" % max(READING.values()))
munotes.in109

Design Principles: Distributed Organisation and In-network Processing

transmissions without aggregation: 33
transmissions with aggregation:    15 (one record per node)

partial state at the sink: SUM = 463, COUNT = 15, AVERAGE = 30.8667
averaging averages instead gives:      31.0370
MAX needs no count: 36

Transmissions. Without aggregation, each reading travels as many hops as its node's depth: three nodes at depth 1, six at depth 2 and six at depth 3 give 3 + 12 + 18 = 33 transmissions. With aggregation, each node sends one record: 15. The saving grows with the depth of the tree, and it is largest for the nodes near the sink, which would otherwise carry everyone's readings.

The right average. The sink receives (SUM, COUNT) records and ends with SUM = 463 and COUNT = 15, so the average is 463 / 15, about 30.8667 degrees.

The wrong one. If each node forwards only the plain average of itself and its children, the sink gets 31.0370. The error is that a subtree of five nodes and a subtree of one are given equal weight. On this tree the error is small; on a lopsided tree it can be large, and nothing in the output warns that it is wrong. This is why the partial state record carries the COUNT.

Distinctions

AggregationCompressionCollaborative signal processing
CombinesValues into one summary (sum, maximum)Correlated readings into fewer bitsSeveral nodes' raw signals into one decision
Keeps detailNo, only the summaryYes, approximatelyNo, only the decision
ExampleAverage temperature of a floorTwo neighbours' near-identical readingsThe direction of a sound
Partial state of AVERAGEPartial state of MAXPartial state of MEDIAN
Holds(SUM, COUNT)The largest value so farEvery value
SizeTwo numbersOne numberGrows with the subtree
Saving from aggregationLargeLargeNone

What it does not mean

In-network processing is not free. It costs processing energy (small next to the radio's), delay while nodes wait for their children, and the loss of detail.

An average of averages is not the average. Use (SUM, COUNT).

munotes.in110

Design Principles: Distributed Organisation and In-network Processing

Distributed does not mean random. Nodes follow the same protocol; what they lack is a central authority and a global view.

Adaptive fidelity is not lower quality. It is the right quality for the moment, higher when it matters and lower when it does not.

Quick revision

  • Principles (Karl and Willig 3.3), part one: distributed organisation; in-network processing (aggregation, distributed source coding and compression, distributed and collaborative signal processing, mobile code); adaptive fidelity and accuracy.
  • TinyDB aggregates: initializer, merging function (commutative, associative), evaluator; AVERAGE's partial state is (SUM, COUNT); MEDIAN does not aggregate.
  • Aggregation is useful only when it does not hinder collaboration; alarms are not averaged.
  • The sink merges the last records, evaluates, and starts each epoch.
  • Worked tree of 15: 33 transmissions without aggregation, 15 with; average 463 / 15, about 30.8667; averaging averages gives 31.0370, wrong.

Test yourself

1. What is in-network processing, and name its four forms. Processing data inside the network rather than only at the edge, so less is transmitted: aggregation; distributed source coding and compression; distributed and collaborative signal processing; mobile code or agents.

2. How does TinyDB compute an AVERAGE inside the network? Each node's reading x becomes the partial state record (x, 1); records are merged by adding the sums and adding the counts; at the sink the evaluator divides SUM by COUNT.

3. Why can a node not simply forward the average of its own and its children's values? Because that gives equal weight to subtrees of different sizes. The correct method carries the count with the sum. In the worked example the correct average is about 30.8667 and averaging averages gives 31.0370.

4. A tree has 3 nodes at depth 1, 6 at depth 2 and 6 at depth 3. How many transmissions does one round cost with and without aggregation? Without: 3 × 1 + 6 × 2 + 6 × 3 = 33. With: one record per node, 15.

5. What is adaptive fidelity? Give an example. Deliberately trading the accuracy of results against energy and changing the trade as conditions change: sample rarely when all is quiet and often when a fire is suspected.

Contents This chapter on its own page

munotes.in111

Chapter Twenty

Design Principles: Data Centricity, Location, Activity and Heterogeneity

Syllabus topic Module 1, "Introduction and Overview of WSNs: Design principles for WSNs"

In one line

A sensor network should ask for data by what it is and where it is, not by which node holds it; it should use what it knows about positions, activity and its own mixture of nodes; and its protocol layers should share information rather than hide it.

In the wording a student can write in an examination: the design principles of WSNs continue with (4) data centricity: communication is addressed to data described by its attributes (type, value, location, time) rather than to node addresses, with interests, publish/subscribe and database-style queries as implementation options; (5) exploiting location information: using node positions for routing, scoping and addressing; (6) exploiting activity patterns: long quiet periods punctuated by bursts; (7) exploiting heterogeneity: giving demanding roles to better-equipped nodes; and (8) component-based protocol stacks and cross-layer optimisation, instead of strict layering.

Principle 4: data centricity

Address-centric against data-centric

The Internet is address-centric: every host has an address, and communication is between two addresses. That makes sense when the user wants to talk to a particular machine. Karl and Willig explain why it does not make sense in a sensor network: nodes are "deployed redundantly to protect against node failures or to compensate for the low quality of" a single node's own sensing equipment, so "the identity of the particular node supplying data becomes irrelevant. What is important are the answers and values themselves, not which node has provided them."

A data-centric network is organised around the data instead. The user asks for "the average temperature in a given location area, as opposed to requiring temperature readings from individual nodes", in Karl and Willig's example, or sets a condition such as "raise an alarm if temperature exceeds a threshold". The network finds whichever nodes can answer.

Naming data by attributes

To route by data, the data needs names. The protocol that made this concrete, directed diffusion, describes itself as "data-centric in that all communication is for named data", and names data by attribute-value pairs. Its authors' own example is a vehicle-tracking task, sent into the network as an interest:

AttributeValueMeaning
typewheeled vehicledetect vehicle location
interval20 mssend events every 20 ms
duration10 sfor the next 10 s
rect[-100, 100, 200, 400]from sensors within the rectangle

A sensor that detects such a vehicle names its data the same way, with its type, its instance, its location and the time, and the network delivers data whose attributes match an interest. How diffusion then routes by these interests (gradients and reinforcement) is [Directed Diffusion and Rumour Routing].

The survey makes the same point in its account of the application layer: queries are "generally not issued to particular nodes"; "the locations of the nodes that sense temperature higher than 70°F" is attribute-based, "temperatures read by the nodes in region A" is location-based.

munotes.in112

Design Principles: Data Centricity, Location, Activity and Heterogeneity

Three ways to build a data-centric network

  • Interests routed through the network, as in directed diffusion: the interest spreads towards the data, and data flows back along the paths it set up.
  • Publish and subscribe: consumers subscribe to a description of the data they want; producers publish data; the network matches the two, so neither needs to know the other.
  • The network as a database: the user writes a declarative query (SELECT AVG(temperature) FROM sensors WHERE room = 7), and the system plans how to answer it inside the network. This is TinyDB, in [Middleware Approaches, and TinyDB].

Worked example: matching data against an interest

A vineyard sink wants to know about hot spots in the north-west quarter of a 100 m field: temperature above 30 degrees, from the rectangle x from 0 to 50 and y from 50 to 100. Eight nodes have readings. A data-centric network forwards only what matches; the program applies the interest to each reading.

# Data-centric matching: which readings satisfy an attribute-based interest?
interest = {"type": "temperature", "min_value": 30.0,
            "rect": (0, 50, 50, 100)}                      # x0, x1, y0, y1

readings = [  # node, type, value, x, y
    (1, "temperature", 31.5, 12, 88), (2, "temperature", 28.0, 40, 70),
    (3, "moisture",    34.0, 20, 60), (4, "temperature", 33.2, 45, 95),
    (5, "temperature", 36.1, 70, 90), (6, "temperature", 30.4, 5, 55),
    (7, "temperature", 29.9, 30, 80), (8, "temperature", 32.0, 48, 40),
]

def matches(r, want):
    node, kind, value, x, y = r
    x0, x1, y0, y1 = want["rect"]
    return (kind == want["type"] and value > want["min_value"]
            and x0 <= x <= x1 and y0 <= y <= y1)

hits = [r for r in readings if matches(r, interest)]
for node, kind, value, x, y in hits:
    print("node %d reports %s %.1f at (%d, %d)" % (node, kind, value, x, y))
print("%d of %d readings match; the other %d are never sent"
      % (len(hits), len(readings), len(readings) - len(hits)))
node 1 reports temperature 31.5 at (12, 88)
node 4 reports temperature 33.2 at (45, 95)
node 6 reports temperature 30.4 at (5, 55)
3 of 8 readings match; the other 5 are never sent

What the matching did. Node 3 is in the rectangle but reports moisture, the wrong type. Node 5 is hot but outside the rectangle. Nodes 2 and 7 are in the rectangle but too cool; node 7's 29.9 fails "above 30" by a tenth of a degree. Node 8 is hot but south of y = 50. The user never needed to know any node's number; the answer names nodes only because the data carries its source.

munotes.in113

Design Principles: Data Centricity, Location, Activity and Heterogeneity

Principle 5: exploit location information

Sensor data is usually meaningless without where, so nodes need their positions anyway ([Time Synchronisation and Localisation]). Once they have them, location can do more work:

  • Location-based routing: forward towards the destination's position instead of looking up a route ([Geographic Routing: Greedy Forwarding and GPSR]).
  • Location-based addressing and scoping: send a query to "the nodes in region A", as the interest above did, so that nodes elsewhere never hear it. This is what makes a query's cost scale with the region instead of the network ([Figures of Merit: Scalability, Robustness and Measuring a Network]).
  • Location-aware clustering and sleeping: nodes that are close together see the same thing, so only some of them need to be awake.

Principle 6: exploit activity patterns

Karl and Willig describe the typical traffic of a sensor network: "very low data rates over a large timescale, but can have very bursty traffic when something happens (a phenomenon known from real-time systems as event showers or alarm storms)". Months of quiet can alternate with seconds of intense activity.

Designing for this pattern means two things at once:

  • For the quiet: sleep as much as possible, because there is almost nothing to send. Duty-cycling MAC protocols ([Duty Cycling: Preamble Sampling, B-MAC and X-MAC]) exist for this.
  • For the storm: when an event happens, dozens of nodes detect it at once and all want to report. The network must suppress duplicate reports, aggregate them, and wake up enough capacity to carry what matters, without collapsing under collisions.

A design tuned only for the average load fails exactly when the network is needed.

Principle 7: exploit heterogeneity

Not all nodes need be the same. Some may have larger batteries or mains power, faster processors, a second, longer-range radio, or better sensors. A heterogeneous design gives each role to the node suited to it: cluster heads on the mains-powered nodes (the office building in [The Architectural Elements of a Sensor Network]), gateways on the nodes with a cellular radio, heavy computation on the nodes with processors to spare. Even in a network of identical nodes, heterogeneity appears over time as batteries run down unevenly, and a good protocol shifts work away from the weakest nodes.

Principle 8: component-based stacks and cross-layer optimisation

Component-based stacks. Instead of one fixed protocol stack, a node's software is built from small components with clear interfaces, chosen and wired together for each application. TinyOS is built this way ([nesC: Modules, Configurations, Interfaces and Wiring]), and it lets a node carry exactly the protocols its application needs and nothing more.

munotes.in114

Design Principles: Data Centricity, Location, Activity and Heterogeneity

Cross-layer optimisation. The layered model of [The Sensor Network Protocol Stack and Its Three Planes] keeps each layer ignorant of the others. In a sensor network that ignorance is expensive. Karl and Willig note that the simplicity sensor nodes need "may also require breaking with conventional layering rules for networking software, since layering abstractions typically cost time and space". Examples of layers sharing information:

  • The MAC tells routing which links have been reliable, and routing prefers them.
  • Routing tells the MAC which neighbours it will send to, so only those need to stay awake.
  • The application tells the transport and MAC layers how urgent a message is, so an alarm is sent at once while a routine reading waits to be combined with others.

The cost is that cross-layer designs are harder to understand, test and change, because a change in one layer can break another. The principle is to share information across layers deliberately and visibly, not to abolish them.

Distinctions

Address-centricData-centric
Communication is betweenNamed hostsThe user and whatever data matches
A request saysSend me node 145's readingSend me temperatures above 30 in region A
Needs global node identitiesYesNo
SuitsThe InternetSensor networks with redundant nodes
ExampleIPDirected diffusion, publish and subscribe, TinyDB
Strict layeringCross-layer design
Layers know about each otherNoSelected information is shared
Energy and delayHigherLower
Easy to change and testYesHarder

What it does not mean

Data-centric does not mean nodes have no identities. Data may carry its source, as in the worked example. What changes is that requests and routing do not depend on the identities.

Location-based addressing does not need GPS on every node. Positions can be estimated from neighbours, or set when nodes are placed.

Heterogeneity is not a fault to be designed away. It is a resource to be used, and it appears even in identical nodes as their batteries drain unevenly.

Cross-layer optimisation is not the absence of structure. It is layers sharing chosen information through defined interfaces.

Quick revision

  • Principles (Karl and Willig 3.3), part two: data centricity; exploit location; exploit activity patterns; exploit heterogeneity; component-based stacks and cross-layer optimisation.
  • Data-centric: "What is important are the answers and values themselves, not which node has provided them." Data named by attribute-value pairs; implementations: interests (directed diffusion), publish and subscribe, database queries (TinyDB).
  • Directed diffusion's interest: type = wheeled vehicle, interval = 20 ms, duration = 10 s, rect = [-100, 100, 200, 400].
  • Worked matching: 3 of 8 readings match; 5 are never sent.
  • Activity: "event showers or alarm storms" after long quiet; design for both.
  • Cross-layer: layering costs "time and space"; share link quality, neighbour and urgency information across layers.
munotes.in115

Design Principles: Data Centricity, Location, Activity and Heterogeneity

Test yourself

1. Distinguish address-centric and data-centric networking. Address-centric networks communicate between named hosts, so a request names a node. Data-centric networks communicate about data described by its attributes, so a request describes the data wanted (type, value, region, time) and any node holding matching data answers. Data centricity suits sensor networks because redundant nodes make the identity of the answering node irrelevant.

2. How does directed diffusion name data? Give its example. By attribute-value pairs. Its vehicle-tracking interest is type = wheeled vehicle, interval = 20 ms, duration = 10 s, rect = [-100, 100, 200, 400]: detect vehicles, send events every 20 ms for the next 10 s, from sensors within the rectangle.

3. What does it mean for a sensor network to exploit activity patterns? Designing for traffic that is very low most of the time but bursts when an event happens: sleep aggressively in the quiet, and in the burst suppress duplicates, aggregate and avoid collisions so the critical reports get through.

4. Give two examples of cross-layer optimisation in a sensor network. The MAC passing link-quality information to routing so routes use reliable links; routing telling the MAC which neighbours will be used so only they stay awake; the application marking alarms as urgent so lower layers send them at once.

5. Why should a sensor network exploit heterogeneity? Because nodes that have more energy, processing or radio capability can take the demanding roles (cluster head, gateway, heavy computation), which spreads the load to where it can be carried and extends the network's lifetime.

Contents This chapter on its own page

munotes.in116

Chapter Twenty-One

Service Interfaces of a WSN

Syllabus topic Module 1, "Introduction and Overview of WSNs: Service interfaces of WSNs"

In one line

A service interface is the doorway between an application and the sensor network: through it the application says what information it wants, where, when and how accurately, and receives answers and events; for a sensor network it must speak about data and regions, not about addresses and byte streams.

In the wording a student can write in an examination: the service interface of a WSN is the set of operations the network offers its applications. Because WSN interaction is data-centric and scoped, the conventional socket interface (connect to an address, send and receive bytes) is inadequate. A WSN service interface must be able to express: one-shot queries (request and reply) and long-lived interests (subscriptions); asynchronous event notification with conditions; addressing by attribute and by location, with node identities only where needed; scoping in space and time; the accuracy or quality required and the energy and accuracy trade-off; in-network processing such as aggregation; and management and security operations.

Why a new interface is needed

On the Internet, a program opens a socket to an address and a port and exchanges a stream of bytes with one other program. Everything the application wants is expressed as who I talk to and what bytes I send.

A sensor network's user wants to say something else entirely: tell me the average temperature in the north-west quarter, every 10 minutes, to within half a degree, for the next week, or wake me the moment anyone crosses the fence. Karl and Willig put the conclusion in one sentence: departing from an address-centric view of the network "requires new programming interfaces that go beyond the simple semantics of the conventional socket interface and allow concepts like required accuracy, energy/accuracy trade-offs, or scoping."

What a WSN service interface must be able to express

Each requirement below comes straight from how applications use sensor networks, and an answer should give the reason with each.

  1. Simple request and response. A one-shot query, in Karl and Willig's term: what is the temperature in cold store 3 now?, answered once.
  2. Long-lived interests. Standing requests that produce a stream of answers: report every 10 minutes. Karl and Willig note that interactions can be "long-lasting relationships between many sensors and many sinks".
  3. Asynchronous event notification. The application registers a condition and is told when it becomes true, without polling: notify me if any reading exceeds 50 degrees. The condition may combine several readings: two neighbours both detecting motion.
  4. Addressing by attribute and by location. Requests are addressed to data and regions ("temperatures read by the nodes in region A" and "the locations of the nodes that sense temperature higher than 70°F", in the survey's examples), and only rarely to a particular node.
  5. Scoping in space and time. "These interactions can be scoped both in time and in space": only from a region, only for a period, only between certain hours.
  6. Accuracy and quality, and the trade-off with energy. The application should be able to say how accurate an answer must be, or how quickly it must arrive, so that the network can spend no more energy than that needs.
  7. In-network processing. The request should be able to say what processing to do on the way: return the average, the maximum, the count, not every reading ([Design Principles: Distributed Organisation and In-network Processing]).
  8. Changing requirements at run time. Karl and Willig require that "sinks have to have a means to inform the sensors of their requirements at runtime".
  9. Management and security. Setting parameters, reprogramming, turning nodes on and off, and authentication: the survey's Sensor Management Protocol covers exactly these ([The Sensor Network Protocol Stack and Its Three Planes]).
munotes.in117

Service Interfaces of a WSN

Three interfaces that meet the requirements

Publish and subscribe, with interests

The application subscribes by handing the network an interest, a set of attribute-value pairs describing the data it wants, and a function to call when matching data arrives. Sensor nodes publish data named the same way. Directed diffusion's vehicle-tracking interest ([Design Principles: Data Centricity, Location, Activity and Heterogeneity]) is exactly such a subscription: what (wheeled vehicles), how often (every 20 ms), for how long (10 s) and where (a rectangle). Requirements 2, 3, 4 and 5 are met directly.

A declarative query

The application writes what it wants in a query language and lets the system decide how to get it. TinyDB's queries look like SQL with a sampling clause, for example a query ending in SAMPLE PERIOD 1s FOR 10s asks for readings every second for ten seconds. The system plans the sampling, the routing and the aggregation. Requirements 1, 2, 5 and 7 are met, and the application needs no knowledge of the network at all. [Middleware Approaches, and TinyDB] gives the language.

Events and scripts: the survey's SQTL

The survey describes the Sensor Query and Tasking Language (SQTL), which supports three kinds of event, "defined by keywords receive, every," and expire: receive, generated when a node receives a message; every, occurring periodically when a timer times out; and expire, occurring when a timer has expired. A node that receives a message intended for it that contains a script executes the script. This is an interface that sends programs into the network, and meets requirement 8 above all.

The survey's other two application-layer proposals are interfaces too: TADAP for spreading a user's interest into the network or letting nodes advertise their data, and SQDDP for issuing queries and collecting replies, where queries are attribute-based or location-based.

munotes.in118

Service Interfaces of a WSN

Worked example: one request, three ways

The request: Every 10 minutes for the next 24 hours, give me the average temperature of the nodes in the north-west quarter of the vineyard; and tell me at once if any node there reads above 40 degrees.

As a subscription (interest), in attribute-value pairs:

AttributeValue
typetemperature
rect[0, 50, 50, 100]
interval10 min
duration24 h
aggregateaverage

and a second interest with type = temperature, rect as above, and a condition value greater than 40, delivered as an event.

As a declarative query, in TinyDB's style:

SELECT AVG(temperature) FROM sensors
WHERE x < 50 AND y > 50
SAMPLE PERIOD 10 min FOR 24 h

with a second query for the alarm condition, whose answers arrive only when a reading exceeds 40.

As a socket-style exchange, the application would have to learn which node numbers are in the north-west quarter, open a connection to each, ask each for its reading every 10 minutes, average them itself, and poll each one continually for the alarm. Every one of those steps is something the data-centric interfaces do inside the network, and every one costs radio transmissions if done from outside.

The comparison is the answer to why not sockets?: the socket version is longer, knows node identities it should not need, cannot tell the network to aggregate, and polls where it should be notified.

Distinctions

Socket interfaceWSN service interface
AddressesA host and a portData by attribute and by location
InteractionA byte stream between two programsQueries, subscriptions and events
Scoping in space and timeNoneBuilt in
Accuracy and energyNot expressibleExpressible
In-network processingNoneAggregation and conditions
One-shot querySubscriptionEvent notification
AnswersOnceRepeatedly, on a scheduleWhen a condition becomes true
ExampleTemperature now in cold store 3Average every 10 minutesAny reading above 40

What it does not mean

A service interface is not the radio's interface. It is between the application and the whole network, above every protocol layer.

Data-centric interfaces do not forbid naming a node. Sometimes a node must be addressed (to reprogram it, or to read a particular instrument); the interface should allow it without making it the normal way.

Declarative does not mean slow or wasteful. Saying what rather than how lets the system choose the cheapest how, including aggregation inside the network.

Quick revision

  • Service interface: how applications ask the network for information and receive it.
  • Sockets fail: the network needs "new programming interfaces that go beyond the simple semantics of the conventional socket interface" (Karl and Willig).
  • Must express: one-shot queries; long-lived interests; asynchronous events with conditions; addressing by attribute and location; scoping in space and time; accuracy and the energy trade-off; in-network processing; run-time changes; management and security.
  • Interfaces: publish and subscribe with interests; declarative queries (TinyDB, SAMPLE PERIOD); events and scripts (SQTL: receive, every, expire); the survey's TADAP and SQDDP.
munotes.in119

Service Interfaces of a WSN

Test yourself

1. What is a service interface of a WSN, and why is the socket interface not suitable? It is the set of operations through which applications request information from the network and receive results and events. Sockets address a host and exchange bytes, while WSN applications need to request data by attribute and location, scope requests in space and time, state accuracy, ask for aggregation and be notified of events.

2. List six things a WSN service interface must be able to express. Any six of: one-shot queries; long-lived subscriptions; asynchronous event notification with conditions; addressing by attribute and location; scoping in space and time; required accuracy and the energy trade-off; in-network processing; changes of requirements at run time; management and security.

3. What are SQTL's three kinds of event? Receive (a node receives a message), every (a timer times out periodically) and expire (a timer has expired). A node receiving a message intended for it that carries a script executes the script.

4. Express average temperature in region A every 10 minutes for a day as an interest. type = temperature, rect = region A's coordinates, interval = 10 min, duration = 24 h, aggregate = average.

Contents This chapter on its own page

munotes.in120

Chapter Twenty-Two

Gateway Concepts

Syllabus topic Module 1, "Introduction and Overview of WSNs: Gateway concepts" (and the paired practical, "Mote-to-PC Serial Communication Simulation")

In one line

A gateway is the node that joins a sensor network to another network, usually the Internet, translating between two different ways of addressing, routing and talking so that data can cross in both directions.

In the wording a student can write in an examination: a gateway connects a WSN to other networks. It is needed because the two sides differ in protocols (low-power radio and WSN routing against IP), in addressing (data-centric naming against IP addresses), and in availability (sleeping nodes against always-on hosts). Gateway concepts cover (1) WSN-to-Internet communication, where a node's data or alarm must reach an Internet host; (2) Internet-to-WSN communication, where an Internet user queries or tasks the WSN; and (3) WSN tunnelling, where two separate WSNs are joined through the Internet as if they were one.

Why a gateway is needed

If a sensor network spoke the Internet's protocols natively, every node could simply be a host. For most of the subject's history it could not, and even now there are reasons the two sides differ.

  1. Different protocols. The nodes run a low-power radio and a routing protocol built for the field; the outside world runs IP, TCP and HTTP. Something must translate.
  2. Different addressing. Inside the WSN, requests are data-centric (temperature in region A); outside, everything is addressed to a host. A node cannot be expected to know Internet addresses, and an Internet user should not need to know node numbers.
  3. Different availability. An Internet host answers whenever it is asked. A sensor node is asleep most of the time and may take seconds to minutes to answer; a TCP connection held open to a node would waste its battery.
  4. A boundary. Security, access control and accounting belong at the point where the WSN meets the outside world, and so does caching: a gateway can answer many users from one set of readings.

WSN to Internet

The situation. A node detects an event and the report must reach someone outside: an email to the farm manager, a message to a monitoring server, a record in a cloud database.

The problems, and how a gateway solves them.

  • Finding the gateway. The node does not know where the gateway is. In practice the gateway is (or sits beside) the sink, and the collection routing already leads every node towards it. With several gateways, the routing leads to the nearest.
  • Naming the destination. The node cannot hold Internet addresses for every possible receiver. So the node reports data-centrically (alarm, zone 7), and the gateway maps that to Internet destinations according to its configuration: this kind of alarm goes to this server and this person.
  • Translating the message. The gateway turns the WSN packet into an Internet message (an HTTP request, an email, a message to a broker), adding what the Internet side needs and the node could not afford to send, such as full timestamps and identifiers.
munotes.in121

Gateway Concepts

Internet to WSN

The situation. A user on the Internet wants something from the sensor network: the latest readings, or a new task.

The problems, and how a gateway solves them.

  • Which gateway to ask. The user must know an Internet address for the WSN. The gateway presents one: a web page, a web service, or a named endpoint.
  • Expressing the request. The user's request (a web request for average temperature, north-west quarter) must be turned into the WSN's own terms: an interest, a query or a task, sent into the field. The gateway is therefore an application-level gateway: it understands the meaning of requests, not only their packets.
  • Waiting for sleeping nodes. A request may take a long time to be answered inside the WSN. The gateway can answer from its cache of recent readings, and can collect one answer for many users, so the nodes are asked once.
  • Protecting the field. The gateway can refuse, rate-limit or authenticate requests, so that a busy website does not drain the batteries of a forest.

WSN tunnelling

The situation. Two sensor networks are far apart, a sensor field on each bank of a river, or two buildings of one campus, and should behave as one network. No radio link joins them.

The idea. Each island has a gateway on the Internet. When a packet in one island is addressed to something in the other, its gateway encapsulates the whole WSN packet inside an Internet packet and sends it to the other gateway, which unwraps it and injects it into its own island. The Internet is used as a long virtual link, a tunnel. To the WSN protocols the two islands look like one network with one long hop.

The humblest gateway: a base station mote on a serial cable

In a laboratory, and in the practical, the gateway is a base station mote connected to a PC by a USB or serial cable. The mote listens to the radio and passes what it hears to the PC; the PC runs the program that stores, displays or forwards the data.

TinyOS ships this as an application. Its own source comment says it plainly: "BaseStationP bridges packets between a serial channel and the radio." Packets from the radio are passed to the serial port; packets from the PC are sent out over the radio. Messages going from serial to radio are tagged with the group identity compiled into the base station, and radio messages are filtered by that same group, so a base station hears only its own network.

munotes.in122

Gateway Concepts

How the bytes cross the cable

A serial line carries a stream of bytes, with nothing to mark where one packet ends and the next begins. TinyOS's serial stack, specified in TEP 113, solves this in three levels: encoding, framing and a protocol level with acknowledgements and a CRC. Its framing uses the same encoding as the HDLC protocol:

  • 0x7e is reserved as the frame delimiter: it marks the start and end of every packet.
  • 0x7d is reserved as the escape byte.
  • If the data itself contains 0x7e or 0x7d, the sender sends 0x7d followed by the byte XORed with 0x20. TEP 113's example: "0x7e becomes 0x7d 0x5e".
  • The receiver, on seeing 0x7d, drops it and XORs the next byte with 0x20 to recover the original.

A CRC (cyclic redundancy check) at the protocol level lets the receiver detect a packet damaged on the cable.

Worked example: framing a packet

The program frames a five-byte payload that happens to contain both reserved bytes, then decodes it again. (The CRC is left out so that the framing can be seen on its own.)

# HDLC-style framing as TinyOS's serial stack uses it (TEP 113):
# 0x7e marks frame boundaries, 0x7d escapes, an escaped byte is XORed with 0x20.
FLAG, ESC = 0x7E, 0x7D

def frame(payload):
    out = [FLAG]
    for b in payload:
        if b in (FLAG, ESC):
            out += [ESC, b ^ 0x20]
        else:
            out.append(b)
    return out + [FLAG]

def unframe(stream):
    body, escaped = [], False
    for b in stream[1:-1]:                  # drop the two delimiters
        if escaped:
            body.append(b ^ 0x20)
            escaped = False
        elif b == ESC:
            escaped = True
        else:
            body.append(b)
    return body

payload = [0x12, 0x7E, 0x34, 0x7D, 0x56]
line = frame(payload)
print("payload:", " ".join("%02x" % b for b in payload))
print("on the wire:", " ".join("%02x" % b for b in line))
print("decoded:", " ".join("%02x" % b for b in unframe(line)))
print("round trip correct:", unframe(line) == payload)
payload: 12 7e 34 7d 56
on the wire: 7e 12 7d 5e 34 7d 5d 56 7e
decoded: 12 7e 34 7d 56
round trip correct: True

What happened to the two reserved bytes. 0x7e in the data became the pair 0x7d 0x5e (0x7e XOR 0x20 is 0x5e), exactly TEP 113's example, and 0x7d became 0x7d 0x5d. Only the two real delimiters, at the ends, appear as bare 0x7e, so the receiver always knows where a frame starts and stops. The frame is two bytes longer for the delimiters and two more for the escapes: framing costs a little size for certainty.

A real gateway: Great Duck Island

The deployment of [Applications of Wireless Sensor Networks] shows the full chain. Each patch of motes had a gateway node; a longer radio link, the transit network, carried data from the gateway to a base station with a database and a wide-area connection; and remote users read a replica of that database over the Internet. Every layer kept some persistent storage, because a disconnection could happen at any level. That is Internet-to-WSN communication done through a cache, which is how most real systems do it.

munotes.in123

Gateway Concepts

Distinctions

SinkGatewayRouter inside the WSN
IsWhere the WSN's data is collectedThe bridge to another networkA node forwarding packets
SpeaksThe WSN's protocolsBoth sides' protocolsThe WSN's protocols
TranslatesNoYes: protocols, addresses, requestsNo
Often combined withThe gatewayThe sinkEvery node
WSN to InternetInternet to WSNWSN tunnelling
StartsInside the fieldOutside, with a userIn one island, bound for another
Gateway's key jobMap data to Internet destinationsTranslate requests, cache answersEncapsulate and unwrap WSN packets
ExampleAn alarm emailed to the managerA web page showing the average moistureTwo fields on either side of a river as one network

What it does not mean

A gateway is not only a router. A router forwards packets of one protocol; a gateway translates between protocols and between ways of addressing, often at the level of meaning.

Internet-to-WSN does not mean a TCP connection to each node. The gateway answers for the network, from a cache or by collecting one answer for many users.

Tunnelling does not make the Internet part of the WSN. The Internet only carries the WSN's packets, wrapped, between two gateways.

The serial escape is not encryption. It only keeps the reserved bytes from being mistaken for delimiters.

Quick revision

  • Gateway: joins the WSN to another network; needed for different protocols, different addressing, different availability, and a boundary for security and caching.
  • WSN to Internet: node reports data-centrically to the gateway; gateway maps to Internet destinations and translates.
  • Internet to WSN: gateway gives the WSN an Internet address, translates requests into queries or interests (an application-level gateway), caches answers, protects the field.
  • WSN tunnelling: gateways encapsulate WSN packets in Internet packets to join two WSN islands.
  • Base station mote: TinyOS BaseStation "bridges packets between a serial channel and the radio", filtered by group identity.
  • TEP 113 framing: 0x7e delimiter, 0x7d escape, escaped byte XOR 0x20 (0x7e becomes 0x7d 0x5e), plus a CRC.

Test yourself

1. Why does a WSN need a gateway to reach the Internet? Because the two sides differ in protocols (low-power radio and WSN routing against IP), addressing (data-centric names against host addresses) and availability (sleeping nodes against always-on hosts), and because security, access control and caching belong at the boundary.

munotes.in124

Gateway Concepts

2. Explain Internet-to-WSN communication and the problems a gateway solves in it. An Internet user requests data or tasks the WSN. The gateway gives the WSN an Internet address, translates the request into the WSN's own query or interest, answers from a cache or gathers one answer for many users because nodes sleep, and protects the field by authenticating and rate-limiting requests.

3. What is WSN tunnelling? Joining two separate WSN islands through the Internet: each island's gateway encapsulates WSN packets bound for the other island in Internet packets, and the other gateway unwraps and injects them, so the Internet acts as one long virtual link.

4. How does TinyOS frame a packet on a serial line, and how is a data byte of 0x7e sent? With HDLC-style framing: 0x7e marks the start and end of each frame and 0x7d is the escape byte. A data byte of 0x7e is sent as 0x7d followed by 0x7e XOR 0x20, which is 0x5e.

5. What does the TinyOS BaseStation application do? It bridges packets between the radio and the serial port: radio packets of its group are passed to the PC, and packets from the PC are sent over the radio, tagged with the base station's group identity.

Contents This chapter on its own page

munotes.in125

Chapter Twenty-Three

Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP

Syllabus topic Module 1, "Introduction and Overview of WSNs: Gateway concepts", as widened by the paper's row 1 ("Wireless sensor networks form the backbone of modern IoT")

In one line

6LoWPAN lets IPv6 packets travel over tiny 802.15.4 frames by compressing their headers and fragmenting them, RPL builds routes towards a border router in such a lossy network, and CoAP gives each sensor a small web interface, so a sensor node can be an ordinary Internet host.

In the wording a student can write in an examination: WSNs become part of the Internet of Things (IoT) when their nodes speak Internet protocols. 6LoWPAN (IPv6 over Low-Power Wireless Personal Area Networks, RFC 4944 and RFC 6282) is an adaptation layer between IPv6 and the 802.15.4 MAC that provides header compression and fragmentation and reassembly. RPL (the IPv6 Routing Protocol for Low-Power and Lossy Networks, RFC 6550) builds a Destination-Oriented Directed Acyclic Graph (DODAG) rooted at the border router, using DIO, DIS and DAO messages and a Rank computed by an Objective Function. CoAP (the Constrained Application Protocol, RFC 7252) is a web transfer protocol over UDP with a 4-byte fixed header, Confirmable and Non-confirmable messages and the methods GET, PUT, POST and DELETE.

Why put IP on a sensor node at all

A translating gateway ([Gateway Concepts]) works, but every application needs its own translation, and every sensor network speaks its own dialect. If the nodes speak IP, then any Internet host can reach any node through an ordinary router, the border router, and all the Internet's tools (addressing, security, web protocols) work end to end. The obstacles were size and energy, and the three protocols in this chapter are the answers to them.

The problem 6LoWPAN solves: a 1,280-octet packet in a 127-octet frame

IPv6 requires every link to carry packets of at least 1,280 octets. RFC 4944 sets out what an 802.15.4 frame can actually carry:

StageOctets leftWhere the octets go
Largest 802.15.4 physical packet127aMaxPHYPacketSize
After the largest MAC frame overhead102aMaxFrameOverhead is 25
After link-layer security at its largest81AES-CCM-128 adds 21
After an uncompressed IPv6 header41the IPv6 header is 40
After a UDP header33UDP uses 8
Or after a TCP header instead21TCP uses 20

So in the worst case, as RFC 4919 puts it, a UDP application has 33 octets for its data out of a 127-octet frame, and a full IPv6 packet does not fit in one frame at all. Two things are needed, and 6LoWPAN provides both in an adaptation layer below IP.

Fragmentation and reassembly. A packet larger than a frame is cut into fragments, each with a small header saying which packet it belongs to and where it goes, and put back together at the other end. RFC 4944 requires this layer "below IP", in keeping with IPv6's own rules.

munotes.in126

Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP

Header compression. Most of an IPv6 header can be worked out from the 802.15.4 header (the addresses are often derived from the link-layer addresses) or takes an expected value (the version, the usual hop limits). Compression leaves those fields out. RFC 8376's summary of the result: in the best case the IPv6 header for link-local communication is reduced to 2 bytes, for global communication to 3 bytes in the most extreme case, and in more practical situations to 11 bytes (one address prefix compressed) or 19 bytes (both).

Worked: what compression buys. Take the practical case of a 19-byte compressed header in place of 40. In the worst-case frame of 81 octets, the UDP application then has 81 - 19 - 8 = 54 octets instead of 33, and a typical sensor reading of a few tens of octets fits in a single frame instead of needing fragments. Every fragment avoided is a transmission, and an energy cost, saved.

RPL: routing towards the border router

Sensor networks running IP still need routing, and ordinary Internet routing protocols expect stable links and plenty of memory. RPL was designed for low-power and lossy networks, and its structure is the collection tree of [Design Principles: Distributed Organisation and In-network Processing] made into a standard.

The DODAG. RPL organises the network as a Destination-Oriented Directed Acyclic Graph: RFC 6550 defines it as "a DAG rooted at a single destination", the DODAG root, which is usually the border router. Every node has one or more parents closer to the root, so there are no loops and every path leads up to the root.

Rank. Each node has a Rank, which RFC 6550 describes as "a scalar representation of the location of that node within a DODAG Version"; it grows with distance from the root and "is used to avoid and detect loops". How Rank is computed is left to an Objective Function, which is how an application chooses what "best parent" means: fewest hops, best link quality, least energy.

The messages.

  1. DIO (DODAG Information Object): sent by the root and then by every node, advertising the DODAG and the sender's Rank. A node hearing DIOs chooses its parents and computes its own Rank, then sends its own DIOs. This is how the graph grows outward from the root.
  2. DIS (DODAG Information Solicitation): sent by a node that wants to join and has heard no DIO; it asks neighbours to send one.
  3. DAO (Destination Advertisement Object): sent upward, advertising the node's address so that routes can be built down from the root towards the node.

Trickle. RPL does not send DIOs at a fixed rate. It uses the Trickle algorithm (RFC 6206): when nothing changes, the interval between DIOs doubles, so a stable network falls almost silent; when an inconsistency is detected, the interval is reset to its minimum and the network repairs itself quickly. It is an application of the principle [Design Principles: Data Centricity, Location, Activity and Heterogeneity] states as exploiting activity patterns: a network that has nothing to say should cost nothing to run.

munotes.in127

Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP

Traffic and modes. RFC 6550 names three traffic flows: multipoint-to-point (from the nodes to a central point, the main case), point-to-multipoint (from the root to the nodes) and point-to-point. For routes downward it has two modes: in storing mode nodes keep routing tables for the nodes below them; in non-storing mode only the root keeps them, and downward packets carry their route, which suits nodes with little memory.

CoAP: a web protocol for a sensor

HTTP over TCP is too heavy for a node with 10 kB of RAM: TCP's connection set-up and its 20-byte header alone are costly, and the text headers of HTTP are larger than a sensor's whole reading. CoAP keeps the web's model (resources named by URIs, requests with methods, responses with codes) and makes everything small.

  • Over UDP, not TCP, with the CoAP default port 5683.
  • A fixed-size 4-byte header, followed by a token, options and the payload.
  • Methods GET, PUT, POST and DELETE, which RFC 7252 says behave in a way "similar" to HTTP's.
  • Reliability without TCP: a Confirmable message must be acknowledged and is retransmitted until it is; a Non-confirmable message is sent once, for readings where the next one will do.
  • URIs of the form coap://host:port/path, such as RFC 7252's own example, coap://example.com:5683/~sensors/temp.xml.

So a web application can ask a node directly: a GET on the node's temperature resource, answered in one small UDP datagram, fitting in one 802.15.4 frame.

The stack, put together

LayerInternet host6LoWPAN sensor node
ApplicationHTTPCoAP
TransportTCPUDP
NetworkIPv6 with ordinary routingIPv6 with RPL
AdaptationNone needed6LoWPAN: compression and fragmentation
MAC and physicalEthernet, Wi-FiIEEE 802.15.4

The border router sits between the two columns. Unlike a translating gateway, it does not need to understand the application: it forwards IPv6 packets, expanding compressed headers on the way out and compressing them on the way in.

The adaptation layer, computed

The arithmetic above deserves to be run rather than quoted, because one of its conclusions is the opposite of the obvious one.

# What header compression is worth, in fragments and in transmissions.
PHY, MAC_OVERHEAD, SECURITY = 127, 25, 21    # RFC 4944's worst case, in octets
UDP, FRAG1, FRAGN = 8, 4, 5                  # UDP, and 6LoWPAN's two fragment headers
IPV6_MTU = 1280

frame = PHY - MAC_OVERHEAD - SECURITY
HDRS = ((40, "uncompressed, 40"), (19, "both prefixes compressed, 19"),
        (11, "one prefix compressed, 11"), (3, "global, best case, 3"),
        (2, "link local, best case, 2"))

def frames_for(hdr, app_octets):
    """Frames needed to carry app_octets of application data, RFC 4944."""
    if hdr + UDP + app_octets <= frame:
        return 1
    first = frame - FRAG1 - hdr - UDP            # application octets in fragment one
    rest = frame - FRAGN                         # and in each one after it
    return 1 + -(-(app_octets - first) // rest)

print("A worst-case 802.15.4 frame leaves %d octets above link-layer security." % frame)
print()
print("What the compressed header buys, in payload and in frames:")
print("  IPv6 header                    payload in    a full %d octet" % IPV6_MTU)
print("                                 one frame     packet needs")
for hdr, label in HDRS:
    payload = frame - hdr - UDP
    full = frames_for(hdr, IPV6_MTU - 40 - UDP)
    print("  %-30s %6d octets %8d frames" % (label, payload, full))
print()
print("Read that last column again. Compression barely changes it, because it only")
print("shortens the FIRST fragment: every fragment after it carries raw payload and is")
print("the same size whatever the header was. Compression is not a way to fragment less")
print("when the packet is large.")
print()
print("Where it does change everything is the packet a sensor actually sends:")
print("                      frames needed, with an IPv6 header of")
print("  application data    " + "".join("%12s" % ("%d octets" % hdr) for hdr, _ in HDRS))
for app in (20, 33, 40, 54, 60, 71, 80):
    print("  %6d octets     " % app
          + "".join("%9d fr" % frames_for(hdr, app) for hdr, _ in HDRS))
print()
print("Compression does not make the packet smaller. It makes the payload that fits in")
print("one frame larger, and for a reading of a few tens of octets that is the whole")
print("difference between one transmission and two.")

# What a fragment costs, and why losing one is worse than it looks.
print()
print("Why a lost fragment is worse than a lost packet:")
for n in (2, 4, 8, 16):
    for per_frag in (0.01, 0.05):
        ok = (1 - per_frag) ** n
        print("  %2d fragments at %2.0f%% loss each: the packet arrives %5.1f%% of the time"
              % (n, per_frag * 100, ok * 100))
print("  and a packet that loses one fragment loses all of them, because there is no")
print("  selective repeat below IP: every fragment is sent again.")

# CoAP against HTTP, in bytes on the wire for one reading.
print()
print("Asking a sensor for a reading, in octets on the wire:")
COAP_FIXED, COAP_TOKEN = 4, 2                 # RFC 7252's fixed header, plus a short token
uri = "/s/t"
coap_req = COAP_FIXED + COAP_TOKEN + 1 + len(uri)
coap_rsp = COAP_FIXED + COAP_TOKEN + 1 + 4
http_req = len("GET /sensors/temperature HTTP/1.1\r\nHost: node.example\r\n\r\n")
http_rsp = len("HTTP/1.1 200 OK\r\nContent-Type: text/plain\r\nContent-Length: 4\r\n\r\n21.4")
print("  CoAP over UDP: request %3d, response %3d, total %3d" % (coap_req, coap_rsp, coap_req + coap_rsp))
print("  HTTP over TCP: request %3d, response %3d, total %3d" % (http_req, http_rsp, http_req + http_rsp))
print("  a factor of %.1f, before TCP's three way handshake and its acknowledgements,"
      % ((http_req + http_rsp) / float(coap_req + coap_rsp)))
print("  and the CoAP exchange fits in one frame in each direction while the HTTP one")
print("  does not fit in the worst-case frame at all.")
munotes.in128

Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP

A worst-case 802.15.4 frame leaves 81 octets above link-layer security.

What the compressed header buys, in payload and in frames:
  IPv6 header                    payload in    a full 1280 octet
                                 one frame     packet needs
  uncompressed, 40                   33 octets       17 frames
  both prefixes compressed, 19       54 octets       17 frames
  one prefix compressed, 11          62 octets       17 frames
  global, best case, 3               70 octets       17 frames
  link local, best case, 2           71 octets       17 frames

Read that last column again. Compression barely changes it, because it only
shortens the FIRST fragment: every fragment after it carries raw payload and is
the same size whatever the header was. Compression is not a way to fragment less
when the packet is large.

Where it does change everything is the packet a sensor actually sends:
                      frames needed, with an IPv6 header of
  application data       40 octets   19 octets   11 octets    3 octets    2 octets
      20 octets             1 fr        1 fr        1 fr        1 fr        1 fr
      33 octets             1 fr        1 fr        1 fr        1 fr        1 fr
      40 octets             2 fr        1 fr        1 fr        1 fr        1 fr
      54 octets             2 fr        1 fr        1 fr        1 fr        1 fr
      60 octets             2 fr        2 fr        1 fr        1 fr        1 fr
      71 octets             2 fr        2 fr        2 fr        2 fr        1 fr
      80 octets             2 fr        2 fr        2 fr        2 fr        2 fr

Compression does not make the packet smaller. It makes the payload that fits in
one frame larger, and for a reading of a few tens of octets that is the whole
difference between one transmission and two.

Why a lost fragment is worse than a lost packet:
   2 fragments at  1% loss each: the packet arrives  98.0% of the time
   2 fragments at  5% loss each: the packet arrives  90.2% of the time
   4 fragments at  1% loss each: the packet arrives  96.1% of the time
   4 fragments at  5% loss each: the packet arrives  81.5% of the time
   8 fragments at  1% loss each: the packet arrives  92.3% of the time
   8 fragments at  5% loss each: the packet arrives  66.3% of the time
  16 fragments at  1% loss each: the packet arrives  85.1% of the time
  16 fragments at  5% loss each: the packet arrives  44.0% of the time
  and a packet that loses one fragment loses all of them, because there is no
  selective repeat below IP: every fragment is sent again.

Asking a sensor for a reading, in octets on the wire:
  CoAP over UDP: request  11, response  11, total  22
  HTTP over TCP: request  57, response  68, total 125
  a factor of 5.7, before TCP's three way handshake and its acknowledgements,
  and the CoAP exchange fits in one frame in each direction while the HTTP one
  does not fit in the worst-case frame at all.
munotes.in129

Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP

Distinctions

Translating gatewayBorder router
Nodes speakTheir own WSN protocolsIPv6 (compressed)
Understands the applicationYes, it mustNo
Adding a new applicationNeeds new translation at the gatewayWorks end to end
munotes.in130

Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP

DIODISDAO
Sent byThe root, then every nodeA node wanting to joinA node, upward
CarriesThe DODAG and the sender's RankA request for a DIOThe node's address, for downward routes
ConfirmableNon-confirmable
AcknowledgedYes, retransmitted until it isNo
SuitsCommands, alarmsPeriodic readings

What it does not mean

6LoWPAN is not a new network layer. It is an adaptation layer below IPv6; the network layer is still IPv6.

IP on the node does not remove the need for a border router. The border router is still the link between the 802.15.4 network and the rest of the Internet, and still the natural place for security.

RPL is not only for collection. Its main flow is multipoint-to-point, but it also builds downward and point-to-point routes.

CoAP is not HTTP compressed. It is a separate protocol over UDP with its own reliability; a proxy can translate between CoAP and HTTP.

Quick revision

  • 6LoWPAN: adaptation layer for IPv6 over 802.15.4, fragmentation and reassembly plus header compression.
  • The squeeze (RFC 4944, RFC 4919): 127 physical, 102 after MAC, 81 after AES-CCM-128, 41 after IPv6, 33 after UDP, 21 after TCP; IPv6 needs 1,280.
  • Compressed IPv6 header (RFC 8376): 2 bytes link-local best case, 3 global extreme, 11 or 19 practical. With 19: 54 octets for UDP data in the worst-case frame.
  • RPL: a DODAG rooted at the border router; Rank from an Objective Function; DIO advertises, DIS solicits, DAO advertises destinations; Trickle timing; storing and non-storing modes; MP2P, P2MP, P2P traffic.
  • CoAP: UDP port 5683, 4-byte header, GET, PUT, POST, DELETE, Confirmable and Non-confirmable messages.
munotes.in131

Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP

Test yourself

1. Why can an IPv6 packet not be sent directly in an 802.15.4 frame, and what does 6LoWPAN do about it? IPv6 needs links to carry 1,280-octet packets, while an 802.15.4 frame is at most 127 octets, 102 after MAC overhead and 81 with the strongest security. 6LoWPAN adds an adaptation layer that fragments packets and reassembles them, and compresses the IPv6 and UDP headers by leaving out fields that can be derived or take expected values.

2. In the worst case, how many octets are left for UDP data in an 802.15.4 frame without compression, and with a 19-byte compressed header? Without: 81 - 40 - 8 = 33. With: 81 - 19 - 8 = 54.

3. Explain how RPL builds its routing structure. The root, usually the border router, sends DIO messages advertising the DODAG. A node hearing DIOs chooses parents with lower Rank, computes its own Rank with the Objective Function, and sends its own DIOs, so the graph grows outward with no loops. A node that hears nothing sends a DIS to ask for a DIO; nodes send DAOs upward so downward routes can be built. Trickle slows DIOs when the network is stable and speeds them when it changes.

4. What are CoAP's main features? A web protocol over UDP (default port 5683) with a fixed 4-byte header, methods GET, PUT, POST and DELETE, URIs as on the web, and its own reliability: Confirmable messages are acknowledged and retransmitted, Non-confirmable ones are not.

5. How does a border router differ from a translating gateway? The border router forwards IPv6 packets between the 6LoWPAN network and the Internet, only compressing and expanding headers, so it does not need to understand applications. A translating gateway must understand each application's messages and convert them between the WSN's protocols and the Internet's.

Contents This chapter on its own page

munotes.in132

Chapter Twenty-Four

Why a Sensor Node Needs an Operating System

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Overview of wireless sensor network operating systems"

In one line

A sensor node needs an operating system to share one tiny processor among many things happening at once (sensing, receiving, forwarding, sending), to hide the hardware behind simple interfaces, and to put everything to sleep whenever possible, all within a few kilobytes of memory.

In the wording a student can write in an examination: an operating system for a WSN node manages concurrency (many flows of events on one microcontroller), provides hardware abstraction for sensors, radio and timers, handles scheduling, memory, communication, power management and reprogramming, and must do so with a tiny footprint. Its design issues are the execution model (event-driven or multithreaded), scheduling, memory management, protocol support, resource sharing, modularity, energy management, dynamic reprogramming, robustness and timing. A desktop operating system cannot be used, because a node has kilobytes of memory, no memory protection and a battery that must last months.

Why not simply program the hardware directly?

A small node could be programmed as one loop: read the sensor, send, sleep, repeat. That works for a node that does only one thing. A real node, even a simple one, does several things whose timing it does not control:

  • a timer says it is time to sample;
  • the ADC finishes a conversion some microseconds later;
  • a packet arrives from a neighbour that must be forwarded;
  • the radio reports that a send has finished, or failed;
  • a new command arrives from the sink.

These are concurrent: they overlap, and any of them can happen while another is being handled. Managing them by hand in one loop quickly becomes unmanageable and wasteful. An operating system provides the structure: a way to respond to each event, to share the processor between them, to schedule deferred work, and to sleep when nothing is left to do.

What the TinyOS designers said a sensor OS must handle

The Berkeley paper that introduced TinyOS in 2000 began by setting out the "requirements that shape the design of network sensor systems". There are five, and they are still the best short answer to why is a sensor OS different?.

  1. Small physical size and low power consumption. Processing, storage and interconnect are "limited and scarce", so the rest of the solution "must be austere and efficient".
  2. Concurrency-intensive operation. The node's main job is "to flow information from place to place with a modest amount of processing on-the-fly, rather than to accept a command, stop, think, and respond". Data is captured, processed and streamed at once, or received and forwarded, with little memory to buffer between flows, and some of the events have real-time requirements.
  3. Limited physical parallelism and controller hierarchy. A desktop spreads work across many controllers and buses. A node has one microcontroller wired directly to its sensors and radio, so the concurrency must all be managed by that one processor.
  4. Diversity in design and usage. Nodes are application-specific and carry only the hardware their application needs, so the software must be modular: it must be easy to assemble "just the software components required", with very efficient interfaces between them.
  5. Robust operation. Nodes are numerous and unattended, and redundancy inside a device is too costly, so each device must be reliable: its components "should be as independent as possible and connected with narrow interfaces".
munotes.in133

Why a Sensor Node Needs an Operating System

Their result was an operating system that, in the paper's words, "fits in 178 bytes of memory".

The design issues a sensor OS must settle

Each of these is a decision, and the systems of [TinyOS: Components, Tasks and the Scheduler] and [Contiki, RIOT and the Other Sensor Operating Systems] settle them differently.

Design issueThe questionWhy it is hard on a sensor node
Execution modelEvents that run to completion, or threads that block?Threads need a stack each; the next chapter weighs the two
SchedulingWho runs next, and can a running job be interrupted?Must be tiny and must let the node sleep as soon as work runs out
Memory managementIs memory allocated statically or at run time?Kilobytes of RAM and no memory protection, so a runaway allocation corrupts everything
Protocol supportHow are the radio, MAC and routing provided?The protocols are part of the application's energy budget, not a fixed layer
Resource sharingHow do tasks share the radio, the ADC and the flash?One of each, used by everything
ModularityHow is a system assembled for one application?Only the components needed should be linked in
Energy managementHow does the OS put the hardware to sleep?The single most important job; idle time must become sleep time automatically
ReprogrammingHow is new code installed in a deployed network?Nodes cannot be collected; code must travel over the radio and should be small
RobustnessWhat happens when one component misbehaves?No protection hardware to contain the damage
TimingCan the OS meet real-time deadlines?Some events (a radio byte, a sample) cannot wait

Reprogramming deserves a word of its own

Contiki's authors put the case: networks of "hundreds or even thousands of nodes" must have code downloaded into them, and "bugs may have to be patched in an operational network", because "it is not feasible to physically collect and reprogram all sensor devices". Sending code costs energy, so the smaller the piece sent, the better. Most embedded systems must send a complete binary image of the whole system; Contiki's answer is to "load and unload individual applications or services at run-time", so only the changed program travels.

munotes.in134

Why a Sensor Node Needs an Operating System

How small is small

The numbers in the two founding papers make the constraint concrete.

  • Contiki's paper describes the typical device of 2004: "8-bit microcontrollers, code memory on the order of 100 kilobytes, and less than 20 kilobytes of RAM". Its own platform, the ESB, had an MSP430 with 2 kilobytes of RAM and 60 kilobytes of ROM, running at 1 MHz.
  • On that platform a Contiki process's whole state was 23 bytes.
  • The first TinyOS fitted in 178 bytes.
  • The Telos mote of [Inside a Sensor Node: The Five Units] was generous by comparison: 10 kB of RAM.

A desktop operating system needs thousands of times more memory just to start.

Worked example: one wake-up of a relay node

A relay node in the vineyard wakes on a timer. In the next few milliseconds, all of this may happen, and the operating system is what keeps it in order.

Time (ms)EventWhat the OS must do
0Timer fires: time to sampleRun the timer's handler; schedule the sampling
0.1Sampling started on the ADCLet the processor do other work, or sleep, while the ADC converts
0.4A neighbour's packet starts arrivingHandle the radio's interrupt at once: the bytes cannot wait
0.6ADC conversion completeRecord the reading; schedule the packet to be built
2.1Neighbour's packet completeSchedule it to be forwarded
2.2Own packet builtHand it to the radio; queue the neighbour's packet behind it
4.8Radio reports own packet sentSend the neighbour's packet
7.5Radio reports forwarded packet sentNothing left to do: power the radio down, put the processor to sleep

Three things in the table are an operating system's whole job in miniature: the radio's interrupt at 0.4 ms could not wait for the sampling to finish, so the OS must let urgent events interrupt; the ADC and the radio worked while the processor did other things, so the OS must let operations start and finish later; and at 7.5 ms the node went back to sleep because the OS knew nothing else was pending.

Distinctions

Desktop operating systemSensor node operating system
MemoryGigabytes, protected, virtualKilobytes, unprotected, physical
Main concernThroughput and responsiveness for usersEnergy, concurrency and footprint
ProcessesMany, isolated from each otherA few components or processes in one address space
Idle timeWastedConverted into sleep
Updating softwareInstall a packageSend code over the radio to thousands of nodes

What it does not mean

A sensor OS is not a small Linux. It usually has no memory protection, no file system in the usual sense and no users, and it is organised around events and energy.

munotes.in135

Why a Sensor Node Needs an Operating System

"178 bytes" is not the size of a whole application. It is the core of the 2000 TinyOS; applications and their components add to it.

An operating system does not make the node faster. It makes it organised and economical: it lets many slow things overlap and sleeps in the gaps.

Quick revision

  • A sensor OS manages concurrency, hardware abstraction, scheduling, memory, communication, power and reprogramming, in a tiny footprint.
  • TinyOS's five requirements (2000): small size and low power; concurrency-intensive operation; limited physical parallelism and controller hierarchy; diversity in design and usage (modularity); robust operation. First TinyOS: 178 bytes.
  • Design issues: execution model, scheduling, memory management, protocol support, resource sharing, modularity, energy management, reprogramming, robustness, timing.
  • Typical 2004 node (Contiki paper): 8-bit MCU, about 100 kB code, under 20 kB RAM; ESB: 2 kB RAM, 60 kB ROM, 1 MHz; a Contiki process state: 23 bytes.
  • Reprogramming over the radio must send as little code as possible; Contiki loads individual programs at run time.

Test yourself

1. Why does a sensor node need an operating system? Because even a simple node handles several concurrent activities (timers, sensor conversions, arriving and departing packets, commands) on one small processor. The OS provides a structure to respond to events, share the processor, defer work, abstract the hardware, manage energy by sleeping whenever possible, and support reprogramming, within a few kilobytes.

2. State the five requirements the TinyOS designers identified for networked sensors. Small physical size and low power consumption; concurrency-intensive operation; limited physical parallelism and controller hierarchy; diversity in design and usage, requiring efficient modularity; robust operation of numerous unattended devices.

3. List the design issues of a WSN operating system. Execution model, scheduling, memory management, communication protocol support, resource sharing, modularity, energy management, dynamic reprogramming, robustness, and meeting timing requirements.

4. Why is dynamic reprogramming important, and how does Contiki make it cheaper? Deployed networks have hundreds or thousands of nodes that cannot be collected, and bugs must be fixed in the field, so code must be sent over the radio, which costs energy. Contiki loads and unloads individual programs at run time, so only the changed program is sent rather than a whole system image.

5. In the relay node's wake-up, what three things did the OS have to make possible? Letting an urgent event (the radio interrupt) interrupt other work; letting slow operations (the ADC conversion, the radio send) start and finish later while the processor did other things; and putting the node to sleep as soon as nothing was pending.

Contents This chapter on its own page

munotes.in136

Chapter Twenty-Five

Event-driven or Multithreaded: The Two Execution Models

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Overview of wireless sensor network operating systems" (and the paired practical, "Understanding TinyOS Architecture and Execution Model")

In one line

In an event-driven system, short handlers run one after another to completion on one shared stack; in a multithreaded system, each activity is a thread with its own stack that can block and be interrupted; sensor systems mostly choose events, because stacks cost memory they do not have.

In the wording a student can write in an examination: an event-driven operating system runs event handlers and tasks that run to completion and cannot block, sharing one stack, with no locking needed because two handlers never run at the same time; it saves memory and suits the event-heavy work of sensor nodes, but long computations must be split up and programs become state machines. A multithreaded operating system gives each thread its own stack, lets threads block and be preempted, and so supports sequential, blocking code, at the cost of a stack per thread (usually over-provisioned) and locking of shared data. TinyOS is event-driven; Contiki has an event-driven kernel with optional per-process threads; Mantis is preemptive multithreaded.

Why this is the first decision

Almost everything else follows from it: how much memory the system needs, how programs are written, how quickly the node can respond to an urgent event, and whether a bug in one activity can freeze all the others. And the constraint that decides it is the one in [Why a Sensor Node Needs an Operating System]: a few kilobytes of RAM.

The multithreaded model

A thread is a sequential flow of execution with its own stack, the memory where its local variables and return addresses live. Threads let a programmer write naturally:

loop:
    wait for the timer            # the thread blocks here
    reading = read the sensor     # blocks until the ADC finishes
    send reading                  # blocks until the radio finishes

While one thread waits, the scheduler runs another; a preemptive scheduler can also interrupt a running thread to run a more urgent one.

The cost, in the Contiki authors' words. "Each thread must have its own stack and because it in general is hard to know in advance how much stack space a thread needs, the stack typically has to be over provisioned." The memory "must be allocated when the thread is created" and "can not be shared between many concurrent threads". And "a threaded concurrency model requires locking mechanisms to prevent concurrent threads from modifying shared resources": if two threads update the same routing table, each must lock it first, and a forgotten lock is a bug that appears only occasionally.

Worked: what stacks cost. Suppose a node has 10 kB of RAM, like the Telos mote, and runs five threads (sensing, radio receive, radio send, routing, the application), each given a 512-byte stack to be safe (an assumed figure; real stack needs vary). The stacks take 5 × 512 = 2,560 bytes, a quarter of the node's RAM, before any data is stored. On Contiki's ESB platform, with 2 kB of RAM, the same five stacks would not fit at all.

munotes.in137

Event-driven or Multithreaded: The Two Execution Models

The event-driven model

In an event-driven system, the program is a set of handlers: pieces of code that run when something happens (a timer fires, a packet arrives, a conversion finishes) and then return. The Contiki paper states the three consequences: "processes are implemented as event handlers that run to completion"; "because an event handler cannot block, all processes can use the same stack, effectively sharing the scarce memory resources between all processes"; and "locking mechanisms are generally not needed because two event handlers never run concurrently with respect to each other".

Split-phase work. A handler cannot wait for the ADC. Instead it starts the conversion and returns; when the conversion finishes, a completion event calls another handler with the result. The same loop as above becomes:

on timer fired:      start the ADC; return
on ADC done(value):  start sending value; return
on send done:        return                       # nothing left: the node sleeps

This is the split-phase style of [Commands, Events and Split-phase Operation].

Deferred work: tasks. A handler must be short, because nothing else runs while it runs. Work that takes longer is posted as a task, a function to be run later by the scheduler. TinyOS's scheduler runs tasks one at a time in the order they were posted, and a task is never interrupted by another task: it runs to completion. That is the non-preemptive scheduler the practical asks about.

The problems, also in the Contiki authors' words. "The state driven programming model can be hard to manage for programmers", and "not all programs are easily expressed as state machines". Their example is "the lengthy computation required for cryptographic operations", which can take several seconds on a small processor, and in a purely event-driven system "a lengthy computation completely monopolizes" the processor.

Worked example: a long task and a timer, run

A node's timer needs a 1 ms handler every 10 ms (for example to sample a sensor on schedule). At time 2 ms, a 35 ms computation is posted (a compression, or a cryptographic operation). In a run-to-completion scheduler the timer's handler is itself a task, so it must wait for whatever task is running. The program simulates the scheduler twice: with the computation as one task, and with it split into seven tasks of 5 ms, each posting the next when it finishes.

# A run-to-completion task scheduler, the model of TinyOS's tasks.
# Tasks run one at a time, in the order they were posted, and none is ever
# interrupted by another. A timer posts a 1 ms handler task every 10 ms; at
# time 2 ms a 35 ms computation is posted, either whole or in 5 ms pieces.
from collections import deque

def simulate(pieces):
    queue = deque()
    ticks = [10, 20, 30, 40, 50]
    pending = list(ticks)                      # timer ticks not yet posted
    queue.append(("compute", pieces[0], 1))    # posted at time 2
    now, latency = 2, {}
    while queue or pending:
        while pending and pending[0] <= now:   # the timer posts its handler
            t = pending.pop(0)
            queue.append(("tick", 1, t))
        if not queue:                          # nothing to do: sleep until the next tick
            now = pending[0]
            continue
        kind, length, info = queue.popleft()
        if kind == "tick":
            latency[info] = now - info
        now += length                          # run to completion
        if kind == "compute" and info < len(pieces):
            while pending and pending[0] <= now:
                t = pending.pop(0)
                queue.append(("tick", 1, t))
            queue.append(("compute", pieces[info], info + 1))   # repost next piece
    return latency, now

for label, pieces in (("one 35 ms task", [35]), ("seven 5 ms tasks", [5] * 7)):
    latency, end = simulate(pieces)
    shown = ", ".join("%d ms" % latency[t] for t in sorted(latency))
    print("%-17s delays of the timer handler: %s (worst %d ms)"
          % (label + ":", shown, max(latency.values())))
munotes.in138

Event-driven or Multithreaded: The Two Execution Models

one 35 ms task:   delays of the timer handler: 27 ms, 18 ms, 9 ms, 0 ms, 0 ms (worst 27 ms)
seven 5 ms tasks: delays of the timer handler: 2 ms, 3 ms, 4 ms, 0 ms, 0 ms (worst 4 ms)

Read the first line. The single 35 ms task runs from 2 ms to 37 ms. The ticks at 10, 20 and 30 ms all wait behind it and then run in a burst: delays of 27, 18 and 9 ms. A sample meant to be taken every 10 ms is taken three times in 3 ms. This is Contiki's warning, measured: the long computation "completely monopolizes" the processor.

Read the second line. Split into 5 ms pieces, each piece posts the next at the back of the queue, so a tick that arrives during a piece runs as soon as that piece ends. The worst delay is 4 ms, and it can never exceed the length of one piece. In a run-to-completion system, the responsiveness of the whole node is set by its longest task. That is why TinyOS programs break long work into short tasks, and it is the lesson of the practical's scheduler exercise.

The middle ways

Neither model is perfect, and real systems mix them.

  • Contiki: an event-driven kernel, with "optional preemptive multithreading that can be applied to individual processes". Threads are a library, linked only into programs that need them, so the rest of the system keeps its single stack.
  • Protothreads, also from Contiki: code written in sequential style, with waits, but compiled into an event-driven state machine that needs no stack of its own ([Contiki, RIOT and the Other Sensor Operating Systems]).
  • TinyOS's two levels: tasks never preempt each other, but hardware interrupts do preempt tasks; the 2000 paper describes this as "two level scheduling". Urgent, tiny work runs in interrupt handlers; everything else runs as tasks.
  • Preemptive multithreading throughout, as in Mantis, which the Contiki paper describes: "every Mantis program must have stack space allocated from the system heap, and locking mechanisms must be used to achieve mutual exclusion of shared variables".
munotes.in139

Event-driven or Multithreaded: The Two Execution Models

Distinctions

Event-drivenMultithreaded
Unit of executionHandler or task, runs to completionThread, can block
StacksOne, sharedOne per thread, over-provisioned
Locking of shared dataGenerally not neededNeeded
Blocking callsNot allowed: split-phase insteadNatural
Long computationsMust be split, or they block everythingPreempted by the scheduler
Programming styleState machineSequential
ExamplesTinyOS; Contiki's kernelMantis; Contiki's optional threads
PreemptiveNon-preemptive (run to completion)
Can the running job be interrupted by another job?YesNo
Worst delay for an urgent jobVery shortThe longest running job
NeedsA stack per job, locksOne stack

What it does not mean

Event-driven does not mean nothing can interrupt. In TinyOS hardware interrupts preempt tasks; what never happens is one task preempting another.

Run to completion does not mean long-running. It means uninterrupted, which is exactly why each task must be short.

Threads are not wrong for sensor nodes. They cost memory. On a node with more RAM, or for one long computation, a thread can be the better tool, which is why Contiki offers them optionally.

Quick revision

  • Event-driven: handlers and tasks run to completion, one shared stack, no locks, split-phase operations; hard for long computations and state-heavy programs.
  • Multithreaded: a stack per thread, over-provisioned; blocking and preemption; locking needed.
  • Contiki (2004): threads need stacks that "typically has to be over provisioned"; event handlers "cannot block", so "all processes can use the same stack"; a lengthy computation "completely monopolizes" the processor.
  • Worked stacks: 5 × 512 = 2,560 bytes, a quarter of 10 kB of RAM.
  • Simulation: one 35 ms task delays the timer handler by 27, 18 and 9 ms; seven 5 ms tasks, at worst 4 ms. Responsiveness is set by the longest task.
  • Middle ways: Contiki's optional per-process threads, protothreads, TinyOS's two-level scheduling (interrupts preempt tasks); Mantis is preemptive throughout.
munotes.in140

Event-driven or Multithreaded: The Two Execution Models

Test yourself

1. Compare the event-driven and multithreaded execution models for sensor nodes. Event-driven: handlers and tasks run to completion without blocking, share one stack and need no locks, which saves memory; but long computations must be split and programs become state machines. Multithreaded: each thread has its own over-provisioned stack and can block and be preempted, which allows sequential code, at the cost of memory and locking.

2. What does "run to completion" mean, and what follows for how tasks are written? A task, once started, runs until it finishes and is never interrupted by another task. So every task must be short, and long work is split into several tasks, each posting the next, or the whole node becomes unresponsive while it runs.

3. In the simulation, why did splitting the computation reduce the timer handler's worst delay from 27 ms to 4 ms? Because in a run-to-completion scheduler a waiting task can only start when the running one ends. With one 35 ms task the timer handler waited for most of it; with 5 ms pieces, each piece reposted the next at the back of the queue, so the handler ran after at most one piece.

4. Why does a multithreaded system need locks, and why does an event-driven one generally not? Threads can be interrupted in the middle of updating shared data, so another thread could see or change it half-updated; locks prevent that. Event handlers never run concurrently with each other, so a handler always sees shared data in a consistent state.

5. How does Contiki combine the two models? Its kernel is event-driven with one stack, and preemptive multithreading is provided as a library linked only into the processes that need it, so only those processes pay for stacks.

Contents This chapter on its own page

munotes.in141

Chapter Twenty-Six

TinyOS: Components, Tasks and the Scheduler

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Examples of WSN operating systems (case/examples)" (and the paired practical, "Understanding TinyOS Architecture and Execution Model")

In one line

TinyOS builds a node's software from small components wired together at compile time; components talk by commands (downwards) and events (upwards), long work is posted as tasks, and a tiny scheduler runs the tasks one at a time to completion and puts the processor to sleep when there are none.

In the wording a student can write in an examination: TinyOS is an event-driven, component-based operating system for sensor nodes, developed at Berkeley, written in nesC. A system is a scheduler and a graph of components. Each component has command handlers, event handlers, a fixed-size frame (its state) and tasks. Higher components call commands on lower ones, and lower components signal events to higher ones, with the hardware at the bottom. Tasks are deferred procedure calls that run to completion and do not preempt one another; events triggered by hardware interrupts can preempt tasks. The standard scheduler is FIFO and puts the microcontroller to sleep when no task is waiting. Memory is statically allocated, so a component's needs are known at compile time.

Why TinyOS is built this way

Every choice in TinyOS answers one of the requirements in [Why a Sensor Node Needs an Operating System]. Components answer the need for modularity: an application links only the components it uses. Events and tasks answer concurrency on one processor without a stack per thread ([Event-driven or Multithreaded: The Two Execution Models]). Static allocation answers the tiny memory. And a scheduler that sleeps whenever it is idle answers the battery.

Components and their four parts

The 2000 paper that introduced TinyOS describes "a tiny scheduler and a graph of components", and gives each component four parts:

  1. Command handlers: the operations the component offers to the components above it ("start the timer", send this packet).
  2. Event handlers: the code that runs when something the component depends on happens (the timer fired, "the packet was sent").
  3. A frame: the component's own state, of fixed size, allocated when the program is compiled.
  4. Tasks (which the 2000 paper calls "threads"): deferred pieces of work the component can post.

Each component also declares the commands it uses and the events it signals, and these declarations are what let components be wired together. The language that expresses them, nesC, is [nesC: Modules, Configurations, Interfaces and Wiring].

Static allocation. The paper explains why a component's frame is fixed: "the use of static memory allocation allows us to know the memory requirements of a component at compile time", and it avoids the overhead of allocating memory while the program runs. There is no heap to fragment and no allocation to fail in the field.

munotes.in142

TinyOS: Components, Tasks and the Scheduler

The component graph: commands down, events up

A component graph with commands going down, events coming up, and the scheduler's task queue

Figure 26.1 TinyOS: components wired in layers, and the task scheduler

Components are composed in layers: "higher level components issue commands to lower level components and lower level components signal events to the higher level components", with the physical hardware as the lowest level.

Commands go down. The paper defines them as "non-blocking requests made to lower level components". A command typically records its parameters in the frame and posts a task, or calls a lower command, and returns at once with a status saying whether the request was accepted. It "must not wait for long or indeterminate latency actions to take place".

Events come up. The lowest components' event handlers are connected directly to hardware interrupts (a timer, a counter, an external signal). An event handler can store information in its frame, post tasks, signal higher events or call lower commands. The paper's image is exact: a hardware event "triggers a fountain of processing that goes upward through events and can bend downward through commands".

One rule prevents loops: "commands cannot signal events". Both commands and events are meant to do "a small, fixed amount of work".

Tasks

A task is TinyOS's unit of deferred work. TEP 106 defines tasks as "a form of deferred procedure call", which "enable a program to defer a computation or operation until a later time", and gives their essential property: "TinyOS tasks run to completion and do not pre-empt one another", so that "tasks are atomic with respect to other tasks".

In nesC a task is declared with the keyword task and requested with the keyword post:

task void computeAverage() {   // declared in a module's implementation
  // the deferred work: runs later, to completion
}
...
post computeAverage();         // returns SUCCESS or FAIL at once

Three rules about tasks, all from TEP 106 or the 2000 paper, are the ones examiners like:

  1. A task runs to completion. Once started it is never interrupted by another task. It can be interrupted by a hardware event, which the 2000 paper puts as tasks "can be preempted by events".
  2. A task must never block or spin. In the 2000 paper's words, tasks "must never block or spin wait or they will prevent progress in other components".
  3. A task is in the queue at most once. In TinyOS 2 the scheduler must accept a post "unless it is not the first call ... since that task's ... runTask() event has been signaled". So posting a task that is already waiting returns FAIL, and the task still runs once.

Why tasks exist at all. Because commands and event handlers must be short, and some work is not. A timer event handler that needs to average fifty readings posts a task to do it and returns immediately, so the next event is not held up.

munotes.in143

TinyOS: Components, Tasks and the Scheduler

The scheduler

TinyOS 2's standard scheduler is a component, SchedulerBasicP. TEP 106 gives its signature: it provides the Scheduler interface and a parameterised TaskBasic interface (one instance per task, each identified by a number the compiler assigns with nesC's unique() function), and it uses McuSleep, the interface for putting the microcontroller into a low-power state. The basic task is "parameterless and FIFO": tasks run in the order they were posted.

It sleeps when idle. TEP 106 requires that the scheduler's task loop put "the MCU into a low power state when the processor is idle". The 2000 paper explains why this is safe: the processor sleeps "but leaves the peripherals operating, so that any of them can wake up the system", and "once the queue is empty, another thread can be scheduled only as a result of an event", so nothing is lost by sleeping until an interrupt arrives.

It is replaceable. Because the scheduler is a component, an application that needs a different policy (priorities, earliest deadline first) can wire in a different scheduler. TEP 106 adds the rule that follows: components "MUST NOT assume a FIFO policy".

How it all runs: from interrupt to sleep

Put the pieces together and one wake-up of a node looks like this.

  1. The node is asleep; the scheduler's queue is empty.
  2. A hardware timer interrupt fires. The lowest timer component's handler runs (preempting nothing, since nothing is running) and signals an event up to the application.
  3. The application's handler for that event posts a task to take a reading, and returns.
  4. The interrupt is over. The scheduler finds a task in its queue and runs it to completion: the task calls a command to start the ADC, which returns at once.
  5. The queue is empty again, so the scheduler puts the processor to sleep.
  6. The ADC finishes and interrupts. Its component signals a completion event; the application's handler posts a task to send the reading. The scheduler wakes, runs it, and the story continues until the queue empties and the node sleeps again.

The whole design is built so that the processor is awake only while there is work, and every piece of work is short.

Worked example: how big is it?

The 2000 TinyOS fitted, in its paper's words, "in 178 bytes of memory". A small TinyOS 2 application of our own, BlinkTask, in which three timers each post a task that toggles one LED, wired to the standard timer, LED and boot components (the program is printed in full in [Blink: A TinyOS Application Read Line by Line]), was compiled for the TelosB mote by this book's checker, nesc-check.py, inside a TinyOS 2.1.2 laboratory with nescc 1.3.5. The compiler reported 2,682 bytes of program memory (ROM) and 60 bytes of RAM, scheduler, timer and LED drivers included.

munotes.in144

TinyOS: Components, Tasks and the Scheduler

Set that against the MSP430F1611's 48 kB of flash and 10 kB of RAM from [Inside a Sensor Node: The Five Units]: the whole operating system and application use about 5.5 per cent of the flash and well under 1 per cent of the RAM, leaving the rest for the real application's protocols and data. That is what static allocation and linking only the components used buy.

Distinctions

CommandEventTask
DirectionDown, to a lower componentUp, to a higher componentNot a direction: deferred work
Called byA higher component (call)A lower component or hardware (signal)The scheduler, after a post
Must beShort, non-blocking, returns a statusShortRun to completion, never block
Can preemptNoYes, if from an interruptNo other task
TinyOS 1.x schedulerTinyOS 2.x scheduler
WhereC functions in one fileA component, SchedulerBasicP
PolicyFIFO onlyFIFO by default, replaceable
Posting a waiting task againAllowedRefused (the task is in the queue at most once)

What it does not mean

Tasks are not threads, although the 2000 paper used that word. They have no stack of their own and cannot block, wait or be preempted by other tasks.

Run to completion does not mean nothing can interrupt a task. Hardware events can; other tasks cannot.

Components are not run-time objects. They are wired together when the program is compiled; there is no dynamic loading in TinyOS (Contiki's difference, in [Contiki, RIOT and the Other Sensor Operating Systems]).

FIFO is not a promise. It is the default policy, and TEP 106 forbids components from relying on it.

Quick revision

  • TinyOS: event-driven, component-based, written in nesC; a system is "a tiny scheduler and a graph of components".
  • A component's four parts: command handlers, event handlers, a fixed-size frame, tasks.
  • Commands go down ("non-blocking requests"), events come up (from hardware interrupts); commands cannot signal events.
  • Tasks: deferred procedure calls that run to completion and do not pre-empt one another; can be preempted by events; must never block; in the queue at most once.
  • Scheduler: SchedulerBasicP, provides Scheduler and TaskBasic, uses McuSleep; FIFO, replaceable; sleeps when the queue is empty.
  • Static allocation: memory needs known at compile time.
  • Size: the 2000 TinyOS 178 bytes; BlinkTask 2,682 bytes ROM, 60 bytes RAM on TelosB.

Test yourself

1. Describe the architecture of TinyOS. A tiny scheduler and a graph of components, wired at compile time. Each component has command handlers, event handlers, a fixed-size frame and tasks. Higher components call commands on lower ones and lower ones signal events upward, with hardware interrupts at the bottom. Long work is posted as tasks, which the FIFO scheduler runs one at a time to completion, sleeping when none is waiting.

munotes.in145

TinyOS: Components, Tasks and the Scheduler

2. Distinguish commands, events and tasks in TinyOS. Commands are non-blocking requests from a higher component to a lower one, returning a status at once. Events are signalled upward, starting from hardware interrupts, to report that something has happened. Tasks are deferred work posted to the scheduler, run later to completion without preempting one another.

3. What does "tasks run to completion" mean, and why is it important? Once a task starts, no other task interrupts it until it finishes, although hardware events can. It means all tasks can share one stack and need no locks, and it requires tasks to be short and never block.

4. What happens when a component posts a task that is already waiting in the queue? In TinyOS 2 the post returns FAIL, because a task can be in the queue only once; the task still runs once when its turn comes.

5. How does the TinyOS scheduler save energy? When its task queue is empty it puts the microcontroller into a low-power state through the McuSleep interface, leaving the peripherals running so that any interrupt can wake the system; since new tasks can only come from events, nothing is lost by sleeping.

Contents This chapter on its own page

munotes.in146

Chapter Twenty-Seven

Commands, Events and Split-phase Operation

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Examples of WSN operating systems (case/examples)" (and the paired practical, "Implementation of Events, Commands, and Tasks in TinyOS ... to demonstrate split-phase execution")

In one line

In TinyOS a long operation is split into two phases: a command that starts it and returns at once, and an event that reports, later, that it has finished; nothing ever waits.

In the wording a student can write in an examination: split-phase operation divides an operation of long or unpredictable duration into a request phase, a command that starts the operation and returns immediately with a status (SUCCESS means accepted, not finished), and a completion phase, an event signalled when the operation completes, carrying its result. Because commands never block, a component can have many operations in progress, the processor can run other tasks or sleep meanwhile, and all of this needs only one stack. Commands are invoked with call, events with signal, and deferred work is started with post. The nesC manual states it for TinyOS: "all lengthy commands in TinyOS (e.g. send packet) are non-blocking; their completion is signaled through an event (send packet done)".

Why split an operation at all

Reading an analogue sensor takes microseconds to milliseconds; sending a radio packet takes milliseconds and may fail; switching the radio on takes about a millisecond. In a blocking system the program would call read() and wait. In TinyOS nothing may wait, for the reason of [Event-driven or Multithreaded: The Two Execution Models]: a waiting handler would hold up every other handler and task, because they all share the one processor and the one stack. So every such operation is split, and the waiting time becomes time for other work or for sleep.

The shape of a split-phase operation

A split-phase read: the command returns at once, the event comes later

Figure 27.1 Split-phase: request with a command, completion with an event

  1. The application calls a command: call Read.read().
  2. The component starts the hardware and returns at once, with SUCCESS if it accepted the request (or FAIL if, say, a read was already in progress).
  3. The application's handler returns too. The scheduler runs other tasks, or puts the processor to sleep.
  4. The hardware finishes and raises an interrupt.
  5. The component signals the completion event: signal Read.readDone(SUCCESS, value), which runs the application's handler with the result.
  6. The handler keeps its work short, typically by posting a task to use the value.

The interface ties the two halves together. nesC interfaces are, in the manual's word, bidirectional: they specify "a set of functions to be implemented by the interface's provider (commands) and a set to be implemented by the interface's user (events)". A component that uses the Read interface to call read() must implement readDone(), and the compiler refuses the program otherwise. The manual gives the reason: "The interface forces a component that calls the 'send packet' command to provide an implementation for the 'send packet done' event."

munotes.in147

Commands, Events and Split-phase Operation

Three real split-phase interfaces from TinyOS 2:

InterfaceRequest (command)Completion (event)
Read, for a sensorread()readDone(error_t result, value)
AMSend, for a radio packet (TEP 116)send(addr, msg, len)sendDone(msg, error)
SplitControl, to switch a service on or offstart() and stop()startDone(error) and stopDone(error)

The Timer interface (TEP 102) has the same pattern with a twist: startPeriodic(dt) is a command that returns immediately, and fired() is an event that comes back every dt milliseconds, not once.

Worked example: read a sensor and broadcast the reading

The application below does one thing: every second it reads a sensor and broadcasts the value by radio. Three of its operations are split-phase, switching the radio on, reading the sensor and sending the packet, and every one of them is written as a command that returns and an event that completes. It is two files: a module with the logic and a configuration that wires it to TinyOS's components.

module SenseSendC {
  uses interface Boot;
  uses interface SplitControl as RadioControl;
  uses interface Timer<TMilli>;
  uses interface Read<uint16_t>;
  uses interface AMSend;
  uses interface Packet;
}
implementation {
  message_t packet;          // the one buffer, allocated at compile time
  bool busy = FALSE;         // is a send in progress?
  uint16_t reading;

  event void Boot.booted() {
    call RadioControl.start();          // phase 1: ask for the radio
  }
  event void RadioControl.startDone(error_t err) {
    if (err == SUCCESS) call Timer.startPeriodic(1000);
    else call RadioControl.start();     // phase 2: radio on, or try again
  }
  event void RadioControl.stopDone(error_t err) { }

  event void Timer.fired() {
    call Read.read();                   // phase 1: start the conversion
  }
  task void sendReading() {
    uint16_t* payload;
    if (busy) return;
    payload = (uint16_t*)call Packet.getPayload(&packet, sizeof(uint16_t));
    if (payload == NULL) return;
    *payload = reading;
    if (call AMSend.send(AM_BROADCAST_ADDR, &packet, sizeof(uint16_t)) == SUCCESS)
      busy = TRUE;                      // phase 1: send accepted, not finished
  }
  event void Read.readDone(error_t result, uint16_t value) {
    if (result == SUCCESS) {            // phase 2: the value has arrived
      reading = value;
      post sendReading();
    }
  }
  event void AMSend.sendDone(message_t* msg, error_t err) {
    if (msg == &packet) busy = FALSE;   // phase 2: the buffer is free again
  }
}
configuration SenseSendAppC { }
implementation {
  components MainC, SenseSendC, ActiveMessageC;
  components new TimerMilliC();
  components new DemoSensorC() as Sensor;
  components new AMSenderC(6);

  SenseSendC.Boot -> MainC;
  SenseSendC.RadioControl -> ActiveMessageC;
  SenseSendC.Timer -> TimerMilliC;
  SenseSendC.Read -> Sensor;
  SenseSendC.AMSend -> AMSenderC;
  SenseSendC.Packet -> AMSenderC;
}

Reading it as split-phase pairs.

  • Radio: Boot.booted() calls RadioControl.start() and returns; the radio takes time to power up; RadioControl.startDone() arrives later and only then starts the timer. If the start failed, it tries again.
  • Sensor: Timer.fired() calls Read.read() and returns; Read.readDone() arrives with the value and posts sendReading.
  • Radio send: the task calls AMSend.send(); SUCCESS means the radio accepted the packet, so busy is set; AMSend.sendDone() arrives when the packet has actually gone, and clears busy.
munotes.in148

Commands, Events and Split-phase Operation

Why the busy flag. Between send() and sendDone() the packet buffer belongs to the radio stack, which is still reading it. TEP 116 even fixes the rule: AMSend has "an explicit queue of depth one", and a second send "MUST return EBUSY if a prior call to send returned SUCCESS but no sendDone event has been signaled yet". If the task wrote a new reading into it before sendDone(), it would corrupt a packet in flight. The flag is how a split-phase caller remembers that an operation is outstanding. It is the most common bug in a student's first TinyOS program, and the reason the practical asks for split-phase to be demonstrated rather than described.

Why the work is in a task. readDone() only stores the value and posts sendReading; the building and sending of the packet happen in the task. Event handlers stay short.

What the compiler said. Compiled for the TelosB by this book's checker, inside the TinyOS 2.1.2 laboratory, the application built with no warning from its own files and reported 15,448 bytes of program memory and 462 bytes of RAM. Against the toggling application of the previous chapter (2,462 and 38), the difference is the radio stack and the sensor driver the wiring pulled in: TinyOS includes exactly the components a configuration names, and their dependencies, and nothing else.

What may call what: the nesC rules

The nesC reference manual sets out the concurrency rules, and they are the other half of this topic.

Synchronous and asynchronous code. The manual divides a program in two:

  • Synchronous code (SC): "code (functions, commands, events, tasks) that is only reachable from tasks".
  • Asynchronous code (AC): "code that is reachable from at least one interrupt handler".

Tasks never preempt each other, so synchronous code cannot race with other synchronous code. But an interrupt can arrive in the middle of a task, so asynchronous code can race with anything.

The async keyword. A command or event that can run in interrupt context must be declared async. The manual: "nesC reports a compile-time error for any command or event that is AC and that was not declared with async", which "ensures that code that was not written to execute safely in an interrupt handler is not called inadvertently". The events of the application above are all synchronous: TinyOS delivers them from tasks, which is why the program needs no atomic sections.

The two race-free invariants. The compiler warns if either is broken:

  1. If a variable is written in asynchronous code, all accesses to it must be in atomic sections.
  2. If a variable is read in asynchronous code, all writes to it must be in atomic sections.
munotes.in149

Commands, Events and Split-phase Operation

Atomic statements. An atomic block is executed "as-if" no other computation occurred simultaneously. It is how interrupt-level code and task-level code share a variable safely, and the manual warns that "atomic sections should be short".

The usual pattern. Do the minimum in an async event (copy the data, set a flag), then post a task to do the rest in synchronous code. Posting is allowed from async code; the task then runs later, safely, in synchronous context.

Distinctions

Blocking callSplit-phase operation
The callerWaits until the operation finishesContinues at once
Result deliveredAs the return valueIn a later event
Other work meanwhileOnly by other threadsOther tasks, or sleep
StackOne per waiting threadOne for everything
callsignalpost
InvokesA command, downwardsAn event, upwardsA task, to the scheduler
RunsNowNowLater, to completion
Synchronous codeAsynchronous code
Reachable fromTasks onlyAt least one interrupt handler
Can be preempted byInterruptsOther interrupts
Must be declared asyncNoYes, or the compiler refuses it
Shared variablesSafe with other synchronous codeNeed atomic sections

What it does not mean

SUCCESS from a split-phase command does not mean the operation succeeded. It means the request was accepted. Whether it worked is in the completion event.

A split-phase command does not always produce its event. A completion event belongs to a request that was accepted; if the command did not return SUCCESS, the request was refused, and the caller must handle that at once rather than wait for an event.

Atomic does not mean fast. It means uninterrupted, so an atomic block should be short, or it delays every interrupt behind it.

Events are not all interrupts. An event is a callback from a lower component; only some events run in interrupt context, and those are the async ones.

Quick revision

  • Split-phase: a command requests and returns at once; an event later reports completion with the result. "All lengthy commands in TinyOS (e.g. send packet) are non-blocking."
  • Bidirectional interfaces: the user of an interface must implement its events, so a caller of send() must write sendDone().
  • Examples: Read.read() / readDone(), AMSend.send() / sendDone(), SplitControl.start() / startDone(); Timer.startPeriodic() / fired() repeatedly.
  • call a command, signal an event, post a task.
  • SC: reachable only from tasks. AC: reachable from an interrupt handler; must be async; shared variables need atomic sections (the two race-free invariants).
  • Worked application: three split-phase pairs, a busy flag for the buffer, the work in a task; compiled for TelosB: 15,448 bytes ROM, 462 bytes RAM.
munotes.in150

Commands, Events and Split-phase Operation

Test yourself

1. What is split-phase operation? Explain with the example of reading a sensor. Dividing a long operation into a request and a completion. The application calls Read.read(), which starts the conversion and returns SUCCESS at once; the processor meanwhile runs other tasks or sleeps; when the conversion ends, the sensor component signals Read.readDone(result, value) with the reading, and the handler posts a task to use it.

2. Why does TinyOS use split-phase operations instead of blocking calls? Because TinyOS is event-driven with one shared stack and run-to-completion tasks: a blocking call would stop every other handler and task. Splitting lets many operations be outstanding at once, lets the node sleep while hardware works, and needs no stack per operation.

3. In the SenseSend application, why is the busy flag needed? Because between AMSend.send() and AMSend.sendDone() the packet buffer belongs to the radio stack. The flag stops the task from writing a new reading into a buffer still being transmitted; sendDone() clears it.

4. What are synchronous and asynchronous code in nesC, and what must asynchronous code do? Synchronous code is reachable only from tasks; asynchronous code is reachable from at least one interrupt handler. Asynchronous commands and events must be declared async, and variables shared with them must be accessed in atomic sections according to the race-free invariants.

5. State the two race-free invariants. If a variable is written in asynchronous code, all accesses to it must be in atomic sections; if a variable is read in asynchronous code, all writes to it must be in atomic sections.

Contents This chapter on its own page

munotes.in151

Chapter Twenty-Eight

nesC: Modules, Configurations, Interfaces and Wiring

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Examples of WSN operating systems (case/examples)" (and the paired practical, "Implementation of nesC Programming Model ... using modules and configurations to understand component wiring and interface binding")

In one line

nesC is C extended with components: modules hold code, configurations wire components together, interfaces are two-way contracts of commands and events, and the whole program is assembled at compile time from its wiring.

In the wording a student can write in an examination: nesC is the language TinyOS is written in: an extension to C "designed to embody the structuring concepts and execution model of TinyOS", in its manual's words. A nesC program is built from components, of two kinds: modules, which implement their behaviour in C-like code, and configurations, which build a component from other components by wiring them together. Components communicate only through interfaces, which are bidirectional: an interface declares commands, implemented by the component that provides it, and events, implemented by the component that uses it. Wiring uses -> (a used interface to a provided one), <- (the same, reversed) and = (equating a configuration's own interface with an inner one). Generic components are instantiated with new. Components are statically linked, which allows whole-program compilation and compile-time detection of data races.

Why a language of components

[Why a Sensor Node Needs an Operating System] listed "diversity in design and usage": every application needs a different set of drivers and protocols, and the node has no room for any it does not need. nesC's answer is to make the application a wiring diagram. The manual lists the basic idea first: "separation of construction and composition: programs are built out of components, which are assembled ('wired') to form whole programs". Only the components that are wired in are compiled in.

Interfaces: two-way contracts

An interface is a named set of commands and events. The manual explains the two directions: "Interfaces may be provided or used by the component. The provided interfaces are intended to represent the functionality that the component provides to its user, the used interfaces represent the functionality the component needs to perform its job."

And they are two-way: "Interfaces are bidirectional: they specify a set of functions to be implemented by the interface's provider (commands) and a set to be implemented by the interface's user (events)." So wiring a user to a provider connects both directions at once: the user's calls go to the provider's commands, and the provider's signals come back to the user's event handlers. This is what makes split-phase operation ([Commands, Events and Split-phase Operation]) safe: a user of an interface cannot forget to handle its completion event, because the compiler requires it.

Modules and configurations

A module has a specification (the interfaces it provides and uses) and an implementation in C with nesC additions: the command handlers for what it provides, the event handlers for what it uses, its variables (its frame), and its tasks.

munotes.in152

nesC: Modules, Configurations, Interfaces and Wiring

A configuration has a specification too, but its implementation contains no code: it names components and wires them. A configuration is how components are assembled into bigger components, and a whole application is one top-level configuration.

Wiring: the three statements

The manual gives exactly three wiring statements.

  1. endpoint1 -> endpoint2, a link wire: connects a used interface (on the left) to a provided interface (on the right), both on components inside the configuration. "If these two conditions do not hold, a compile-time error occurs."
  2. endpoint1 <- endpoint2: the same connection written the other way round.
  3. endpoint1 = endpoint2, an equate wire: connects one of the configuration's own interfaces to an inner component's interface, making them "effectively ... equivalent". This is how a configuration exports an interface that one of its components really implements.

Two further rules from the manual: the two ends must be compatible (the same interface type, or commands and events with the same signature), and "a configuration's external specification elements must all be wired or a compile-time error occurs". If an interface name is left out, as in SenseSendC.Boot -> MainC, the compiler finds the only interface of the right type on MainC.

Two more features every TinyOS program uses

Generic components. Some components need one private copy per user: every module that wants a timer needs its own timer. These are generic components, instantiated in a configuration with new. The manual: generic components "must be instantiated in a configuration before they can be used", while "components without parameters exist as a single instance which is implicitly instantiated". So new TimerMilliC() creates a timer for this module alone, while MainC and LedsC exist once and are shared.

Fan-out and fan-in. One interface may be wired to several others. Then, in the manual's words, the multiple wiring "will lead to multiple signalers ('fan-in') for the events" and "multiple functions being executed ('fan-out') when commands ... are called". A boot signal wired to three components reaches all three.

Parameterised interfaces. A component can provide many copies of one interface, told apart by a number: the scheduler's TaskBasic interface in [TinyOS: Components, Tasks and the Scheduler] is declared that way, and each task gets its number from the unique() function at compile time.

Worked example: an application that uses every feature

The program below has five files. It defines an interface, Tick, with one command and one event; a module, TickerP, that provides Tick by counting timer fires and signalling when it reaches a limit; a configuration, TickerC, that exports Tick and wires TickerP to a generic timer; a module, CountC, that uses Tick; and the top-level configuration, CountAppC, that wires the application together.

munotes.in153

nesC: Modules, Configurations, Interfaces and Wiring

CountAppC's wiring: arrows run from each used interface to its provider

Figure 28.1 The wiring of the worked application, drawn from its source

The interface. The provider implements start; the user implements reached.

// A bidirectional interface: the provider implements the command,
// the user implements the event.
interface Tick {
  command void start(uint16_t limit);
  event void reached(uint16_t count);
}

The provider module. It provides Tick and uses a timer. Its command starts the timer; its handler for the timer's event counts, and signals Tick's event when the count reaches the goal.

module TickerP {
  provides interface Tick;
  uses interface Timer<TMilli>;
}
implementation {
  uint16_t count = 0;
  uint16_t goal = 0;

  command void Tick.start(uint16_t limit) {
    goal = limit;
    count = 0;
    call Timer.startPeriodic(100);
  }

  event void Timer.fired() {
    count++;
    if (count == goal) {
      call Timer.stop();
      signal Tick.reached(count);
    }
  }
}

The exporting configuration. TickerC offers Tick to the outside world, but TickerP is the one that implements it, so the two are equated with =. TickerP's used Timer is linked with -> to a fresh timer made with new.

configuration TickerC {
  provides interface Tick;
}
implementation {
  components TickerP, new TimerMilliC();

  Tick = TickerP.Tick;              // export: TickerC's Tick is TickerP's
  TickerP.Timer -> TimerMilliC;     // wire: TickerP uses what the timer provides
}

The user module. It uses Tick, so it must implement Tick's event, reached. On boot it asks for four ticks; each time the goal is reached it lights an LED and asks for twice as many.

module CountC {
  uses interface Boot;
  uses interface Tick;
  uses interface Leds;
}
implementation {
  event void Boot.booted() {
    call Tick.start(4);
  }

  event void Tick.reached(uint16_t count) {
    call Leds.led0On();
    call Tick.start(count * 2);     // go again, waiting twice as long
  }
}

The application. The top-level configuration names the components and links each used interface of CountC to a provider.

configuration CountAppC { }
implementation {
  components MainC, LedsC, CountC, TickerC;

  CountC.Boot -> MainC.Boot;
  CountC.Tick -> TickerC.Tick;
  CountC.Leds -> LedsC.Leds;
}

What the compiler checked, and what it said. Built for the TelosB inside the TinyOS 2.1.2 laboratory, the application compiled with no warning from its own files and reported 2,388 bytes of program memory and 40 bytes of RAM. In compiling it, nesC verified every rule of this chapter: that CountC implements reached because it uses Tick; that TickerP implements start because it provides Tick; that every -> runs from a used interface to a provided one of the same type; and that TickerC's own interface is wired. Break any of those and the program does not compile.

Following one call through the wiring. When CountC runs call Tick.start(4), the wiring sends it to TickerC's Tick, which the = has made the same as TickerP's Tick, so TickerP's start command runs. When TickerP runs signal Tick.reached(count), the same wires carry it back, and CountC's reached handler runs. Neither module names the other anywhere: they know only the interface. Replace TickerP with a different module that provides Tick, change one line of TickerC, and CountC does not change at all.

munotes.in154

nesC: Modules, Configurations, Interfaces and Wiring

Distinctions

ModuleConfiguration
ContainsCode: handlers, variables, tasksComponents and wiring only
ImplementsCommands of provided interfaces, events of used onesNothing directly; delegates by wiring
In the exampleTickerP, CountCTickerC, CountAppC
providesuses
The component offersThe interface's commandsNothing; it needs the interface
The component must implementThe commandsThe events
-> (link)= (equate)
ConnectsA used interface to a provided one, both insideThe configuration's own interface to an inner one
PurposeConnect componentsExport an interface
In the exampleTickerP.Timer -> TimerMilliCTick = TickerP.Tick
Generic componentNon-generic component
InstancesOne per newExactly one
Examplenew TimerMilliC()MainC, LedsC

What it does not mean

A configuration is not a module with no code. It has no code because its job is different: assembling components, not implementing behaviour.

Wiring is not a function call. It is a compile-time connection; the calls it enables cost no more than an ordinary function call, because the compiler resolves them.

provides does not mean "calls". A component that provides an interface implements its commands and signals its events; the component that uses it calls the commands and handles the events.

new does not allocate memory at run time. A generic component is instantiated when the program is compiled; nothing is created while it runs.

Quick revision

  • nesC: C with components; programs "built out of components, which are assembled ('wired')".
  • Module: code. Configuration: components and wiring only.
  • Interfaces are bidirectional: commands implemented by the provider, events by the user.
  • Wiring: -> used to provided (link); <- reversed; = equate, to export an interface; ends must be compatible; a configuration's own interfaces must all be wired.
  • Generic components with new (one per instance, e.g. TimerMilliC); others are single instances (MainC, LedsC).
  • Fan-out and fan-in from multiple wiring; parameterised interfaces numbered with unique().
  • Worked CountApp: interface Tick, TickerP provides, TickerC exports with =, CountC uses; compiled for TelosB: 2,388 bytes ROM, 40 bytes RAM.

Test yourself

1. Distinguish a module and a configuration in nesC. A module implements behaviour in C-like code: command and event handlers, variables and tasks. A configuration contains no code; it names components and wires their interfaces together, building a larger component or a whole application.

2. What does it mean that nesC interfaces are bidirectional? An interface contains commands, which the providing component implements, and events, which the using component implements. Wiring a user to a provider connects both directions, so the provider can call back into the user, which is how split-phase completion events reach the caller.

munotes.in155

nesC: Modules, Configurations, Interfaces and Wiring

3. Explain the three wiring statements. endpoint1 -> endpoint2 links a used interface to a provided interface inside a configuration; <- is the same written in reverse; endpoint1 = endpoint2 equates one of the configuration's own interfaces with an inner component's, exporting it.

4. In the worked example, why does TickerC use = for Tick but -> for Timer? Tick is TickerC's own provided interface, implemented by the inner module TickerP, so it is equated to export it. Timer is an interface TickerP uses and an inner timer component provides, so it is linked from user to provider.

5. What is a generic component? Give an example. A component with parameters that must be instantiated in a configuration with new, each instantiation a separate copy, such as new TimerMilliC(), which gives each user its own timer. A non-generic component such as MainC exists once.

Contents This chapter on its own page

munotes.in156

Chapter Thirty

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Examples of WSN operating systems (case/examples)" (and the paired practical, "Simulate a single sensor node in TOSSIM and observe execution logs to understand the TinyOS runtime environment" and "Simulate packet transmission between two motes and analyze signal strength and packet loss under varying radio gain values")

In one line

TOSSIM runs an unchanged TinyOS program on a PC, many motes at once, as a discrete event simulation driven from a Python script; its radio is a gain in dB on each link plus recorded noise, which is enough to show a link going from receiving every packet to receiving none across a band of signal strength a few dB wide: the transitional region.

In the wording a student can write in an examination: TOSSIM is the simulator for TinyOS. The application is compiled with make micaz sim, which replaces the components that touch hardware with simulation implementations, so the same nesC code runs in the simulator and on motes. TOSSIM is a discrete event simulator: it takes events from a queue sorted by time and executes them, an event standing for a hardware interrupt or a system event such as a packet arriving, and tasks also run as events. TOSSIM is a library controlled by a Python (or C++) script, which boots the motes, describes the radio links, chooses which dbg debugging channels to print, and runs the events. In TinyOS 2.1.2 the radio model gives each directed link a gain in dB, gives each node noise generated from a recorded noise trace (the CPM model), and receives each packet with a probability that depends on its signal-to-noise ratio (SNR). Links with a high SNR deliver almost every packet (the connected region), links with a low SNR almost none (the disconnected region), and between them lies the transitional region, where delivery is unreliable and varies from link to link.

Why simulate a sensor network

The paper that introduced TOSSIM in 2003 argues that simulation lets a designer "study system design alternatives in a controlled environment, explore system configurations that are difficult to physically construct, and observe interactions that are difficult to capture in a live system". A sensor network adds its own reason: it is "closely connected to the physical world", which "adds noise, variation, and uncertainty to execution".

It then sets four requirements for a TinyOS simulator:

  1. Scalability: it "must be able to handle large networks of thousands of nodes".
  2. Completeness: it must cover as many interactions as possible, because "the reactive nature of sensor networks requires simulating complete applications".
  3. Fidelity: it "must capture the behavior of the network at a fine grain", and "must reveal unanticipated interactions, not just those a developer suspects".
  4. Bridging: it "must bridge the gap between algorithm and implementation, allowing developers to test and verify the code that will run on real hardware". The paper's warning is practical: if the simulator and the motes are programmed differently, "one has to implement algorithms twice".
munotes.in161

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

TOSSIM's answer is to run the real program: "Compiling unchanged TinyOS applications directly into its framework, TOSSIM can simulate thousands of motes running complete applications."

How TOSSIM runs a TinyOS program

It replaces the components that touch hardware. In the words of TinyOS's own tutorial, "TOSSIM simulates entire TinyOS applications. It works by replacing components with simulation implementations." A simulated millisecond timer, for example, replaces the real one. Everything above those components, the application and most of TinyOS, is the same code that would run on a mote.

It is a discrete event simulator. "When it runs, it pulls events of the event queue (sorted by time) and executes them." An event can stand for a hardware interrupt or for a higher-level happening such as a packet's reception, and "tasks are simulation events, so that posting a task causes it to run a short time (e.g., a few microseconds) in the future". Simulated time does not tick along steadily: it jumps from one event to the next. TinyOS 2.1.2's simulator counts it in ticks of which there are 10,000,000,000 per simulated second, a tenth of a nanosecond each.

It is a library, driven by a script. "TOSSIM is a library: you must write a program that configures a simulation and runs it." The program can be written in Python or in C++. Python "allows you to interact with a running simulation dynamically, like a powerful debugger", while C++ runs faster. The Python interface is built on the C++ one, which calls the simulated program through plain C functions, because "C++ doesn't understand nesC, and nesC doesn't understand C++".

It is built with make micaz sim. The sim option does the work in five steps: it writes an XML description of the application (app.xml, which names every variable and its type), compiles the application for simulation into sim.o, compiles the C++ and Python support, links all three into a shared library, _TOSSIMmodule.so, and copies the Python module TOSSIM.py into the directory. The TinyOS 2.0 tutorial is plain about the platform: "Currently, the only platform TOSSIM supports is the micaz."

A note for the laboratory. On the current Ubuntu that this book's TinyOS 2.1.2 laboratory runs, the simulator build needed an older C compiler, gcc 10, and two extra compiler flags. They go in the application's Makefile; the first line names the application's top-level configuration:

COMPONENT=BlinkTaskAppC
override GCC = gcc-10
export GCC
PFLAGS += -fgnu89-inline -fsigned-char
include $(MAKERULES)

One mote: BlinkTask's execution log

The first simulation exercise asks to "Simulate a single sensor node in TOSSIM and observe execution logs". The program to simulate is BlinkTask from [Blink: A TinyOS Application Read Line by Line]: three timers at 250, 500 and 1,000 binary milliseconds, each of whose events posts a task that toggles one LED and counts the toggles, with a dbg statement in each event and each task. After make micaz sim, this Python script, run.py, simulates one mote for 1.1 simulated seconds:

munotes.in162

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

from TOSSIM import *
import sys

t = Tossim([])
t.addChannel("BlinkTask", sys.stdout)
m = t.getNode(1)
m.bootAtTime(0)
while t.time() < 1.1 * t.ticksPerSecond():
    t.runNextEvent()

Line by line. The import loads the module the build made. Tossim([]) creates the simulator; the empty list means the script will not inspect any of the program's variables. addChannel sends the debugging channel named BlinkTask to the screen. getNode(1) is the mote with ID 1, and bootAtTime(0) boots it at tick 0. The loop runs events, one at a time, until the simulated clock passes 1.1 seconds. Run with python2 run.py, it printed:

DEBUG (1): booted at 0:0:0.000000000
DEBUG (1): Timer0 fired at 0:0:0.244140635
DEBUG (1): task toggle0 ran, toggle 1
DEBUG (1): Timer0 fired at 0:0:0.488281260
DEBUG (1): task toggle0 ran, toggle 2
DEBUG (1): Timer1 fired at 0:0:0.488281280
DEBUG (1): task toggle1 ran, toggle 3
DEBUG (1): Timer0 fired at 0:0:0.732421885
DEBUG (1): task toggle0 ran, toggle 4
DEBUG (1): Timer0 fired at 0:0:0.976562510
DEBUG (1): task toggle0 ran, toggle 5
DEBUG (1): Timer1 fired at 0:0:0.976562530
DEBUG (1): task toggle1 ran, toggle 6
DEBUG (1): Timer2 fired at 0:0:0.976562550
DEBUG (1): task toggle2 ran, toggle 7

What the log shows about the TinyOS runtime.

  1. Every line carries the node's ID, "DEBUG (1)". With a hundred motes in one simulation, that is how the lines are told apart.
  2. Time is in binary milliseconds. Timer0's period is 250 binary milliseconds, 250 / 1,024 = 0.244140625 seconds. The log says 0.244140635, ten billionths of a second later, far below anything a millisecond timer can notice. By 1,000 binary milliseconds, 1,000 / 1,024 = 0.9765625 seconds, the counter reads 7, exactly as the previous chapter worked out (4 + 2 + 1 = 7).
  3. An event posts, and its task runs after it. Every "fired" line is followed by its task's line. At 0.488 seconds Timer0 and Timer1 are both due; the log shows Timer0's event, then its task, then Timer1's event 20 nanoseconds later, then its task. The events are short, the work is deferred, and each task runs to completion before the next event is handled.
  4. Nothing happens between events. The next line would be Timer0 at 1,250 binary milliseconds, about 1.22 seconds, after the script has stopped. Between the lines there is simply no event, which on a real mote is time spent asleep.

dbg: how a simulated mote prints

TOSSIM gives a program four printing calls: dbg prints a line preceded by the node ID; dbg_clear prints without the ID (for printing something long, such as a packet, over several calls); dbgerror and dbgerror_clear do the same for errors, whose lines begin ERROR instead of DEBUG. The first argument names the channel, and the rest are exactly those of printf. BlinkTask prints on the channel BlinkTask.

munotes.in163

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

A channel is just a string, and nothing is printed unless the script connects the channel to an output with addChannel. A statement can name several channels separated by commas, and "A channel can have multiple outputs", for instance the screen and a log file. Channels are how a large simulation stays readable: switch on the channel you are debugging and leave the rest silent. dbg statements print only in the simulator; BlinkTask compiled for the TelosB with every one of them in its source.

How TOSSIM models the radio

TinyOS 1: a graph of bit errors. The 2003 TOSSIM modelled the network as "a directed graph, in which each vertex is a node, and each edge has a bit error probability". Because the edge from u to v is separate from the edge from v to u, links could be asymmetric, and a missing edge could express the hidden terminal problem: edges from a to b and from b to c but none between a and c.

TinyOS 2: gains, noise and the SNR. The TinyOS 2 simulator works on packets and signal strengths instead. Its tutorial states the starting point: "When you start TOSSIM, no node can communicate with any other." The script builds the network by adding links to the Radio object:

  1. r.add(src, dest, gain) adds a directed link: "When src transmits, dest will receive a packet attenuated by the gain value." Adding the link from 1 to 2 does not add the link from 2 to 1, so asymmetric links come free.
  2. The gain is the received power. TinyOS 2.1.2's packet model puts every packet on the air at 0 dBm, so a gain of -80 dB means the receiver hears the packet at -80 dBm. The tutorial's topology file says the same in its own example: "1 2 -54.0 means that when 1 transmits 2 hears it at -54 dBm".
  3. The noise comes from a recording. The model TinyOS 2.1.2 wires in is CpmModelC, whose source describes it: "CPM (closest-pattern matching) is a wireless noise simulation model based on statistical extraction from empirical noise data", "exploiting time-correlated noise characteristics", after Lee and Levis's "Improving Wireless Simulation through Noise Modeling" (IPSN 2007). The script gives each node readings from a noise trace with addNoiseTraceReading and then calls createNoiseModel. TinyOS ships traces in tos/lib/tossim/noise, among them casino-lab.txt and meyer-heavy.txt.
  4. Each packet is a coin toss weighted by its SNR. The SNR is the signal minus the noise, in dB. The model converts it into a probability of reception with a function whose comment says it is "Based on CC2420 measurement", and then draws a random number: below the probability the packet is received, above it the packet is lost.
munotes.in164

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

A packet is judged twice. Reading CpmModelC.nc shows the draw is made when the packet starts to arrive, against the noise at that moment, and again when it finishes, against the noise plus the power of any other packets arriving at the same time. A packet is also lost if the receiver is already receiving one, if the receiver is transmitting, or if a newer packet arrives so much stronger that the older one fails its draw against it. That is how collisions arise in the simulation.

Before sending, the channel is checked. The default MAC is "for a CSMA protocol": a node transmits only when the channel is clear, and CpmModelC calls it clear when what the node hears, its noise plus any packets arriving, is below -72 dBm, a threshold its comment says "comes from the CC2420 data sheet". The medium access ideas behind this are in [Contention: ALOHA and CSMA].

The reception curve, computed

The next program computes the chance of receiving a packet at each SNR by two published rules: the one TOSSIM uses, copied from CpmModelC.nc, and the one Zuniga and Krishnamachari derived for the MICA2 mote's radio (non-coherent FSK, NRZ encoding) in their equation 6, for a 50-byte frame.

# The chance that a packet is received, against its signal-to-noise
# ratio (SNR), by two published rules.
import math

def tossim(snr_db):
    # TinyOS 2.1.2, CpmModelC.nc: "Based on CC2420 measurement"
    pse = 0.5 * math.erfc(0.9794 * (snr_db - 2.3851) / math.sqrt(2))
    return (1 - pse) ** (23 * 2)

def mica2(snr_db, f=50):
    # Zuniga and Krishnamachari, equation 6: MICA2's FSK radio, NRZ
    # encoding, a frame of f bytes
    snr = 10 ** (snr_db / 10)            # decibels to a plain ratio
    return (1 - 0.5 * math.exp(-snr / 2 / 0.64)) ** (8 * f)

print("SNR (dB)   TOSSIM, CC2420   MICA2 model, 50 bytes")
for snr_db in range(0, 13):
    print("%8d   %14.3f   %21.3f" % (snr_db, tossim(snr_db), mica2(snr_db)))
SNR (dB)   TOSSIM, CC2420   MICA2 model, 50 bytes
       0            0.000                   0.000
       1            0.000                   0.000
       2            0.000                   0.000
       3            0.000                   0.000
       4            0.068                   0.000
       5            0.786                   0.000
       6            0.991                   0.000
       7            1.000                   0.018
       8            1.000                   0.235
       9            1.000                   0.668
      10            1.000                   0.922
      11            1.000                   0.989
      12            1.000                   0.999

Reading the curves. Both rise from almost nothing to almost certain across a few dB. TOSSIM's goes from 0.068 at 4 dB to 0.991 at 6 dB; the MICA2 model from 0.018 at 7 dB to 0.989 at 11 dB. Below the rise a link is useless, above it perfect. Its steepness is why the next experiment moves from every packet to none within a few dB of gain. The 0.9794, the 2.3851 and the power 23 × 2 are constants written into TinyOS's code; the chapter uses them as the simulator does and does not claim more for them than the code's own comment.

munotes.in165

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

Two motes, ten gains: packet loss against signal strength

The second exercise asks to "Simulate packet transmission between two motes and analyze signal strength and packet loss under varying radio gain values". Rather than repeat the run ten times, our program RadioTest runs one sender and ten receivers at once, each receiver on its own link with its own gain, so a single simulation sweeps the whole range. The receivers never transmit, so they cannot disturb one another.

RadioTestC.nc, the module:

module RadioTestC {
  uses interface Boot;
  uses interface SplitControl as RadioControl;
  uses interface Timer<TMilli> as SendTimer;
  uses interface Timer<TMilli> as ReportTimer;
  uses interface AMSend;
  uses interface Receive;
}
implementation {
  message_t packet;
  bool busy = FALSE;
  uint16_t sent = 0;          // node 1 counts the packets it sends
  uint16_t received = 0;      // every other node counts what it hears

  event void Boot.booted() {
    call RadioControl.start();
    call ReportTimer.startOneShot(11000);
  }

  event void RadioControl.startDone(error_t err) {
    if (err != SUCCESS) call RadioControl.start();
    else if (TOS_NODE_ID == 1) call SendTimer.startPeriodic(100);
  }
  event void RadioControl.stopDone(error_t err) { }

  event void SendTimer.fired() {
    if (busy || sent == 100) return;
    if (call AMSend.send(AM_BROADCAST_ADDR, &packet, 2) == SUCCESS) {
      busy = TRUE;
      sent++;
    }
  }
  event void AMSend.sendDone(message_t* msg, error_t err) { busy = FALSE; }

  event message_t* Receive.receive(message_t* msg, void* payload, uint8_t len) {
    received++;
    return msg;
  }

  event void ReportTimer.fired() {
    if (TOS_NODE_ID == 1) dbg("Radio", "sent %u packets\n", sent);
    else dbg("Radio", "received %u of them\n", received);
  }
}

RadioTestAppC.nc, the configuration:

configuration RadioTestAppC { }
implementation {
  components MainC, RadioTestC, ActiveMessageC;
  components new TimerMilliC() as SendTimerC;
  components new TimerMilliC() as ReportTimerC;
  components new AMSenderC(7);
  components new AMReceiverC(7);

  RadioTestC.Boot -> MainC;
  RadioTestC.RadioControl -> ActiveMessageC;
  RadioTestC.SendTimer -> SendTimerC;
  RadioTestC.ReportTimer -> ReportTimerC;
  RadioTestC.AMSend -> AMSenderC;
  RadioTestC.Receive -> AMReceiverC;
}

What it does. Every node starts its radio, a split-phase operation that ends in startDone ([Commands, Events and Split-phase Operation]). Node 1, and only node 1, then starts a periodic timer and broadcasts a packet with 2 bytes of payload every 100 binary milliseconds until it has sent 100; the busy flag keeps it from sending a second packet before the first one's sendDone. Every other node counts what it receives. Eleven thousand binary milliseconds after booting, each node reports on the channel Radio. The last packet goes out after about 100 × 100 = 10,000 binary milliseconds, about 9.77 seconds, so the reports come after the sending is over. The program also compiled for the TelosB, to 11,120 bytes of program memory and 370 bytes of RAM, radio stack included.

munotes.in166

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

The script, run.py, takes the name of a noise trace as its argument:

from TOSSIM import *
import os
import sys

trace = sys.argv[1]                      # which recorded noise to use
gains = [-60, -80, -88, -90, -91, -92, -93, -94, -95, -96]

t = Tossim([])
t.addChannel("Radio", sys.stdout)
t.randomSeed(1)                          # the same random draws every run
r = t.radio()
for i, g in enumerate(gains):
    r.add(1, i + 2, g)                   # node i + 2 hears node 1 at g dBm
print("node: " + "".join("%5d" % (i + 2) for i in range(len(gains))))
print("gain: " + "".join("%5d" % g for g in gains))

folder = os.path.join(os.environ["TOSROOT"], "tos/lib/tossim/noise")
noise = [int(v) for v in open(os.path.join(folder, trace)).read().split()[:1000]]
s = sorted(noise)
print("noise: half at or below %d dBm, 9 in 10 at or below %d dBm, loudest %d dBm"
      % (s[499], s[899], s[-1]))
for n in range(1, len(gains) + 2):
    m = t.getNode(n)
    for v in noise:
        m.addNoiseTraceReading(v)        # what this node's radio will hear
    m.createNoiseModel()
    m.bootAtTime(1000 * n)

while t.time() < 12 * t.ticksPerSecond():
    t.runNextEvent()

Line by line. The links go only from node 1 to each receiver, one gain each, from -60 dB down to -96 dB, closely spaced where the loss is expected to change. Every node is given the same first 1,000 readings of the chosen trace, and the script prints a summary of them. randomSeed(1) fixes the random draws, so the same run gives the same counts every time; another seed gives slightly different ones. Node n boots at 1,000 × n ticks, a tenth of a microsecond apart, and the simulation runs for 12 simulated seconds.

Run with python2 run.py casino-lab.txt, the quieter of the two traces:

node:     2    3    4    5    6    7    8    9   10   11
gain:   -60  -80  -88  -90  -91  -92  -93  -94  -95  -96
noise: half at or below -98 dBm, 9 in 10 at or below -97 dBm, loudest -54 dBm
DEBUG (1): sent 100 packets
DEBUG (2): received 100 of them
DEBUG (3): received 100 of them
DEBUG (4): received 100 of them
DEBUG (5): received 100 of them
DEBUG (6): received 100 of them
DEBUG (7): received 94 of them
DEBUG (8): received 47 of them
DEBUG (9): received 1 of them
DEBUG (10): received 0 of them
DEBUG (11): received 0 of them
munotes.in167

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

The same script with the noisier trace:

$ python2 run.py meyer-heavy.txt
node:     2    3    4    5    6    7    8    9   10   11
gain:   -60  -80  -88  -90  -91  -92  -93  -94  -95  -96
noise: half at or below -98 dBm, 9 in 10 at or below -81 dBm, loudest -39 dBm
DEBUG (1): sent 100 packets
DEBUG (2): received 98 of them
DEBUG (3): received 77 of them
DEBUG (4): received 59 of them
DEBUG (5): received 66 of them
DEBUG (6): received 54 of them
DEBUG (7): received 46 of them
DEBUG (8): received 26 of them
DEBUG (9): received 2 of them
DEBUG (10): received 0 of them
DEBUG (11): received 0 of them
Reception ratio against gain for the two noise traces, with the 0.9 and 0.1 lines that bound the transitional region

Figure 30.1 Packets received out of 100, per link, in the two runs above (the -60 dB link, off the chart, received 100 and 98)

Worked example: reading the two runs

Out of 100 packets, the reception ratio of a link is the count divided by 100.

Gain (dB)Quiet traceRatioNoisy traceRatio
-601001.00980.98
-801001.00770.77
-881001.00590.59
-901001.00660.66
-911001.00540.54
-92940.94460.46
-93470.47260.26
-9410.0120.02
-9500.0000.00
-9600.0000.00

The quiet trace. The noise is at or below -97 dBm nine times in ten. A packet heard at -90 dBm then has an SNR of 7 dB or more, the top of TOSSIM's curve: all 100 arrive. At -93 dBm the SNR is about 4 to 5 dB, the steep part: 47 arrive. At -94 dBm, about 3 to 4 dB, the foot of the curve: 1 arrives. The whole fall, from 94 to 1, happens between -92 and -94 dBm, 2 dB.

The noisy trace. The typical reading is the same, -98 dBm, but one reading in ten or more is -81 dBm or louder, and the loudest is -39 dBm. When a burst of noise overlaps a packet, the packet's SNR drops, and if the burst comes within a few dB of the packet's own strength the packet is lost. A burst at -81 dBm ruins a -88 dBm packet but not a -60 dBm one, which only the rare, loudest bursts can reach. So the -60 dB link loses 2 packets, the -80 dB link loses 23, and the fall is spread over more than 13 dB, from 77 at -80 dBm to 26 at -93 dBm, before the link dies at -94 dBm like the quiet one.

The counts wobble. The noisy run received 59 at -88 dBm but 66 at the weaker -90 dBm. Each packet's fate is a random draw, so a count out of 100 varies by several either way; with another seed the order of these two could reverse. A real measurement repeats the run with several seeds and reports the average.

munotes.in168

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

The transitional region

Zuniga and Krishnamachari summarise what measurements of real sensor networks found: "three distinct reception regions in a wireless link: connected, transitional, and disconnected". "The transitional region is often quite significant in size, and is generally characterized by high-variance in reception rates and asymmetric connectivity." To give the regions edges, they "bound the connected region to PRRs greater than 0.9, and the transitional region to values between 0.9 and 0.1", PRR being the packet reception rate, the reception ratio above. Below 0.1 is the disconnected region.

The runs, classified. In the quiet trace, links from -60 to -92 dB are connected (0.94 or better), only the -93 dB link is transitional (0.47), and from -94 dB they are disconnected. In the noisy trace only the -60 dB link is connected (0.98); every link from -80 to -93 dB is transitional (0.77 down to 0.26); and from -94 dB they are disconnected. Same radio, same gains: the noise made the transitional region at least six times wider, from under 2 dB to at least 13 dB. The paper names the noise floor as "Another important element that determines the transitional region".

Why a transitional region exists at all. The paper identifies two causes. The radio: reception does not jump from 0 to 1 at one SNR but climbs over a few dB, as the curves above showed. And the channel: the signal at a given distance is not fixed, because of shadowing and multipath, so two links of the same length can differ by several dB. Its key finding is that for narrow-band radios the transitional region "is not an artifact of the radio non-ideality, as it would exist even with perfect-threshold receivers because of multi-path fading".

Where it begins and ends, computed. The paper models the channel with the log-normal shadowing path loss model, PL(d) = PL(d0) + 10 n log10(d / d0) + X, where X is a random amount in dB with standard deviation σ. Its equation 10 gives the SNR needed for PRRs of 0.9 and 0.1, and equation 13 the distances at which the transitional region begins and ends, taking the signal to lie within two standard deviations of its average. The program below evaluates them with the paper's section III values: a 50-byte frame, 0 dBm transmitted, a noise floor of -115 dBm, 55 dB of loss at the 1 m reference distance, and a path loss exponent of 4.

munotes.in169

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

# Where the transitional region begins and ends: Zuniga and
# Krishnamachari, equations 10, 13 and 14, with their section III values.
import math

f, Pt, Pn, PL0, n = 50, 0, -115, 55, 4   # bytes, dBm, dBm, dB, exponent

def snr_needed(prr):                     # equation 10, in dB
    return 10 * math.log10(-1.28 * math.log(2 * (1 - prr ** (1 / (8 * f)))))

gU, gL = snr_needed(0.9), snr_needed(0.1)
print("PRR 0.9 needs %.2f dB of SNR; PRR 0.1 needs %.2f dB" % (gU, gL))
for sigma in (0, 2, 4):                  # shadowing, in dB
    ds = 10 ** ((Pn + gU - Pt + PL0 + 2 * sigma) / (-10 * n))
    de = 10 ** ((Pn + gL - Pt + PL0 - 2 * sigma) / (-10 * n))
    print("sigma %d dB: transitional from %4.1f m to %4.1f m, Gamma %.2f"
          % (sigma, ds, de, (de - ds) / ds))
PRR 0.9 needs 9.85 dB of SNR; PRR 0.1 needs 7.57 dB
sigma 0 dB: transitional from 17.9 m to 20.4 m, Gamma 0.14
sigma 2 dB: transitional from 14.2 m to 25.7 m, Gamma 0.81
sigma 4 dB: transitional from 11.3 m to 32.4 m, Gamma 1.86

Reading it. The radio alone needs 9.85 dB of SNR for a PRR of 0.9 and 7.57 dB for 0.1, a band only 2.28 dB wide. With no shadowing (σ = 0) that band is a thin ring, from 17.9 m to 20.4 m. With the paper's σ of 4 dB it stretches from 11.3 m to 32.4 m, the paper's own figures. Γ, the transitional region coefficient, is "the ratio of the radius of the transitional and connected regions", (de - ds) / ds, and "The lower the coefficient the better": 0.14 without shadowing, 1.86 with it, which the paper rounds to 1.9. Beyond 11.3 m, in other words, a link of the same length may be excellent or useless depending on where exactly the nodes stand.

Why it matters. The paper warns that "a large number of the links in the network (even higher than 50%) can be unreliable due to the transitional region" in a dense deployment, and that "the idealized perfect-reception-within-range models used in common network simulation tools can be very misleading". It cites the finding that protocols using the minimum hop-count metric "perform poorly in terms of throughput" and that ETX, the "expected number of transmissions", does best. A protocol that picks the longest hop it can hear will pick transitional links. So sensor network routing measures each link's delivery and prefers links that deliver, which is [Routing Tables and What Happens When the Topology Changes].

From positions to gains. TOSSIM itself knows nothing of positions: the script gives it gains. The TinyOS 2.1.2 distribution includes a tutorial by Zuniga, "Building a Network Topology for TOSSIM", whose program applies this same log-normal model, with a path loss exponent, a shadowing deviation, a noise floor and hardware variation, to nodes on a grid or at random, and writes a file of gain lines ready for the script. Its example parameters for a football field are a path loss exponent of 4.7 and a shadowing deviation of 3.2 dB, and for the aisle of a building 3.3 and 5.5 dB.

munotes.in170

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

What TOSSIM cannot tell you

  1. CPU time. "TOSSIM's run-instantly execution model does not capture CPU time." Code takes no simulated time to run, so a task that would take too long on a mote looks fine in TOSSIM.
  2. Preemption. Since interrupts are discrete events, "TOSSIM does not model preemption and the resulting possible TinyOS data races". A race between an interrupt and a task cannot appear in the log.
  3. Energy. The TinyOS 2 tutorial: "TOSSIM currently does not support gathering power measurements." The energy budget of [How Long a Node Lasts: The Energy Budget Worked Out] is not something TOSSIM measures.
  4. Any radio but the one you describe. The gains and the noise are inputs. A simulation with optimistic gains will deliver packets the real field will not.

Distinctions

On a moteIn TOSSIM
The programThe nesC sourceThe same nesC source, hardware components replaced
Built withmake telosb (or micaz)make micaz sim
How many motesOne per deviceMany in one process
TimeReal timeTicks, 10,000,000,000 per simulated second, jumping from event to event
What you seeLEDs, packets, serial outputdbg channels, variables, packets
RadioReal propagation and noiseGains you set, noise from a trace
CPU time, preemption, energyRealNot modelled
TOSSIM for TinyOS 1 (2003)TOSSIM for TinyOS 2.1.2
Radio modelDirected graph, a bit error probability per edgeDirected links, a gain in dB per link
LossEach edge's bit error probabilityNoise generated from a recorded trace (CPM), and the SNR
ReceptionBit by bitPacket by packet, a random draw weighted by SNR
Controlled byExternal programs over a TCP socket, such as the TinyViz GUIA Python or C++ script, with TOSSIM as a library
RegionPRR (Zuniga and Krishnamachari)Behaviour
ConnectedAbove 0.9Almost every packet arrives
Transitional0.9 to 0.1Unreliable, varies from link to link, often asymmetric
DisconnectedBelow 0.1Almost nothing arrives

What it does not mean

A simulated success is not a deployment success. TOSSIM runs the real code, which is its strength, but through the radio the script describes. The noisy run shows how different the same gains become in a different environment.

munotes.in171

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

A gain is not a distance. TOSSIM has no positions; a distance becomes a gain only through a propagation model such as the one in [Radio Technology in WSNs: The Sensor Radio and Its Link Budget], and shadowing makes that conversion uncertain.

A radio's range is not a circle. The disc model, every packet received inside a range and none outside, is the model the transitional region paper calls "very misleading". The quiet run looks nearly disc-like only because its noise is steady and its gains were chosen by hand.

The transitional region is not caused by bad radios. Even a receiver with a perfect threshold would see one, because shadowing and multipath spread the signal strength of equally long links.

Older tutorials set noise differently. The TinyOS 2.0 tutorial sets a node's noise with setNoise(node, mean, variance), a Gaussian. The model TinyOS 2.1.2 wires in, CpmModelC, takes its noise from a trace instead, which is why the script above feeds readings to every node.

Quick revision

  • TOSSIM: the TinyOS simulator; runs unchanged TinyOS applications; replaces the components that touch hardware with simulation implementations; a discrete event simulator (events from a queue sorted by time; tasks are events too).
  • Built with make micaz sim (micaz the only platform), in five steps: app.xml, sim.o, the C++ and Python support, _TOSSIMmodule.so, TOSSIM.py.
  • A library driven by a Python (or C++) script: Tossim, getNode, bootAtTime, addChannel, runNextEvent, radio().add(src, dest, gain), addNoiseTraceReading, createNoiseModel, randomSeed.
  • dbg(channel, format, ...) prints only in the simulator, only for channels the script adds; lines start DEBUG (node ID).
  • The 2003 paper's four requirements: scalability, completeness, fidelity, bridging.
  • TinyOS 2.1.2 radio: directed links with a gain in dB, packets sent at 0 dBm; noise from a trace via CPM (closest-pattern matching); reception a random draw weighted by SNR; clear channel below -72 dBm.
  • Transitional region: between the connected (PRR above 0.9) and disconnected (below 0.1) regions; high variance, asymmetric links; caused by the radio's gradual curve and by shadowing and multipath, widened by noise.
  • Our runs: quiet noise, the fall from 94 to 1 packets within 2 dB; heavy noise, transitional from -80 to -93 dB.
  • Zuniga and Krishnamachari, 50-byte frame: 9.9 dB and 7.6 dB for PRR 0.9 and 0.1; with σ = 4 dB, transitional from 11.3 m to 32.4 m, Γ about 1.9.
  • TOSSIM does not model CPU time, preemption or energy.

Test yourself

1. What is TOSSIM, and how does it simulate a TinyOS application? TOSSIM is the TinyOS simulator. It compiles the unchanged application with make micaz sim, replacing only the components that touch hardware with simulation implementations, and runs it as a discrete event simulation: events are taken from a queue sorted by time and executed, and tasks are run as events too. It is a library controlled by a Python or C++ script, which boots the motes, sets up the radio links and noise, chooses the debugging channels and runs the events.

munotes.in172

TOSSIM: Simulating Motes, Radio Gain and Packet Loss

2. Write a Python script to simulate one mote in TOSSIM and print its debugging messages. Import TOSSIM, create the simulator with Tossim([]), connect the application's dbg channel to standard output with addChannel, get the node with getNode(1), boot it with bootAtTime(0), and call runNextEvent in a loop until time() passes the chosen number of ticksPerSecond(). The BlinkTask script in this chapter is a complete example.

3. How does TOSSIM in TinyOS 2 model the radio? Each directed link has a gain in dB, set by the script with add(src, dest, gain); packets are sent at 0 dBm, so the gain is the power received. Each node has noise generated from a recorded trace by the CPM model. Each packet's SNR, signal minus noise, gives a probability of reception, and a random draw decides the packet; other packets arriving at the same time count as noise, which produces collisions.

4. Define the transitional region. What causes it? It is the range of links between the connected region, where more than 90 per cent of packets arrive, and the disconnected region, where fewer than 10 per cent do. Its links are unreliable, vary widely and are often asymmetric. It is caused by the radio, whose reception rises gradually over a few dB of SNR, and by the channel, where shadowing and multipath make equally long links differ in signal strength; varying noise widens it.

5. In the two RadioTest runs, the -80 dB link received 100 packets with one noise trace and 77 with the other. Explain. In the quiet trace the noise stays near -98 dBm, so a -80 dBm packet has an SNR of about 18 dB and is always received. In the noisy trace one noise reading in ten or more is -81 dBm or louder, and a packet that overlaps such a burst has almost no SNR and is lost. The link is the same; the environment moved it from the connected region into the transitional region.

6. What can TOSSIM not tell you? How long code takes to run (it does not model CPU time), races between interrupts and tasks (it does not model preemption), how much energy a node uses (it does not measure power), and how the real radio environment will behave (the gains and noise are only what the script supplies).

Contents This chapter on its own page

munotes.in173

Chapter Thirty-One

Contiki, RIOT and the Other Sensor Operating Systems

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Examples of WSN operating systems (case/examples)"

In one line

TinyOS is one answer among several: Contiki keeps an event-driven kernel but loads programs at run time and offers threads and protothreads, RIOT gives real prioritised threads on a tickless kernel, and MANTIS, Nano-RK and LiteOS choose threads, real-time scheduling and a Unix-like environment, each paying for convenience in memory.

In the wording a student can write in an examination: Contiki is a lightweight, open source operating system for sensor nodes, written in C, built around an event-driven kernel with optional preemptive multithreading as a library, dynamic loading of individual programs and services at run time, lightweight protothreads for sequential-style code, and the uIP TCP/IP stack; it is simulated with Cooja. RIOT is an open source operating system for low-end IoT devices with a minimal kernel, multithreading with fixed priorities and preemption, a tickless scheduler, and support for the standard IoT protocols (6LoWPAN, IPv6, UDP, CoAP). MANTIS is a multithreaded, layered OS with preemptive priority scheduling. Nano-RK is a real-time OS with resource reservations and rate monotonic scheduling. LiteOS is a Unix-like OS with a shell, a hierarchical file system and C++ support. Other approaches include virtual machines such as Maté and SOS, another event-driven system.

How sensor operating systems differ

The questions every sensor OS must answer were set out in [Why a Sensor Node Needs an Operating System], and the choice between events and threads in [Event-driven or Multithreaded: The Two Execution Models]. The systems in this chapter differ first in their architecture. The survey by Farooq and Kunz names four classic ones:

  • Monolithic: all services bundled into a single system image. Module interaction is cheap and the image small, but the system is hard to understand, modify and maintain.
  • Microkernel: "minimum functionality is provided inside the kernel", the rest in separate servers. More reliable and easier to extend, at the cost of crossings between user and kernel.
  • Virtual machine: programs run on a virtual machine that resembles hardware. "The key advantage is its portability and a main disadvantage is typically a poor system performance."
  • Layered: services in layers, each using the one below. Manageable and easy to understand, but not very flexible.

Its conclusion for sensor nodes: an OS "should have an architecture that results in a small kernel size", that allows the kernel to be extended, and that is flexible, so that "only application-required services get loaded onto the system". The survey classes TinyOS as monolithic (the whole system is one image built at compile time), Contiki and LiteOS as modular, and MANTIS as layered.

Contiki

Contiki came from the Swedish Institute of Computer Science in 2004. Its paper introduces it as "a lightweight operating system with support for dynamic loading and replacement of individual programs and services", which "is built around an event-driven kernel but provides optional preemptive multithreading that can be applied to individual processes". It is written in C and was ported to the MSP430 and the Atmel AVR.

munotes.in174

Contiki, RIOT and the Other Sensor Operating Systems

A running system. "A running Contiki system consists of the kernel, libraries, the program loader, and a set of processes." A process is an application program or a service, and "All processes, both application programs and services, can be dynamically replaced at run-time." A process is defined by an event handler function and an optional poll handler, all processes share one address space, and "Interprocess communication is done by posting events."

The core and the loaded programs. A Contiki system is split, at compile time, into two parts. The core (typically the kernel, the program loader, the most used parts of the C run-time and libraries, and the communication stack with its radio drivers) is one binary image stored in the node before deployment and normally not changed afterwards. Loaded programs are brought in later by the program loader, either over the network or from attached storage such as EEPROM. This split is what lets a deployed network be reprogrammed one program at a time.

Contiki's ROM and RAM, with the core at the bottom and a loaded program above it

Figure 31.1 Contiki: the core, fixed before deployment, and a program loaded at run time

The kernel: events, polling and one stack

The Contiki kernel "consists of a lightweight event scheduler that dispatches events to running processes and periodically calls processes' polling handlers". It never preempts an event handler, so event handlers must run to completion. It has two kinds of event:

  1. Asynchronous events are "a form of deferred procedure call": queued by the kernel and delivered to the target process later.
  2. Synchronous events immediately schedule the target process, and control returns to the poster only when the target has finished, like a procedure call between processes.

Polling is the third mechanism: polls "can be seen as high priority events that are scheduled in-between each asynchronous event", used by processes close to the hardware to check a device's status. The kernel uses a single shared stack for all processes, rewound between event handlers.

Interrupts never post events. Contiki "never disables interrupts", so that it can run on top of a real-time executive. Posting an event from an interrupt handler would then race with the event handlers, so Contiki does not allow it: an interrupt sets a polling flag, and the poll handlers do the work. The survey sums it up: since events run to completion and interrupts cannot post events, Contiki "provides serialized access to all resources".

Power saving is left to the application. "The Contiki kernel contains no explicit power save abstractions." Instead, the event scheduler exposes the size of its event queue, so the application can put the processor to sleep when no event is waiting.

munotes.in175

Contiki, RIOT and the Other Sensor Operating Systems

Loadable programs, services and libraries

Loadable programs carry relocation information. The loader allocates memory for the program (or aborts if it cannot), relocates it, and calls its initialisation function, which may start or replace processes.

Services are processes that implement something other processes use: communication stacks, sensor drivers, data-handling algorithms. The paper describes a service as "a form of a shared library" that can be replaced at run time. A service layer next to the kernel keeps track of the running services, each named by a text string. A service's interface is a version number and a table of function pointers; an application calls it through a small stub that finds the service once, caches its process ID, and checks the version. When a service is replaced, the kernel tells the old one to remove itself, and the old one can hand its state to its replacement, tagged with its version number.

Libraries can be linked in three ways: statically with the core, statically with a loaded program, or called as a service. Frequently used functions (the paper's example is memcpy()) belong in the core; rarely used ones (atoi()) travel inside the program that needs them.

Communication is a service too, so a stack or a routing protocol can be replaced at run time, and two stacks can be loaded at once for comparison. Because synchronous event handlers run to completion, the whole stack can share one packet buffer, with no copying. According to the 2011 survey, Contiki provides uIP, "a TCP/IP protocol stack for small 8 bit micro-controllers", a lighter layered stack called Rime, and ContikiRPL, an implementation of the RPL routing protocol of [Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP]. Contiki networks are simulated with Cooja.

Threads, when a program really needs them

Preemptive multithreading in Contiki "is implemented as a library on top of the event-based kernel", linked only into programs that ask for it. Each thread gets its own stack, and preemption uses a timer interrupt that saves the registers and switches back to the kernel's stack. The platform-specific part is small: for the MSP430 it "consists of 25 lines of C code". A thread uses six calls: mt_yield(), mt_post(), mt_wait() and mt_exit() from inside a running thread, and mt_start() and mt_exec() to set one up and run it; mt_exec() is called from an event handler. A thread suits one long computation, such as cryptography, that would otherwise hold up every event.

munotes.in176

Contiki, RIOT and the Other Sensor Operating Systems

Worked example: how big is Contiki?

Table 1 of the Contiki paper gives the compiled size of each part of the system, with an example service (a sensor data replicator), for the two processors it ran on. In bytes:

PartAVRMSP430
Kernel1,044810
Service layer128110
Program loader(not ported)658
Multi-threading library678582
Timer library9060
Replicator stub18298
Replicator service1,7521,558
Total3,8743,876

The AVR column adds up as 1,044 + 128 + 678 + 90 + 182 + 1,752 = 3,874, and the MSP430 column as 810 + 110 + 658 + 582 + 60 + 98 + 1,558 = 3,876. On the MSP430 the kernel and service layer together are 810 + 110 = 920 bytes.

The paper places this between its neighbours: Contiki's code size is "larger than that of TinyOS" but "smaller than that of the Mantis system". Its kernel is larger than TinyOS's because it offers more: TinyOS's kernel has only a FIFO task queue, while Contiki's has events and prioritised poll handlers, and its flexibility at run time needs code that TinyOS resolves when it compiles.

What run-time loading bought in practice. The authors built a 40-node alarm system. Its application was about 6 kilobytes, and the complete system image with the core and the C library nearly 30 kilobytes. Reprogramming one node by wire took just over 30 seconds, and the whole network "at least 30 minutes of work". Loading one changed component over the radio took about two minutes, "a reduction in an order of magnitude", with the nodes left where they were.

Protothreads

Event-driven code has one great weakness: a handler cannot wait. Anything that needs several steps with waits in between (switch the radio on, wait, send, wait for an acknowledgement, switch it off) must be written as a state machine, with the current step stored in a variable and a switch on it in every handler. The protothreads paper says this style "makes many programs difficult to write, maintain, and debug".

The idea. A protothread lets the programmer write the steps in order, with a blocking wait, PT_WAIT_UNTIL(condition), between them, while the program remains event-driven underneath. "A protothread is stackless": it has no stack of its own, "all protothreads in a system run on the same stack, which is rewound every time a protothread blocks". It is driven by repeated calls to the function it lives in, and each call carries on from where the last one waited.

What it costs. The memory overhead is "only two bytes per protothread". In the programs the authors rewrote, "the majority of the state machines could be entirely removed", and "the number of lines of code was reduced by one third", for an execution time overhead "on the order of a few processor cycles".

munotes.in177

Contiki, RIOT and the Other Sensor Operating Systems

How it works. In the paper's portable version, the two bytes hold a line number. PT_BEGIN opens a C switch statement on that number, with case 0 for a fresh start. Each wait stores its own line number (the C macro __LINE__) and places a case label for that number right there. So a later call jumps straight back to the wait and tests the condition again: if it is still false, the function returns; if true, it carries on. A case label inside a loop inside a switch is legal C, a trick the paper traces to Duff's device.

The program. Here is that mechanism in full, written for this book: two protothreads, a radio that is on for 2 ticks and off for 6, and a sensor that takes a reading every 5 ticks and must wait for the radio to be on before sending it. The loop in main is the event loop, calling each protothread once per tick.

#include <stdio.h>

/* All a protothread keeps between calls: the line to carry on from. */
struct pt { unsigned short resume; };

#define PT_BEGIN(p)  switch ((p)->resume) { case 0:
#define PT_WAIT_UNTIL(p, cond)                                            \
    (p)->resume = __LINE__; __attribute__((fallthrough)); /* meant */   \
    case __LINE__: if (!(cond)) return 0
#define PT_END(p)    } (p)->resume = 0; return 1

static int now;             /* the clock, in ticks, advanced by main */
static int radio_on;        /* shared by the two protothreads */

/* The radio's duty cycle: on for 2 ticks, off for 6, for ever. */
static int radio(struct pt *p)
{
    static int until;       /* static, so it survives a wait */
    PT_BEGIN(p);
    for (;;) {
        radio_on = 1;
        printf("%2d  radio on\n", now);
        until = now + 2;
        PT_WAIT_UNTIL(p, now >= until);
        radio_on = 0;
        printf("%2d  radio off\n", now);
        until = now + 6;
        PT_WAIT_UNTIL(p, now >= until);
    }
    PT_END(p);
}

/* A reading every 5 ticks, sent as soon as the radio is on. */
static int sensor(struct pt *p)
{
    static int next = 1, reading;
    PT_BEGIN(p);
    for (;;) {
        PT_WAIT_UNTIL(p, now >= next);
        reading = 20 + now % 7;
        printf("%2d  sensor reads %d and waits for the radio\n", now, reading);
        PT_WAIT_UNTIL(p, radio_on);
        printf("%2d  sensor sends %d\n", now, reading);
        next += 5;
    }
    PT_END(p);
}

int main(void)
{
    struct pt r = {0}, s = {0};
    printf("a protothread's state: %zu bytes\n", sizeof r);
    for (now = 0; now < 18; now++) {    /* one call of each per tick */
        radio(&r);
        sensor(&s);
    }
    return 0;
}
a protothread's state: 2 bytes
 0  radio on
 1  sensor reads 21 and waits for the radio
 1  sensor sends 21
 2  radio off
 6  sensor reads 26 and waits for the radio
 8  radio on
 8  sensor sends 26
10  radio off
11  sensor reads 24 and waits for the radio
16  radio on
16  sensor sends 24
16  sensor reads 22 and waits for the radio
16  sensor sends 22
munotes.in178

Contiki, RIOT and the Other Sensor Operating Systems

Reading the run. The sensor's code reads like a thread: read, wait for the radio, send, repeat. At tick 1 the radio happens to be on, so the reading goes at once. At tick 6 it is off, so the sensor protothread returns at its wait on every call until tick 8, when the radio comes on and the same call carries straight on to the send. At tick 16 two readings go together: the one taken at 11, which waited 5 ticks, and the one due at 16. Neither function has a stack of its own; between calls each protothread is 2 bytes, the line to carry on from. The reading itself is made up (20 plus the tick modulo 7), to keep the example short.

The limits, stated by the paper. "Automatic variables are not saved across a blocking wait", because the stack is rewound; a variable that must survive a wait has to be static, as until, next and reading are above. The switch-based version also "limits the use of the C switch statement together with protothread statements", since the protothread's own switch would be confused by another. And a protothread can block only in its own function, never inside a function it calls.

Protothreads against threads, in numbers. The MANTIS figures later in this chapter give a default thread stack of 128 bytes on a MICA2 mote with 4 KB of RAM. Five threads would reserve 5 × 128 = 640 bytes of stack, which is 640 / 4,096 = 0.15625, about 16 per cent of the RAM; five protothreads need 5 × 2 = 10 bytes, 10 / 4,096 = 0.00244140625, about 0.24 per cent.

RIOT

RIOT is a free and open source operating system for low-end IoT devices, written from scratch for them. Its 2018 overview paper says it runs in memory of the order of 10 kilobytes, "on devices with neither MMU (memory management unit) nor MPU (memory protection unit)", on 8-bit, 16-bit and 32-bit microcontrollers. It calls such devices "very constrained", citing RFC 7228, whose device classes appear in [Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP].

Structure. RIOT is a modular system "built around a minimalistic kernel". Modules are chosen when the system is compiled: core (the kernel), hardware abstraction (cpu, boards, drivers and periph), sys (system libraries such as networking and file systems), pkg (third-party libraries) and the application. The minimal configuration "requires 3.2 kBytes of ROM and 2.8 kBytes of RAM on 32-bit Cortex-M platforms", and one compliant with 6LoWPAN "requires 38.5 kBytes of ROM and 10 kBytes of RAM".

munotes.in179

Contiki, RIOT and the Other Sensor Operating Systems

Real threads. "A thread in RIOT is akin to a thread in Linux", each with a priority. The price of a thread is its control block, its stack and its saved registers; on a Cortex-M the thread control block is 36 bytes (12 without messaging), and a thread with simple logic can run "starting from 128 bytes of RAM in total". The kernel provides mutexes, semaphores and messaging, each compiled only if used. Multithreading is optional: a single-threaded application can drop most of the scheduler.

Scheduling. The scheduler is "based on fixed priorities and preemption with O(1) operations, allowing for soft real-time capabilities". The highest-priority ready thread runs, interrupted only by interrupt service routines, and threads of equal priority are not time-sliced. The scheduler is tickless: "It does not depend on CPU time slices and periodic system timer ticks", so the system wakes only when something happens. When no thread is ready, the lowest-priority idle thread runs and puts the processor into the most energy-saving mode possible.

Standards and licence. RIOT follows open network standards, supporting the low-power IP stack (6LoWPAN, IPv6, UDP and CoAP) in its default network stack, GNRC, and it aims to comply with ANSI C (C99). Its code is under the LGPLv2.1, "a non-viral copyleft license", while applications built on it may use other licences.

Why RIOT is not Contiki. Contiki keeps an event kernel and adds threads as a library; RIOT makes preemptive, prioritised threads the kernel's basic abstraction, and relies on keeping each thread small instead of avoiding threads.

MANTIS, Nano-RK and LiteOS

These three appear in the survey by Farooq and Kunz, whose figures are given here as it reports them.

MANTIS (the MultimodAl system for NeTworks of In-situ wireless Sensors) is a multithreaded OS written in C, with a layered architecture: hardware; the kernel and scheduler, the MAC and physical layer (the COMM layer) and the device drivers; the system API; and on top, the network stack, a command server and the user threads. The survey reports a footprint of 500 bytes for the kernel, scheduler and network stack. Scheduling is preemptive and priority-based, with five priority classes (kernel, sleep, high, normal and idle), round robin within a class, and a default time slice of 10 ms. Each thread's stack comes from the heap, 128 bytes by default, and each context switch costs about 60 microseconds, which against a 10 ms slice is 0.06 / 10 = 0.006, under 1 per cent. Threads synchronise with mutexes and semaphores, a Unix-like shell runs on the node, and applications can be tested on a PC before being moved to the mote. The Contiki paper names the cost: "every Mantis program must have stack space allocated from the system heap, and locking mechanisms must be used to achieve mutual exclusion of shared variables".

munotes.in180

Contiki, RIOT and the Other Sensor Operating Systems

Nano-RK is a real-time OS: "a fixed, preemptive multitasking real-time OS for WSNs", in the survey's words. It supports reservations of CPU time, network bandwidth and sensors, so that a task cannot use more than its share, and schedules periodic tasks by rate monotonic scheduling (the shorter a task's period, the higher its priority), with the priority ceiling protocol against priority inversion. Memory is managed statically only. Networking uses a socket-like interface and RT-Link, a TDMA link protocol. The survey reports that it "uses 2 Kb of RAM and 18 Kb of ROM".

LiteOS, from the University of Illinois at Urbana-Champaign, is Unix-like: a thread-based programming model (with callbacks for events), a hierarchical file system (LiteFS), a shell (LiteShell) that runs on a base station or PC and sends commands to the nodes over the radio, and object-oriented programming in LiteC++. The survey's table credits it with dynamic memory management and memory protection for processes.

Other approaches: virtual machines and scripts

The Contiki paper's related work names systems that load code a different way:

  • Maté is "a virtual machine for TinyOS devices"; code for it can be downloaded at run time. Virtual-machine code is smaller, so it costs less energy to send, but running it costs more, and "for long running programs the energy saved during the transport of the binary code is instead spent in the overhead of executing the code".
  • MagnetOS "uses a virtual Java machine to distribute applications across the sensor network".
  • SensorWare "provides an abstract scripting language for programming sensors", for platforms less constrained than Contiki's, and EmStar is likewise meant for larger systems.
  • SOS is, with TinyOS and Contiki, one of the systems the protothreads paper lists as "based on an event-driven model".

Worked comparison: six systems, one table

TinyOSContikiRIOTMANTISNano-RKLiteOS
ArchitectureMonolithic (one image, components)ModularModular, minimal kernelLayeredMonolithicModular
Programming modelEvents and tasks (threads added later)Events; protothreads; optional threadsThreads with prioritiesThreadsThreadsThreads and event callbacks
SchedulingFIFO tasks, run to completionEvents as they come; prioritised pollingFixed priority, preemptive, ticklessPriority classes, round robinRate monotonic, reservationsPriority-based round robin
MemoryStaticDynamic, loadable programsMostly static structuresHeap for stacksStatic onlyDynamic, protected
NetworkingActive messagesuIP (TCP/IP), Rime6LoWPAN, IPv6, UDP, CoAPCOMM layer, user-level stackSockets, RT-LinkFile-based
Real timeNoNoSoft real timeLimitedYesNo
LanguagenesCCCCCLiteC++
SimulatorTOSSIMCoojaRIOT native; also CoojaAVRORANone listedAVRORA
munotes.in181

Contiki, RIOT and the Other Sensor Operating Systems

Sources, row by row: the survey's Tables 1 and 2 for TinyOS, Contiki, MANTIS, Nano-RK and LiteOS, and the RIOT paper for RIOT. RIOT native, in its paper's words, compiles and runs RIOT applications "as user processes in a host OS", so that "up to hundreds of RIOT instances can run in parallel"; the paper also names Cooja among the emulators RIOT can be tested with. For Nano-RK the survey lists no simulator.

How to use the table in an answer. Pick the rows the question is about. A question asking to compare TinyOS and Contiki is answered from the architecture, programming model, memory and networking rows; one asking which OS suits a real-time application is answered from the real-time row, which points to Nano-RK, and to RIOT for soft real time.

Distinctions

TinyOSContiki
KernelEvent-driven, FIFO tasksEvent-driven, events plus polling
ThreadsNone in the basic modelOptional library, per process; protothreads
Changing code in the fieldWhole image replacedIndividual programs and services loaded at run time
LanguagenesCC
NetworkingActive messagesuIP TCP/IP, Rime
SizeSmaller, in the Contiki paper's own comparisonLarger than TinyOS, smaller than Mantis; 3,876 bytes on MSP430 with an example service
SimulatorTOSSIMCooja
A threadA protothreadAn event handler
StackIts ownShared, rewoundShared
Can waitAnywhereOnly in its own function, with PT_WAIT_UNTILNo
State keptStack and registers2 bytesWhatever the programmer stores
Local variables across a waitKeptLost (use static)Not applicable

What it does not mean

Contiki is not multithreaded at its core. Its kernel is event-driven; threads are a library that only the programs needing them link in.

Protothreads are not threads. They cannot be preempted, have no stack, and lose their local variables at a wait. They are a way of writing an event-driven state machine in sequential style.

Dynamic loading does not mean the whole system can be replaced cheaply. The core stays fixed; the saving comes from sending one program instead of the whole image.

More features are not free. Every row of the comparison table that adds convenience (threads, dynamic memory, a file system, a shell) adds bytes of RAM or ROM, which is why TinyOS, the smallest, stayed the reference for the most constrained motes.

Quick revision

  • Architectures (survey): monolithic, microkernel, virtual machine, layered; TinyOS monolithic, Contiki and LiteOS modular, MANTIS layered.
  • Contiki (2004): event-driven kernel, optional preemptive multithreading as a library, dynamic loading of programs and services; core fixed, programs loaded later; asynchronous and synchronous events, polling; one shared stack; interrupts never post events; services with version number and function table; uIP, Rime, ContikiRPL; Cooja.
  • Contiki sizes: 3,874 bytes (AVR), 3,876 bytes (MSP430) with an example service; reprogramming a 40-node network by wire at least 30 minutes, one component over the air about 2 minutes.
  • Protothreads: sequential code with PT_WAIT_UNTIL, stackless, 2 bytes each, lines of code cut by one third; built on a switch and __LINE__; automatic variables not kept across a wait.
  • RIOT: minimal kernel, modules, threads with fixed priorities, preemption, O(1), tickless, idle thread; minimal 3.2 kB ROM, 2.8 kB RAM (Cortex-M), with 6LoWPAN 38.5 kB ROM, 10 kB RAM; 6LoWPAN, IPv6, UDP, CoAP; LGPLv2.1.
  • MANTIS: multithreaded, layered, five priority classes, round robin, 10 ms slice, 128-byte stacks. Nano-RK: real-time, reservations, rate monotonic. LiteOS: Unix-like, LiteFS, LiteShell, LiteC++.
  • Virtual machines: Maté, MagnetOS; scripts: SensorWare.
munotes.in182

Contiki, RIOT and the Other Sensor Operating Systems

Test yourself

1. Describe the architecture of the Contiki operating system. A Contiki system consists of the kernel, libraries, a program loader and processes, which are applications or services. The kernel is a lightweight event scheduler that dispatches asynchronous and synchronous events to processes and calls their poll handlers, all on one shared stack. The system is split into a core, fixed before deployment, and programs loaded at run time over the network or from EEPROM. Services, such as communication stacks, are called through a version-checked interface and can be replaced at run time. Preemptive multithreading is an optional library, and communication is provided by uIP and Rime.

2. Compare TinyOS and Contiki. Both have event-driven kernels whose handlers run to completion on one stack. TinyOS is written in nesC and linked statically into one image, with a FIFO task scheduler and active-message networking; Contiki is written in C, loads and replaces individual programs and services at run time, adds polling and an optional thread library, supports protothreads, and provides TCP/IP through uIP. TinyOS is smaller; Contiki is more flexible. TinyOS is simulated with TOSSIM, Contiki with Cooja.

3. What are protothreads? Give their advantages and limitations. Protothreads let event-driven code be written in a sequential, thread-like style with a blocking wait, PT_WAIT_UNTIL, while using no stack of their own. Each costs two bytes and adds only a few processor cycles; in the authors' programs most state machines disappeared and code shrank by a third. Their limits: local (automatic) variables are not kept across a wait, a switch statement cannot be mixed with them in the switch-based version, and they can block only in their own function.

munotes.in183

Contiki, RIOT and the Other Sensor Operating Systems

4. What features make RIOT suitable for low-end IoT devices? A small modular system built around a minimal kernel, needing 3.2 kB of ROM and 2.8 kB of RAM in its minimal configuration; real multithreading with small thread control blocks; a fixed-priority, preemptive scheduler with O(1) operations for soft real time; a tickless design with an idle thread that puts the processor into its deepest sleep; and support for the standard IoT protocols 6LoWPAN, IPv6, UDP and CoAP.

5. Name three other sensor operating systems and one distinguishing feature of each. MANTIS: a multithreaded, layered OS with preemptive priority scheduling in five classes. Nano-RK: a real-time OS with reservations of CPU, network and sensors and rate monotonic scheduling. LiteOS: a Unix-like OS with a hierarchical file system, a shell on the base station and LiteC++.

6. Five threads with 128-byte stacks, or five protothreads: how much RAM does each need on a 4 KB mote? The threads need 5 × 128 = 640 bytes of stack, about 16 per cent of 4,096 bytes; the protothreads need 5 × 2 = 10 bytes, about 0.24 per cent.

Contents This chapter on its own page

munotes.in184

Chapter Thirty-Two

Ad Hoc Networks: MANETs, and How a Sensor Network Differs

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Ad-hoc networks in WSNs" (and the paired practical, "Simulate a mobile ad-hoc network and visualize packet movement to study dynamic topology formation")

In one line

An ad hoc network is built by its own nodes, with no base station or access point, each node relaying packets for the others; a MANET is an ad hoc network whose nodes move, so its topology keeps changing; a sensor network is also ad hoc, but differs from a MANET in scale, traffic, energy, mobility and purpose.

In the wording a student can write in an examination: an ad hoc network is a network set up for a particular need without any fixed infrastructure: the nodes configure themselves and relay packets for one another, so that communication can span more than one radio hop (multihop). A MANET (mobile ad hoc network) is "an autonomous system of mobile nodes" (RFC 2501) that are free to move, so the network has a dynamic topology: links appear and break as nodes move, and routes must be found and repaired continually. Examples are disaster relief, military and rescue operations, and networks on large construction sites. A wireless sensor network is an ad hoc network too, but it differs from a MANET: it has far more nodes, deployed densely; its traffic flows from many sources to one sink; its data are redundant and addressed by content; its energy is far scarcer and batteries are rarely replaced; its nodes are usually still, so its topology changes mainly through failures and energy depletion; and its nodes may not have global identifiers. So routing protocols designed for MANETs do not suit WSNs as they are.

What "ad hoc" means

Karl and Willig define it from the word itself: "An ad hoc network is a network that is setup, literally, for a specific purpose, to meet a quickly appearing communication need." Their simplest example is a few laptops cabled together in a meeting room, and the point of the example is self-configuration: "the network is expected to work without manual management or configuration".

Usually the term means more than that. A wireless ad hoc network has no infrastructure: no base station, no access point, no cables to a switch. Nodes within radio range talk directly; nodes further apart talk through other nodes, which relay the packets. In Karl and Willig's words, the nodes "together form a network that relays packets between nodes to extend the reach of a single node". That relaying is multihop communication, the same idea as the multihop sensor network of [Single Hop or Multiple Hops: The Energy Argument Worked Out].

An infrastructure network where every node talks to a base station, beside an ad hoc network where nodes relay for each other

Figure 32.1 Infrastructure and ad hoc: through a base station, or through each other

Infrastructure against ad hoc. In an infrastructure network, a mobile phone or a laptop talks only to a base station or access point, which is wired to the rest of the network and does all the forwarding. Planning, power and routing live in the infrastructure. In an ad hoc network every node is also a router, and the network exists only as long as its nodes cooperate.

munotes.in185

Ad Hoc Networks: MANETs, and How a Sensor Network Differs

MANETs

RFC 2501, the IETF's document on mobile ad hoc networking, gives the technology three other names: "Mobile Packet Radio Networking", from military research in the 1970s and 80s, "Mobile Mesh Networking", and "Mobile, Multihop, Wireless Networking", which it calls "perhaps the most accurate term, although a bit cumbersome".

The definition. "A MANET consists of mobile platforms", called nodes, "which are free to move about arbitrarily": in airplanes, ships, trucks, cars, or carried by people. "A MANET is an autonomous system of mobile nodes." It may work in isolation, or through gateways connect to a fixed network, where it acts as a stub network: it carries traffic that starts or ends at its own nodes, but not traffic passing through. Its nodes have wireless transmitters and receivers, and at any moment their positions, antennas, power levels and interference decide which nodes can hear which: the connectivity is "a random, multihop graph", and it "may change with time as the nodes move or adjust their transmission and reception parameters".

Applications. RFC 2501 lists industrial and commercial "cooperative mobile data exchange", mesh networks as robust and inexpensive alternatives or additions to cellular infrastructure, military networks, "wearable" computing, and "fire/safety/rescue operations or other scenarios requiring rapidly-deployable communications". Karl and Willig give firefighters in a disaster relief operation, and large construction sites where installing access points, "let alone cables, is not a feasible option".

The two basic challenges. Karl and Willig name them: "the reorganization of the network as nodes move about and handling the problems of the limited reach of wireless communication". The first is the dynamic topology of the practical: routes that worked a minute ago break when a node on them walks out of range. The second is why nodes must relay at all. The characteristics that follow from these, and the challenges they cause, are the next chapter, [The Characteristics and Challenges of Ad Hoc Networks in a WSN].

Worked example: when does a link break?

Two nodes are 100 m apart, each with a radio range of 150 m, and they walk directly away from each other at 5 m/s each. Their separation grows by 5 + 5 = 10 m every second. The link lasts while the separation is at most 150 m, so it breaks after (150 - 100) / 10 = 5 seconds. Had they walked side by side in the same direction, the separation would not change and the link would last indefinitely. How long a link lives depends on how the nodes move relative to each other, which is why the choice of mobility model matters to every MANET simulation.

munotes.in186

Ad Hoc Networks: MANETs, and How a Sensor Network Differs

A MANET's topology, simulated

The practical asks to "Simulate a mobile ad-hoc network and visualize packet movement to study dynamic topology formation." A network simulator draws the nodes and animates the packets; the program below shows the same thing in numbers.

The mobility model. Simulations of MANETs need a rule for how nodes move, and the most widely used is the random waypoint model, which Camp, Boleng and Davies describe: a node "begins by staying in one location for a certain period of time (i.e. a pause time)", then "chooses a random destination in the simulation area and a speed that is uniformly distributed between [minspeed, maxspeed]", travels there, and on arrival "pauses for a specified time period before starting the process again". The same survey warns that the nodes' starting positions, scattered at random, are "not representative of the manner in which nodes distribute themselves when moving", and suggests discarding "the initial 1000 s of simulation time".

The program. Twelve nodes move in a 500 m square with a radio range of 150 m, at speeds between 1 and 10 m/s, for 3,000 simulated seconds, of which the first 1,000 are discarded. Every second it recomputes which pairs are in range (the links), counts the links that appeared or broke, and searches for a route from node 0 to node 11 with the fewest hops. It runs twice, with no pause and with a 60-second pause at each waypoint, and prints the topology every minute for the first five minutes of the first run.

# A mobile ad hoc network whose topology changes as its nodes move. Each
# node follows the random waypoint model (Camp, Boleng and Davies 2002):
# pause, pick a random destination and a speed, walk there, pause again.
import math
import random
from collections import deque

N, SIDE, RANGE = 12, 500.0, 150.0      # nodes, field side (m), radio range (m)

def links(pos):
    return {(a, b) for a in range(N) for b in range(a + 1, N)
            if math.dist(pos[a], pos[b]) <= RANGE}

def route(ls, src, dst):               # fewest hops, by breadth-first search
    back, todo = {src: None}, deque([src])
    while todo:
        n = todo.popleft()
        for a, b in ls:
            for x, y in ((a, b), (b, a)):
                if x == n and y not in back:
                    back[y] = n
                    todo.append(y)
    if dst not in back:
        return None
    path = [dst]
    while back[path[-1]] is not None:
        path.append(back[path[-1]])
    return path[::-1]

def run(pause, seconds=3000, warmup=1000, show=False, seed=7):
    rnd = random.Random(seed)
    pos = [(rnd.uniform(0, SIDE), rnd.uniform(0, SIDE)) for _ in range(N)]
    goal, speed, wait = list(pos), [0.0] * N, [0.0] * N
    before, changes, reachable = links(pos), 0, 0
    for t in range(1, seconds + 1):
        for i in range(N):
            if wait[i] > 0:
                wait[i] -= 1
                continue
            d = math.dist(pos[i], goal[i])
            if d <= speed[i]:          # arrives this second: pause, then a new goal
                pos[i], wait[i] = goal[i], pause
                goal[i] = (rnd.uniform(0, SIDE), rnd.uniform(0, SIDE))
                speed[i] = rnd.uniform(1, 10)
            else:
                f = speed[i] / d
                pos[i] = (pos[i][0] + f * (goal[i][0] - pos[i][0]),
                          pos[i][1] + f * (goal[i][1] - pos[i][1]))
        now = links(pos)
        if t > warmup:
            changes += len(now ^ before)
            reachable += route(now, 0, N - 1) is not None
            if show and t % 60 == 0 and t <= warmup + 300:
                path = route(now, 0, N - 1)
                print("  t = %d s: %2d links, route 0 to %d: %s" % (t, len(now), N - 1,
                      " > ".join(map(str, path)) if path else "none"))
        before = now
    minutes = (seconds - warmup) / 60
    print("pause %2d s: %.2f link changes per node per minute, route 0 to %d exists %.0f%% of the time"
          % (pause, changes / N / minutes, N - 1, 100 * reachable / (seconds - warmup)))

run(0, show=True)
run(60)
munotes.in187

Ad Hoc Networks: MANETs, and How a Sensor Network Differs

  t = 1020 s: 19 links, route 0 to 11: none
  t = 1080 s: 20 links, route 0 to 11: none
  t = 1140 s: 24 links, route 0 to 11: 0 > 5 > 7 > 11
  t = 1200 s: 24 links, route 0 to 11: 0 > 2 > 8 > 11
  t = 1260 s: 19 links, route 0 to 11: 0 > 2 > 11
pause  0 s: 3.88 link changes per node per minute, route 0 to 11 exists 74% of the time
pause 60 s: 2.21 link changes per node per minute, route 0 to 11 exists 61% of the time

Reading the topology lines. Of the 66 possible pairs among 12 nodes, 19 to 24 are linked at the five moments printed. At the first two there is no route at all from node 0 to node 11: the network is partitioned, split into pieces that cannot reach each other. At 1,140 s a three-hop route 0 > 5 > 7 > 11 has formed; a minute later the route is 0 > 2 > 8 > 11, through different relays; a minute after that it is 0 > 2 > 11, two hops. Nobody configured any of these routes. They exist because the nodes happen to be where they are, and a routing protocol must discover each one and notice when it breaks. That is dynamic topology formation.

munotes.in188

Ad Hoc Networks: MANETs, and How a Sensor Network Differs

Reading the two summary lines. With no pause, links change about 3.88 times per node per minute; with a 60-second pause, 2.21 times. Long pauses make the network calmer, as Camp and colleagues found: "long pause times (i.e. over 20 s) produce a stable network (i.e. few link changes per MN) even at high speeds", MN being a mobile node. A calmer network is not automatically a better connected one: in this run the route from 0 to 11 existed 74 per cent of the time without pauses and 61 per cent with them. How the nodes move decides both numbers, which is why a MANET protocol's results mean little without its mobility model. Measuring protocols on such networks is [Measuring a MANET Protocol: Throughput, Delivery Ratio and Delay].

How a sensor network differs from a MANET

The two are relatives. Sohraby and colleagues note that "both involve multihop communications", and Karl and Willig that both face the same general problems of reorganisation and limited radio reach. But three standard texts each list differences, and together they give the table an examination asks for.

MANETWireless sensor network
PurposeCommunication between people's devices: voice, access to a web serverSensing the environment and reporting it
EquipmentFairly powerful: a laptop or PDA with a comparably large batterySimple, cheap nodes, limited in power, computation and memory
Number of nodesFewerCan be "several orders of magnitude higher", densely and often redundantly deployed
Traffic patternPoint to point, between pairs of nodesMany sources to one sink; broadcast and multicast
Traffic over timeConventional, well-understood applicationsVery low rates for long periods, then bursts when an event happens
DataNot generally redundantRedundant and correlated; protocols can be data-centric
EnergyScarce, but batteries can be recharged or replacedThe critical constraint; batteries rarely replaced, long unattended life
MobilityNodes move; the main cause of topology changeNodes usually still; the phenomenon or the sink may move
Why the topology changesMovementFailures, energy depletion, sleeping nodes, varying links
Node identifiersCan be assumedMay not exist; global IDs cost too much
ReliabilityEach node should be fairly reliableAn individual node is "next to irrelevant"
Quality of serviceSet by traditional applications, such as low jitter for voiceNew notions, taking energy into account
Self-configurationRequiredRequired: "probably most similar" here

Where each row comes from: purpose, equipment, traffic over time, energy, reliability, QoS, data-centricity, identifiers in MANETs and self-configuration from Karl and Willig; traffic pattern, redundancy, energy and the scale of "several orders of magnitude" from Sohraby and colleagues and from Akyildiz and colleagues, who both use the phrase; density, the broadcast paradigm, limited resources and missing global IDs from Akyildiz and colleagues; mobility from Sohraby ("In most scenarios (applications) the sensors themselves are not mobile (although the sensed phenomena may be)") and Karl and Willig, who add that the observed phenomenon and the sinks may move. The row on why the topology changes draws these together, as the next paragraph explains.

munotes.in189

Ad Hoc Networks: MANETs, and How a Sensor Network Differs

A point that looks like a contradiction. Akyildiz and colleagues list "The topology of a sensor network changes very frequently" among the differences, yet Sohraby notes the sensors are usually not mobile. Both hold: a sensor network's topology changes because nodes are "prone to failures", run out of energy, and are switched off to save it, and because links near the edge of the transitional region come and go ([TOSSIM: Simulating Motes, Radio Gain and Packet Loss]). The causes differ from a MANET's; the effect, a topology that will not stay put, is shared.

The consequence for protocols. Sohraby and colleagues draw it plainly: "For these reasons the plethora of routing protocols that have been proposed for MANETs are not suitable for WSNs, and alternative approaches are required." Akyildiz and colleagues say the same of ad hoc protocols in general: "they are not well suited to the unique features and application requirements of sensor networks". The routing protocols built for sensor networks are in [Routing Challenges and Design Issues in WSNs] and the chapters after it. Karl and Willig's summary is the fairest one-line answer: "there are commonalities", but the differences justify "considering WSNs as a system concept distinct from MANETs".

Distinctions

Infrastructure networkAd hoc network
ForwardingBy the base station or access pointBy the nodes themselves
SetupPlanned and installed beforehandSelf-configuring, when needed
RangeOne hop to the infrastructureMultihop, extended by relaying
Failure of one elementThe base station's failure cuts off its whole cellThe loss of one node may only lengthen routes
Ad hoc networkMANETWSN
InfrastructureNoneNoneNone within the field; a sink or gateway at the edge
MobilityNot requiredDefiningUsually none
Main purposeCommunication without infrastructureCommunication between moving devicesSensing and reporting

What it does not mean

Ad hoc does not mean disorganised. It means set up for the need, without prior infrastructure. An ad hoc network organises itself, and its routing can be quite elaborate.

A WSN is an ad hoc network, but usually not a MANET. Its nodes form their own multihop network, but most of them never move.

No infrastructure does not mean no gateway. RFC 2501 allows a MANET to connect to a fixed network through gateways, and a sensor network almost always reports through a sink or gateway ([Gateway Concepts]).

A partitioned network is not a failed one. In the simulation, node 0 had no route to node 11 for minutes at a time and then gained one. Protocols for such networks must expect partitions and survive them.

munotes.in190

Ad Hoc Networks: MANETs, and How a Sensor Network Differs

Quick revision

  • Ad hoc network: set up "for a specific purpose", no infrastructure, self-configuring, nodes relay for one another (multihop).
  • MANET: "an autonomous system of mobile nodes" (RFC 2501), free to move, connectivity "a random, multihop graph"; may be isolated or a stub network behind gateways; other names: mobile packet radio, mobile mesh, mobile multihop wireless networking.
  • MANET applications: disaster relief, rescue, military, construction sites, wearable computing.
  • Two basic MANET challenges (Karl and Willig): reorganisation as nodes move, and the limited reach of wireless communication.
  • Dynamic topology: links form and break as nodes move; routes must be found and repaired; networks may be partitioned.
  • Random waypoint: pause, pick a random destination and a speed in [min, max], travel, pause again; discard the start of a run (1,000 s).
  • Our run: 3.88 link changes per node per minute without pauses, 2.21 with 60 s pauses; route 0 to 11 existed 74 and 61 per cent of the time.
  • WSN versus MANET: many more nodes, dense; many-to-one traffic; redundant, data-centric data; energy critical; nodes usually still; topology changes through failures and energy; no global IDs; single node unimportant; self-configuration the most similar point. MANET routing protocols are not suitable for WSNs as they are.

Test yourself

1. What is an ad hoc network? How does it differ from an infrastructure network? An ad hoc network is set up for a particular need without any fixed infrastructure: the nodes configure themselves and relay packets for one another, so communication can cross several radio hops. In an infrastructure network every node communicates through a base station or access point, which is planned and installed beforehand and does all the forwarding.

2. Define a MANET and give its applications. A MANET is an autonomous system of mobile nodes, free to move arbitrarily, whose wireless connectivity forms a random multihop graph that changes as the nodes move or adjust their transmission. It can work in isolation or connect to a fixed network through gateways. Applications include disaster relief and rescue operations, military networks, large construction sites, cooperative data exchange in industry, and wearable computing.

3. Differentiate between a MANET and a wireless sensor network. (Any six, from the table.) A WSN has many more nodes, densely deployed; its traffic flows from many sources to a sink, where a MANET's is point to point; its data are redundant and it uses data-centric protocols; energy is its critical constraint and batteries are rarely replaced; its nodes are usually static, so its topology changes through failures and energy depletion rather than movement; its nodes may lack global identifiers; and an individual node matters little, while in a MANET each device serves a user.

munotes.in191

Ad Hoc Networks: MANETs, and How a Sensor Network Differs

4. What is meant by dynamic topology in a MANET? Explain with an example. The set of links changes over time because nodes move: two nodes in range become separated and their link breaks, while others come into range and new links form. For example, two nodes 100 m apart with a range of 150 m, moving directly apart at 5 m/s each, lose their link after (150 - 100) / 10 = 5 seconds. Routes that use such links must be discovered again, and the network may split into partitions for a while.

5. Why are MANET routing protocols not directly suitable for sensor networks? They are designed for point-to-point traffic between identified nodes that move, with energy treated as recoverable. A sensor network has many-to-one, data-centric traffic, far more nodes that may lack global identifiers, redundant data that can be aggregated, and an energy supply that must last for months without replacement, so it needs protocols designed around those facts.

Contents This chapter on its own page

munotes.in192

Chapter Thirty-Three

The Characteristics and Challenges of Ad Hoc Networks in a WSN

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Characteristics and challenges of ad-hoc networks in WSNs"

In one line

An ad hoc network's links come and go, carry less than the radio's rate, run on batteries and are easy to attack; in a sensor network that means sharing one radio channel, finding routes with no fixed paths through relays that sleep, organising itself after being scattered, managing itself without a centre, and doing all of it in a few kilobytes of memory.

In the wording a student can write in an examination: RFC 2501 gives four characteristics of mobile ad hoc networks: (1) dynamic topologies, links that change randomly and rapidly and may be unidirectional; (2) bandwidth-constrained, variable-capacity links, whose real throughput is well below the radio's rate, so congestion is normal; (3) energy-constrained operation, where conserving energy may be the most important design criterion; and (4) limited physical security, with eavesdropping, spoofing and denial of service, offset by the robustness of decentralised control; to which it adds scalability. In a sensor network these become challenges: the shared wireless medium (collisions, hidden terminals), multihop routing with no fixed paths and with relays that sleep to save energy, self-organisation of nodes deployed without planning, identifiers and addressing when global IDs are too costly, decentralised management with only local knowledge and little memory, and quality of service under all these limits. Time synchronisation and localisation are two more, taken in the next chapter.

The characteristics, from RFC 2501

RFC 2501 describes MANETs, but each characteristic applies to an ad hoc sensor network, some with more force.

1. Dynamic topologies

"Nodes are free to move arbitrarily; thus, the network topology", which is typically multihop, "may change randomly and rapidly at unpredictable times, and may consist of both bidirectional and unidirectional links."

In a sensor network the nodes rarely move, but the topology is dynamic all the same: nodes fail, run out of energy, and switch their radios off to save it, and links near the edge of the transitional region come and go ([TOSSIM: Simulating Motes, Radio Gain and Packet Loss]). Unidirectional links are common: node A hears B but B does not hear A, because their transmitters, receivers or surroundings differ. A protocol that assumes a node it can hear can also hear it will send acknowledgements into silence.

2. Bandwidth-constrained, variable-capacity links

"Wireless links will continue to have significantly lower capacity than their hardwired counterparts", and the throughput actually achieved, "after accounting for the effects of multiple access, fading, noise, and interference conditions, etc.", "is often much less than a radio's maximum transmission rate". The consequence: "congestion is typically the norm rather than the exception".

In a sensor network the radio's rate is low to begin with: 250 kbit/s for an IEEE 802.15.4 radio at 2.4 GHz ([Radio Technology in WSNs: The Sensor Radio and Its Link Budget]). And relaying divides it further, because every hop uses the same channel.

munotes.in193

The Characteristics and Challenges of Ad Hoc Networks in a WSN

Worked example: a relay chain on one channel. A reading travels from node A through relays B and C to the sink D, three hops, all on one channel. Suppose all four nodes are close enough to hear one another, so only one of them can transmit at a time. Each packet then occupies the channel three times, once per hop, so the chain delivers at most one packet for every three packet-times: 250 / 3, about 83.3 kbit/s at the very best, before any collision, back-off or acknowledgement. In a longer chain, hops far enough apart can transmit together, but a hop and its neighbours still cannot, so the end-to-end rate stays a fraction of the radio's rate. And near the sink every packet in the network converges on the last few hops, which is where congestion appears first ([Transport Protocol Design Issues in WSNs] takes this further).

3. Energy-constrained operation

"Some or all of the nodes in a MANET may rely on batteries or other exhaustible means for their energy. For these nodes, the most important system design criteria for optimization may be energy conservation."

In a sensor network this is not "some or all" but all, and the batteries are rarely replaced. Energy shapes everything else in this chapter: a relay that sleeps cannot relay, and every message spent on organising the network is energy not spent on sensing. The costs are counted in [Energy Efficiency in Ad Hoc Networks: Where the Energy Goes].

4. Limited physical security

"Mobile wireless networks are generally more prone to physical security threats than are fixed-cable nets. The increased possibility of eavesdropping, spoofing, and denial-of-service attacks should be carefully considered." There is one benefit: "the decentralized nature of network control in MANETs provides additional robustness against the single points of failure of more centralized approaches".

In a sensor network nodes also lie unattended in the open, where they can be picked up and tampered with, and every node that relays is a node that could lie about routes. These are [Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks] and [Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood].

And scalability

RFC 2501 adds that some MANETs "may be relatively large (e.g. tens or hundreds of nodes per routing area)", and that while the need for scalability is not unique to them, "the mechanisms required to achieve scalability likely are". A sensor network may have thousands of nodes, so this weighs more heavily still.

munotes.in194

The Characteristics and Challenges of Ad Hoc Networks in a WSN

The challenges these create in a sensor network

The shared wireless medium

All nodes in range of one another share one channel. Two transmissions that overlap at a receiver collide and are both lost. Worse, a node cannot always tell the channel is busy: in the hidden terminal problem, A and C both reach B but not each other, so each senses a free channel and both transmit to B at once. The 2003 TOSSIM paper uses exactly this case to show what its network graph can express: "for nodes a,b,c, there are edges (a, b) and (b, c) but no edge (a, c)". Deciding who may transmit when is the job of the MAC protocol ([Hidden and Exposed Terminals, and RTS and CTS], and the MAC chapters around it).

Routing with no fixed paths, through relays that sleep

Because distant transmission is expensive, Dargie and Poellabauer note, it is better to split a long distance into short hops, "leading to the challenge of supporting multi-hop communications and routing. Multi-hop communication requires that nodes in a network cooperate with each other to identify efficient routes and to serve as relays." With no infrastructure there is no fixed path to follow, and in a dynamic topology a path found now may be gone soon.

Saving energy makes it harder. Many nodes switch their radios off when idle, and "during these down-times, the sensor node cannot receive messages from its neighbors nor can it serve as a relay for other sensors". Dargie and Poellabauer mention two remedies: wake-up on demand, with a second, very low-power radio that listens for wake-up calls, and adaptive duty cycling, in which not all nodes sleep at once and a subset stays awake "to form a network backbone". The routing families are [Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid]; the sleeping schedules are the MAC chapters.

Self-organisation after an ad hoc deployment

Sensor networks are often deployed without planned positions, sometimes dropped from aircraft over a disaster area, where "many sensor nodes may not survive such a drop". The nodes that do survive "must autonomously perform a variety of setup and configuration steps, including the establishment of communications with neighboring sensor nodes, determining their positions, and the initiation of their sensing responsibilities". And they must keep doing it unattended: "configuration, adaptation, maintenance, and repair must be performed in an autonomous fashion".

Dargie and Poellabauer name four forms of this self-management:

  1. Self-organisation: adapting configuration to the state of the system and its surroundings, for example choosing a transmission power that keeps enough neighbours in range.
  2. Self-optimisation: monitoring and optimising the use of the node's own resources.
  3. Self-protection: recognising and resisting intrusions and attacks.
  4. Self-healing: discovering, identifying and reacting to network disruptions.
munotes.in195

The Characteristics and Challenges of Ad Hoc Networks in a WSN

And all of it "must be designed and implemented such that they do not incur excessive energy overheads".

Identifiers and addresses

An Internet host is configured with an address, or asks a server for one. An ad hoc sensor network has no such server, and may not even have unique names: Akyildiz and colleagues note that sensor nodes "may not have global identification (ID) because of the large amount of overhead and large number of sensors", and Karl and Willig that giving every node a unique identifier is costly, so that protocols working without one "might become important in WSNs". Many sensor protocols therefore address data and places instead of nodes ([Design Principles: Data Centricity, Location, Activity and Heterogeneity]), and use identifiers that need only be unique among neighbours. Where IP is used, the 6LoWPAN adaptation of [Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP] builds addresses from the link-layer identifiers the radios already have.

Decentralised management, in small memories

"The large scale and the energy constraints of many wireless sensor networks make it infeasible to rely on centralized algorithms (e.g., executed at the base station)" for topology management or routing. Nodes must instead "collaborate with their neighbors to make localized decisions, that is, without global knowledge". The price is that the results "will not be optimal"; the gain is that they "may be more energy-efficient than centralized solutions". RFC 2501 lists the same property first among what a MANET routing protocol needs: "Distributed operation: This is an essential property, but it should be stated nonetheless."

Memory narrows the choice further: "routing tables that contain entries for each potential destination in a network may be too large to fit into a sensor's memory. Instead, only a small amount of data (such as a list of neighbors) can be stored."

A 5 by 5 grid of nodes, each labelled with its number of hops from the sink in the bottom left corner

Figure 33.1 Hop counts from a corner sink on a 5 by 5 grid: 25 beacons set them all, or 200 transmissions from the centre

Worked example: central against local route setup. Take 25 nodes on a 5 by 5 grid, each linked to its four nearest neighbours, with the sink in one corner. A node in column i and row j (each counted from 0 at the sink) is i + j hops away, so the hop counts run from 0 at the sink to 8 in the far corner.

  • Centrally. Every node reports its neighbours to the sink, and the sink sends each node its route. A report from a node i + j hops away costs i + j transmissions, and so does the reply. Over the whole grid the hop counts add up to 2 × 5 × (0 + 1 + 2 + 3 + 4) = 100, since each of i and j takes each value from 0 to 4 five times. Reports and replies together cost 2 × 100 = 200 transmissions, and all of them again whenever the topology changes.
  • Locally. The sink broadcasts a beacon saying "0 hops". Each node that hears a beacon with the smallest count so far sets its own count one higher and broadcasts once. When the beacons spread outward in order, each of the 25 nodes broadcasts once: 25 transmissions, and every node knows its hop count and a neighbour one hop closer to the sink.
munotes.in196

The Characteristics and Challenges of Ad Hoc Networks in a WSN

200 / 25 = 8: the local method uses an eighth of the transmissions here, and only nodes near a change need to send again. Its routes follow hop counts, not energy or link quality, so they may not be the best, exactly the trade Dargie and Poellabauer describe. The beacons in practice arrive out of order and links are lossy, which is why real protocols refine this idea ([Routing Tables and What Happens When the Topology Changes]).

Quality of service, under all of this

Applications still want their readings on time and complete. But delay grows with every hop and every sleeping relay, delivery falls with every lossy link, and bandwidth is shared. RFC 2501's working group already listed "multicast and QoS extensions for a dynamic, mobile area" among its longer-term issues. For a sensor network the useful notions of quality are about the information (was the event detected? how accurate is the average?) rather than per-packet guarantees ([Optimization Goals: Quality of Service, Energy Efficiency and Lifetime]).

Time and place

A reading is only useful if the network knows when and where it was taken, yet the nodes have no common clock and, often, no GPS. Time synchronisation and localisation are the two challenges an ad hoc sensor network must solve for itself, and they are [Time Synchronisation and Localisation].

What mobility actually does to a route

The characteristic that causes the most trouble is the dynamic topology, and it is the one most often asserted without a number. The program supplies one: ten thousand pairs of nodes are started in range of each other, set walking in independent directions, and timed until the link breaks.

# What mobility does to a route, measured rather than asserted.
import math
import random

RANGE = 250.0                     # metres, a MANET radio
FIELD = 1000.0
STEP = 1.0                        # seconds a step

def link_lifetime(speed, rng):
    """Two nodes start in range and move independently; when do they part?"""
    ax, ay = rng.uniform(0, FIELD), rng.uniform(0, FIELD)
    r = RANGE * math.sqrt(rng.random())
    th = rng.uniform(0, 2 * math.pi)
    bx, by = ax + r * math.cos(th), ay + r * math.sin(th)
    da, db = rng.uniform(0, 2 * math.pi), rng.uniform(0, 2 * math.pi)
    t = 0.0
    while t < 3600:
        ax += speed * math.cos(da) * STEP
        ay += speed * math.sin(da) * STEP
        bx += speed * math.cos(db) * STEP
        by += speed * math.sin(db) * STEP
        t += STEP
        if math.hypot(ax - bx, ay - by) > RANGE:
            return t
    return 3600.0

print("A link between two nodes moving independently, radio range %.0f m." % RANGE)
print("Ten thousand pairs at each speed, each starting in range and walking straight.")
print()
print("  speed                 one link      a 3 hop route   a 6 hop route   a 10 hop route")
for kmh, label in ((2.0, "a slow walk"), (5.0, "walking"), (30.0, "a bicycle"),
                   (60.0, "a car")):
    speed = kmh / 3.6
    rng = random.Random(20260930)
    lives = [link_lifetime(speed, rng) for _ in range(10000)]
    mean = sum(lives) / len(lives)
    # a route of h links lasts as long as its shortest link: the minimum of h draws
    def route_life(h):
        total = 0.0
        for i in range(0, len(lives) - h, h):
            total += min(lives[i:i + h])
        return total / max(1, len(range(0, len(lives) - h, h)))
    print("  %-22s %8.1f s %13.1f s %15.1f s %15.1f s"
          % ("%s, %.0f km/h" % (label, kmh), mean, route_life(3), route_life(6),
             route_life(10)))
print()
print("Two things follow, and both are in the chapter's list of characteristics.")
print("A route is only as durable as its weakest link, so a long route in a moving")
print("network is broken almost as soon as it is found. And the cost of finding it is")
print("paid again every time it breaks:")
REPAIR_S = 0.5
for kmh in (5.0, 60.0):
    speed = kmh / 3.6
    rng = random.Random(20260930)
    lives = [link_lifetime(speed, rng) for _ in range(10000)]
    def route_life(h):
        total = 0.0
        for i in range(0, len(lives) - h, h):
            total += min(lives[i:i + h])
        return total / max(1, len(range(0, len(lives) - h, h)))
    for h in (3, 10):
        life = route_life(h)
        print("  at %4.0f km/h a %2d hop route lasts %6.1f s, so a %.1f s repair is %5.2f%%"
              % (kmh, h, life, REPAIR_S, 100 * REPAIR_S / (life + REPAIR_S)))
        print("     of the time the route is unusable, before any data is carried")
print("  which is the reactive protocols' whole problem, and the reason proactive ones")
print("  exist despite their constant overhead.")
munotes.in197

The Characteristics and Challenges of Ad Hoc Networks in a WSN

A link between two nodes moving independently, radio range 250 m.
Ten thousand pairs at each speed, each starting in range and walking straight.

  speed                 one link      a 3 hop route   a 6 hop route   a 10 hop route
  a slow walk, 2 km/h       529.4 s         151.8 s            82.0 s            52.3 s
  walking, 5 km/h           254.7 s          61.0 s            33.1 s            21.2 s
  a bicycle, 30 km/h         57.5 s          10.6 s             5.9 s             4.0 s
  a car, 60 km/h             31.9 s           5.5 s             3.2 s             2.3 s

Two things follow, and both are in the chapter's list of characteristics.
A route is only as durable as its weakest link, so a long route in a moving
network is broken almost as soon as it is found. And the cost of finding it is
paid again every time it breaks:
  at    5 km/h a  3 hop route lasts   61.0 s, so a 0.5 s repair is  0.81%
     of the time the route is unusable, before any data is carried
  at    5 km/h a 10 hop route lasts   21.2 s, so a 0.5 s repair is  2.30%
     of the time the route is unusable, before any data is carried
  at   60 km/h a  3 hop route lasts    5.5 s, so a 0.5 s repair is  8.26%
     of the time the route is unusable, before any data is carried
  at   60 km/h a 10 hop route lasts    2.3 s, so a 0.5 s repair is 18.04%
     of the time the route is unusable, before any data is carried
  which is the reactive protocols' whole problem, and the reason proactive ones
  exist despite their constant overhead.
munotes.in198

The Characteristics and Challenges of Ad Hoc Networks in a WSN

Distinctions

Characteristic (RFC 2501)What it meansWhat it costs a sensor network
Dynamic topologiesLinks change randomly and rapidly; some are one-wayRoutes break; acknowledgements over one-way links fail
Bandwidth-constrained, variable-capacity linksReal throughput well below the radio's rate; congestion normalRelaying on one channel divides a low rate further; hot spots near the sink
Energy-constrained operationEnergy conservation may be the top design criterionRelays sleep; every control message is paid for in lifetime
Limited physical securityEavesdropping, spoofing, denial of serviceUnattended nodes can be captured; relays can lie
ScalabilityMechanisms specific to ad hoc networksThousands of nodes, with only neighbour-sized memory
Centralised managementDecentralised management
Who decidesThe base station, with global knowledgeEach node, with local knowledge
Quality of the resultCan be optimalUsually not optimal
CostHigh, and paid again on every changeLow, and local to the change
Weak pointThe centre, and the paths to itDecisions made on incomplete information

What it does not mean

Dynamic topology does not require mobility. In a sensor network failures, sleep and fading links change the topology even when nothing moves.

A link that works one way need not work the other. Unidirectional links are one of RFC 2501's own characteristics.

The radio's rate is not the network's throughput. On a shared channel, relaying divides it, and contention and losses divide it again.

Decentralised does not mean uncoordinated. Nodes follow the same local rules, like the hop-count beacons above, and the global structure emerges from them.

munotes.in199

The Characteristics and Challenges of Ad Hoc Networks in a WSN

Quick revision

  • RFC 2501's characteristics: dynamic topologies (random, rapid, some links unidirectional); bandwidth-constrained, variable-capacity links ("congestion is typically the norm"); energy-constrained operation; limited physical security (eavesdropping, spoofing, denial of service; decentralisation avoids single points of failure); plus scalability.
  • Challenges in a WSN: the shared medium (collisions, hidden terminals); multihop routing with no fixed paths and sleeping relays (remedies: wake-up radios, adaptive duty cycling with an awake backbone); self-organisation after ad hoc deployment (self-organisation, self-optimisation, self-protection, self-healing); identifiers (no global IDs; data-centric and local addressing); decentralised management (local decisions, not optimal, cheaper) in small memories; quality of service; time synchronisation and localisation.
  • Chain of 3 hops on one channel, all in range: at most 250 / 3, about 83.3 kbit/s from a 250 kbit/s radio.
  • 5 by 5 grid, corner sink: central setup 200 transmissions against 25 beacons.

Test yourself

1. What are the characteristics of ad hoc networks? Dynamic topologies, in which links change randomly and rapidly and may be unidirectional; bandwidth-constrained, variable-capacity links whose real throughput is well below the radio's maximum, making congestion normal; energy-constrained operation, where conserving energy may be the most important design criterion; and limited physical security, with a higher risk of eavesdropping, spoofing and denial of service, partly offset by decentralised control. Scalability is a further characteristic.

2. Discuss the challenges of ad hoc networks in WSNs. The shared wireless medium causes collisions and hidden terminals; routes must be found through multiple hops with no fixed infrastructure, while relays sleep to save energy; nodes deployed without planning must organise, optimise, protect and heal themselves; nodes may lack global identifiers, so addressing is by data, place or local identifier; management must be decentralised, using local knowledge and small memories; quality of service must be delivered despite delay, loss and shared bandwidth; and the nodes must synchronise their clocks and find their positions themselves.

3. Why does a sensor network have a dynamic topology even when no node moves? Nodes fail and exhaust their batteries, radios are switched off to save energy, and links near the edge of their range vary with noise and fading, so the set of working links keeps changing.

4. Why is decentralised management preferred in WSNs? What is its drawback? Central algorithms need information from every node and must send decisions back, which costs many transmissions and must be repeated whenever the topology changes, and a large, energy-limited network cannot afford that. Local decisions by each node from its neighbours' information are much cheaper and react locally to changes. The drawback is that the results are usually not optimal, because no node sees the whole network.

munotes.in200

The Characteristics and Challenges of Ad Hoc Networks in a WSN

5. Three relays forward readings to a sink over a 250 kbit/s channel, and all four nodes on the path can hear one another. What is the most the path can deliver? Only one node can transmit at a time, and each reading must be transmitted three times, so at most 250 / 3, about 83.3 kbit/s, and less once collisions, back-offs and acknowledgements are counted.

Contents This chapter on its own page

munotes.in201

Chapter Thirty-Four

Time Synchronisation and Localisation

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Characteristics and challenges of ad-hoc networks in WSNs"

In one line

Every node's clock runs at its own rate and every node is placed without a map, so a sensor network must agree on the time by exchanging time-stamped messages (two-way exchanges, or receivers comparing when they heard the same broadcast) and must work out where each node is from a few anchors, by measured distances and trilateration or, more cheaply, by counting hops.

In the wording a student can write in an examination: time synchronisation gives the nodes of a sensor network a common notion of time, needed for fusing readings, for TDMA and duty-cycle schedules, and for ranging by time of flight. Clocks differ by an offset and drift apart because their crystals run at slightly different rates (the skew). A message's delay has four parts (send, access, propagation and receive time), and the variable ones cause synchronisation error. In sender-receiver synchronisation (as in NTP and TPSN) two nodes exchange time-stamped messages and compute the offset and delay from four timestamps; TPSN first builds a hierarchy of levels from a root, then synchronises each node with a node one level up. In receiver-receiver synchronisation (RBS) a node broadcasts a reference beacon and the receivers compare the times at which they heard it, which removes the sender's delays from the error. Localisation finds each node's position from a few anchor nodes that know theirs. Range-based methods measure distances or angles (RSSI, time of arrival, time difference of arrival, angle of arrival) and compute the position by trilateration (multilateration); range-free methods use only connectivity, such as the centroid of the anchors a node can hear, or DV-Hop, which turns hop counts into distances.

Part 1: Time synchronisation

Why the nodes need a common time

RBS's authors list the reasons in one line: synchronisation is critical "for diverse purposes including sensor data fusion, coordinated actuation, and power-efficient duty cycling". Two readings of the same event, from two nodes, can only be combined if their timestamps are comparable. A duty-cycled MAC only works if neighbours wake together, and a TDMA schedule only if every node agrees where the slots begin. TPSN's authors add that the collaborative tasks of a sensor network are "realized by exchanging messages that are timestamped using the local clocks on the nodes". And measuring distance by the time a signal takes to travel needs clocks agreeing to microseconds, as the worked example at the end of this part shows.

Why clocks disagree

A node's clock counts the ticks of a crystal oscillator. Two things make two clocks disagree:

  1. Offset (phase): at a given instant, the difference between their readings. Two nodes switched on at different moments start with an offset.
  2. Skew (drift of rate): each crystal runs at a slightly different frequency, so the offset keeps changing. RBS's authors give the size: typical crystal oscillators are accurate to between one part in ten thousand and one part in a million, so two nodes' clocks drift apart by 1 to 100 microseconds every second. For the Berkeley motes TPSN's authors quote an upper bound of 40 ppm, "i.e. a clock in mote can loose up to 40µs in a second".
munotes.in202

Time Synchronisation and Localisation

Worked example: how fast a clock goes wrong. A clock that is off by 40 ppm gains or loses 40 microseconds every second. In one minute that is 40 × 60 = 2,400 microseconds, 2.4 ms. To stay within 20 microseconds of a correct clock without estimating the drift, it would need resynchronising every 20 / 40 = 0.5 seconds. Hence TPSN's conclusion that "even if we synchronize the whole network once, nodes will go out of sync in a few minutes", and hence RBS's use of a least-squares line through repeated observations, which recovers the drift (the slope) as well as the offset (the intercept), so a node can correct its clock between synchronisations.

Where the error comes from: four parts of a message's delay

Synchronisation means sending a time in a message, and the message takes an uncertain time to arrive. RBS's authors, after Kopetz and Schwabl, split that delay into four parts:

  1. Send time: building the message at the sender, including operating system delays, context switches and handing it to the network interface.
  2. Access time: waiting for the channel, which depends on the MAC protocol: waiting for a clear channel, retransmitting after a collision, or waiting for a TDMA slot.
  3. Propagation time: travelling from sender to receiver. Between neighbours this is tiny, "simply the physical propagation time of the message through the media".
  4. Receive time: the receiver's network interface receiving the message and telling the host.

Send time and access time vary the most, and they are the ones the two families of protocols treat differently. TPSN time-stamps packets "at the moment when they are sent i.e., at MAC layer", which cuts the send time out; RBS removes both, as below.

Sender-receiver synchronisation: the two-way exchange

A two-way message exchange: A sends at T1, B receives at T2 and replies at T3, A receives the reply at T4

Figure 34.1 The two-way exchange: T1 and T4 on A's clock, T2 and T3 on B's

Node A sends a message at time T1, by its own clock. Node B receives it at T2 and replies at T3, both by B's clock, putting T1, T2 and T3 in the reply. A receives the reply at T4, by its own clock. From these four timestamps, RFC 5905, the Network Time Protocol's specification, computes the offset of B relative to A:

munotes.in203

Time Synchronisation and Localisation

theta = ((T2 - T1) + (T3 - T4)) / 2

and the round-trip delay:

delta = (T4 - T1) - (T3 - T2)

TPSN uses the same exchange and the same arithmetic in its equation 1, writing the offset as ((T2 - T1) - (T4 - T3)) / 2 and calling it the clock drift, and the one-way propagation delay as half the round trip. The one assumption is that the delay is the same in both directions.

Worked example. B's clock is actually 6.2 ms ahead of A's, and a message takes 1.1 ms each way. A sends at T1 = 1,000.0 ms. B receives it at T2 = 1,000.0 + 1.1 + 6.2 = 1,007.3 ms on its clock, and replies 2 ms later at T3 = 1,009.3 ms. A receives the reply at T4 = 1,009.3 - 6.2 + 1.1 = 1,004.2 ms on its clock. Then:

  • offset = ((1,007.3 - 1,000.0) + (1,009.3 - 1,004.2)) / 2 = (7.3 + 5.1) / 2 = 6.2 ms, exactly the true offset;
  • round-trip delay = (1,004.2 - 1,000.0) - (1,009.3 - 1,007.3) = 4.2 - 2.0 = 2.2 ms, the two 1.1 ms trips.

When the delays are not equal. Suppose the message takes 1.5 ms going and 0.7 ms returning, the same 2.2 ms round trip. Then T2 = 1,000.0 + 1.5 + 6.2 = 1,007.7 and T4 = 1,009.3 - 6.2 + 0.7 = 1,003.8, and the formula gives (7.7 + 5.5) / 2 = 6.6 ms: wrong by 0.4 ms, which is half the difference between the two delays, (1.5 - 0.7) / 2 = 0.4. The exchange cannot see asymmetry; the variable send and access times are exactly what make the two directions differ.

TPSN: levels first, then pairs

TPSN synchronises a whole network to one root node in two phases.

  1. Level discovery. "The root node is assigned a level 0 and it initiates this phase by broadcasting a level_discovery packet." Each neighbour assigns itself level 1 and broadcasts in turn; each node takes a level one greater than the first it hears and ignores later ones, so every node ends with a level. It is the hop-count beacon of [The Characteristics and Challenges of Ad Hoc Networks in a WSN], used to build a tree.
  2. Synchronisation. "Pair wise synchronization is performed along the edges of the hierarchical structure": each level 1 node runs the two-way exchange with the root and corrects its clock; level 2 nodes then synchronise with level 1 nodes, and so on down. The nodes wait a random time before starting, to avoid contention.

On Berkeley motes TPSN synchronised a pair of neighbours "to an average accuracy of less than 20µs", and its authors argue it "roughly gives a 2x better performance as compared to Reference Broadcast Synchronization (RBS)", which they implemented on the same motes for comparison.

munotes.in204

Time Synchronisation and Localisation

RBS: synchronising the receivers with each other

RBS takes a different route. A node broadcasts a reference beacon, which "does not contain an explicit timestamp; instead, receivers use its arrival time as a point of reference for comparing their clocks". Receivers then exchange the times at which they heard it, and each learns its offset from the others.

Why this helps. "Although the Send Time and Access Time may be unknown, and highly variable from message to message, the nature of a broadcast dictates that for a particular message, these quantities are the same for all receivers." A broadcast is only used "to synchronize a set of receivers with one another", never sender with receiver, so the sender's two most variable delays drop out of the error altogether. What remains is the difference in propagation time between the receivers, which is negligible over tens of metres, and the difference in their receive times.

How well it did. On off-the-shelf 802.11 hardware, RBS achieved "1.85 ± 1.28µsec", and across four hops "3.68 ± 2.57µsec", a significant improvement over NTP under similar conditions. The authors also note that the broadcast "does not even need to be a dedicated timesync packet": any broadcast the network already sends, such as a route discovery packet, can serve.

Worked example: synchronisation for localisation

Ultrasonic ranging measures how long sound takes to cross from one node to another. TPSN's authors turn a timing error into a distance error with sound at 345 m/s: at their average accuracy of 20 microseconds, 345 × 0.00002 = 0.0069 m, that is 0.69 cm. At a millisecond of error, 345 × 0.001 = 0.345 m, about 35 cm. That is how closely localisation and synchronisation are tied.

Part 2: Localisation

Why nodes must find their own positions

A reading is only useful with a place attached, geographic routing needs positions ([Geographic Routing: Greedy Forwarding and GPSR]), and coverage cannot be judged without them. GPS solves the problem outdoors for larger devices, but Dargie and Poellabauer note that the "need for small form factor and low energy consumption also prohibits the integration of many desirable components, such as GPS receivers". So a few nodes, called anchors (also beacons, reference points or landmarks), know their positions, from GPS or because they were placed by hand, and every other node works out its own from them.

The methods fall into two families. Bulusu and colleagues call them fine-grained, inferring "the distance to a reference point based on signal strength or timing measurements", and coarse-grained, inferring only "proximity to a given reference point". The survey by Mesmoudi and colleagues uses the names common today: range-based and range-free.

munotes.in205

Time Synchronisation and Localisation

Range-based: four ways to measure

  1. RSSI (received signal strength). The stronger the signal, the nearer the sender: a propagation model converts strength into distance. It is "the most common techniques, cheapest and simplest", since it needs no extra hardware, but it is "very susceptible to noise and obstacles".
  2. Time of arrival (TOA), or time of flight: distance is travel time multiplied by the signal's speed. It needs the sender and receiver synchronised, which "adds cost and complexity".
  3. Time difference of arrival (TDOA). In one form, a node sends a radio signal and an ultrasound signal together; the radio arrives almost at once, the sound much later, and the difference gives the distance: speed of sound times (sound's travel time minus radio's travel time). It needs extra hardware, and "the ultrasound signal can be stopped by obstacles". In another form, differences in arrival time at pairs of anchors place the node on hyperbolas whose intersection is its position.
  4. Angle of arrival (AOA): the direction a signal comes from, measured with an antenna array or several receivers. Accurate, but needs extra hardware and suffers from multipath.

Worked example: why RSSI ranging is rough. In the log-normal shadowing model, the received signal at a given distance varies around its average by a random amount with standard deviation σ dB, and it falls by 10n dB for every tenfold increase in distance. So a reading σ dB too strong makes the distance look shorter by a factor of 10 raised to σ / (10n). With the values Zuniga and Krishnamachari used, σ = 4 dB and n = 4, that factor is 10 to the power 0.1, about 1.26: one typical shadowing error puts a node about 26 per cent too far or about 21 per cent too near. An error of that size in every range is what the trilateration below has to absorb.

Trilateration

Three anchors, each with a circle of its measured distance, meeting at the unknown node

Figure 34.2 Trilateration: the node lies where the circles of its distances from the anchors meet

A node that knows its distance d1 from an anchor at (x1, y1) lies on a circle around it. Two circles meet in two points; a third anchor picks one. With more anchors, and with measured distances that are never exact, the circles do not meet in one point, and the position is found by least squares: the point that fits all the circles best. The standard trick is algebraic. Each circle's equation is

(x - xi)² + (y - yi)² = di²

munotes.in206

Time Synchronisation and Localisation

and subtracting the last anchor's equation from each of the others cancels the x² and y² terms, leaving equations that are linear in x and y. Those are solved by ordinary least squares. The program below does exactly that. Using more than three anchors is multilateration; the arithmetic is the same.

Range-free: centroid and DV-Hop

Centroid. Bulusu and colleagues place reference points on a grid, each broadcasting beacons periodically. A node listens for a while and computes, for each reference point, a connectivity metric: the percentage of its beacons it received. It treats the reference points whose metric exceeds a threshold ("say 90%") as nearby, and places itself at their centroid, the average of their coordinates. Their measurements found "the accuracy for 90% of our data points is within one-third of the separation distance" between reference points.

Worked example. Four reference points stand at (0, 0), (10, 0), (0, 10) and (10, 10), 10 m apart, with a reliable range of 12 m.

  • A node at (3, 4) is 5 m, about 8.1 m, about 6.7 m and about 9.2 m from them, so it hears all four and places itself at their centroid, (5, 5). Its error is the distance from (3, 4) to (5, 5), 2 m across and 1 m up, the square root of 5, about 2.24 m.
  • A node at (1, 1) is about 12.7 m from (10, 10), out of range, so it hears three. Their centroid is ((0 + 10 + 0) / 3, (0 + 0 + 10) / 3), about (3.33, 3.33), and its error is about 3.30 m, a third of the separation.

The centroid method needs no measurement at all, only counting beacons, and its accuracy is set by how closely the reference points are spaced.

DV-Hop. Niculescu and Nath's DV-Hop works over many hops, where most nodes hear no anchor directly. The survey describes it in three steps:

  1. Hop counts. By a distance-vector exchange, "all nodes in the network get minimal hop-count to every anchor nodes".
  2. Hop size. Each anchor, knowing the true distances to the other anchors and the hop counts to them, computes an average distance per hop: the sum of those distances divided by the sum of those hop counts. It sends this hop size to the nodes around it.
  3. Position. An unknown node multiplies its hop count to each anchor by the hop size, giving an estimated distance to each, and with three or more of these it uses trilateration.

It is simple and "does not depend on range measurement error", but a hop is only an average distance: a straight chain of nodes and a crooked one with the same hop count get the same estimate.

munotes.in207

Time Synchronisation and Localisation

Localisation, computed

The program first trilaterates one node from four corner anchors, with exact and then with slightly wrong distances. It then scatters 100 nodes over a 100 m square with a radio range of 18 m, makes the first 10 of them anchors, and runs DV-Hop's three steps for the other 90.

# Locating nodes from anchors: trilateration from measured distances, and
# DV-Hop (Niculescu and Nath) when a node can only count hops.
import math
import random
from collections import deque

def trilaterate(anchors, dists):
    """Least squares. Subtracting the last circle's equation from each of
    the others leaves equations that are linear in x and y."""
    (xn, yn), dn = anchors[-1], dists[-1]
    rows = [(2 * (xn - xi), 2 * (yn - yi),
             di**2 - dn**2 - xi**2 + xn**2 - yi**2 + yn**2)
            for (xi, yi), di in zip(anchors[:-1], dists[:-1])]
    saa = sum(a * a for a, b, c in rows)
    sab = sum(a * b for a, b, c in rows)
    sbb = sum(b * b for a, b, c in rows)
    sac = sum(a * c for a, b, c in rows)
    sbc = sum(b * c for a, b, c in rows)
    det = saa * sbb - sab * sab
    return (sac * sbb - sbc * sab) / det, (saa * sbc - sab * sac) / det

# 1. Trilateration: a node at (30, 40) and four anchors at the corners.
corners = [(0, 0), (100, 0), (0, 100), (100, 100)]
true = (30, 40)
exact = [math.dist(true, a) for a in corners]
print("exact distances   -> (%.2f, %.2f)" % trilaterate(corners, exact))
errors = [1.05, 0.92, 1.10, 0.97]          # measured 5% long, 8% short, ...
noisy = [d * e for d, e in zip(exact, errors)]
x, y = trilaterate(corners, noisy)
print("distances 3-10%% off -> (%.2f, %.2f), %.2f m from the truth"
      % (x, y, math.dist((x, y), true)))

# 2. DV-Hop: 100 nodes scattered over 100 m x 100 m, radio range 18 m;
#    the first 10 know their positions (the anchors).
rnd = random.Random(5)
pos = [(rnd.uniform(0, 100), rnd.uniform(0, 100)) for _ in range(100)]
RANGE, ANCHORS = 18.0, range(10)
nbrs = [[j for j in range(100) if j != i and math.dist(pos[i], pos[j]) <= RANGE]
        for i in range(100)]

def hops_from(src):                        # step 1: flood hop counts
    h, todo = {src: 0}, deque([src])
    while todo:
        n = todo.popleft()
        for m in nbrs[n]:
            if m not in h:
                h[m] = h[n] + 1
                todo.append(m)
    return h

hops = {a: hops_from(a) for a in ANCHORS}
hop_size = {}                              # step 2: metres per hop, per anchor
for a in ANCHORS:
    others = [b for b in ANCHORS if b != a and b in hops[a]]
    hop_size[a] = (sum(math.dist(pos[a], pos[b]) for b in others)
                   / sum(hops[a][b] for b in others))
errs = []
for n in range(10, 100):                   # step 3: estimate, then trilaterate
    known = [a for a in ANCHORS if n in hops[a]]
    if len(known) < 3:
        continue
    nearest = min(known, key=lambda a: hops[a][n])
    est = [hop_size[nearest] * hops[a][n] for a in known]
    guess = trilaterate([pos[a] for a in known], est)
    errs.append(math.dist(guess, pos[n]))
print("DV-Hop: hop sizes %.1f to %.1f m; %d of 90 nodes located"
      % (min(hop_size.values()), max(hop_size.values()), len(errs)))
print("        mean error %.1f m, that is %.2f of the radio range; worst %.1f m"
      % (sum(errs) / len(errs), sum(errs) / len(errs) / RANGE, max(errs)))
munotes.in208

Time Synchronisation and Localisation

exact distances   -> (30.00, 40.00)
distances 3-10% off -> (36.92, 37.20), 7.46 m from the truth
DV-Hop: hop sizes 9.3 to 12.5 m; 90 of 90 nodes located
        mean error 14.2 m, that is 0.79 of the radio range; worst 47.2 m

Reading it. With exact distances, least squares returns the node's true position. With distances only 3 to 10 per cent wrong, the position moves by 7.46 m: ranging errors grow into position errors, which is why RSSI's rough ranges limit it. DV-Hop needs no ranging hardware at all and still places every one of the 90 nodes, but with a mean error of 14.2 m, 0.79 of the radio range, and a worst case of 47.2 m. The hop sizes, 9.3 to 12.5 m, are only about a half to two thirds of the 18 m range, because a hop in a random network rarely covers the full range, and a node whose shortest paths bend around a gap gets hop counts that overstate its distances. Better anchor placement, more anchors, or ranging where it is affordable all reduce the error; that trade between accuracy and cost is the whole subject of localisation.

Distinctions

Sender-receiver (NTP, TPSN)Receiver-receiver (RBS)
Who is synchronisedThe sender with the receiverThe receivers of one broadcast with each other
MessagesA two-way exchange per pairOne broadcast, then receivers exchange arrival times
Delays in the errorSend, access, propagation and receive times (TPSN stamps at the MAC to cut send time)Only the differences in propagation and receive times
Timestamp in the messageYesNo
OffsetSkew (drift)
What it isThe difference between two clocks at one instantThe difference in their rates
Corrected byOne exchangeRepeated observations, a fitted line
Mote figureAnything at switch-onUp to 40 ppm, 40 microseconds per second (TPSN)
Range-basedRange-free
MeasuresDistance or angle (RSSI, TOA, TDOA, AOA)Only connectivity or hop counts
HardwareSometimes extra (ultrasound, antenna arrays)None beyond the radio
AccuracyHigher, if ranging is goodLower; depends on anchor density
ExamplesTrilateration from measured rangesCentroid, DV-Hop, APIT
munotes.in209

Time Synchronisation and Localisation

What it does not mean

Synchronised does not mean correct. A network synchronised to its root agrees with itself; it need not agree with the real time unless the root does.

One synchronisation is not enough. Clocks drift apart again at up to tens of microseconds per second, so synchronisation is repeated, or the drift is estimated and corrected.

RBS does not need a special packet. Any broadcast that several receivers hear can be used as the reference.

Range-free is not error-free. It avoids measurement error by not measuring, and pays in resolution: a hop or a beacon's range is a coarse unit of distance.

More anchors are not free. Each anchor needs GPS or a manual survey; localisation methods are judged by how few anchors they need for a given accuracy.

Quick revision

  • Synchronisation is needed for data fusion, coordinated actuation, duty cycling (RBS), time-stamped cooperation (TPSN), and ranging.
  • Offset and skew; crystals accurate to 1 part in 10 thousand to 1 part in a million, 1 to 100 microseconds per second; motes up to 40 ppm.
  • Delay components (Kopetz and Schwabl, via RBS): send, access, propagation, receive.
  • Two-way exchange (RFC 5905): offset = ((T2 - T1) + (T3 - T4)) / 2, delay = (T4 - T1) - (T3 - T2); assumes symmetric delay; asymmetry errs by half the difference.
  • TPSN: level discovery from a root, then pairwise synchronisation down the levels; MAC-layer time-stamps; under 20 microseconds; about 2x better than RBS.
  • RBS: a reference broadcast with no timestamp; receivers compare arrival times; removes send and access time; 1.85 ± 1.28 microseconds on 802.11; regression estimates skew.
  • Ranging error = speed of sound × timing error: 345 m/s × 20 microseconds = 0.69 cm.
  • Localisation: anchors; range-based (RSSI, TOA, TDOA, AOA) and range-free (centroid, DV-Hop).
  • Trilateration: circles of measured distance; least squares after subtracting one equation; multilateration with more anchors.
  • Centroid: connectivity metric above 90 per cent; average of heard reference points; 90 per cent of estimates within one third of the spacing.
  • DV-Hop: hop counts to anchors; hop size = sum of anchor distances / sum of hops; distance = hop size × hops; trilaterate.

Test yourself

1. Why is time synchronisation needed in sensor networks? Explain the sources of error. Readings from different nodes must be combined in time order, duty-cycled and TDMA schedules need neighbours to agree when to wake, actions must be coordinated, and ranging by time of flight needs microsecond agreement. Clocks start with different offsets and run at slightly different rates. A synchronisation message's delay has four parts: send time at the sender, access time waiting for the channel, propagation time, and receive time at the receiver; the variation in these, especially send and access time, becomes synchronisation error.

munotes.in210

Time Synchronisation and Localisation

2. Explain the two-way message exchange and derive the offset. A sends at T1 by its clock; B receives at T2 and replies at T3 by its clock; A receives at T4. If the offset of B is theta and the one-way delay d is the same both ways, T2 = T1 + d + theta and T4 = T3 + d - theta. Subtracting, (T2 - T1) - (T4 - T3) = 2 theta, so theta = ((T2 - T1) + (T3 - T4)) / 2, and the round-trip delay is (T4 - T1) - (T3 - T2).

3. Explain TPSN. The Timing-sync Protocol for Sensor Networks works in two phases. In level discovery, the root takes level 0 and broadcasts; each node takes a level one greater than the first level it hears and rebroadcasts, building a hierarchy. In the synchronisation phase, each node performs a two-way exchange with a node one level above and corrects its clock, level by level from the root down. It time-stamps messages at the MAC layer, and achieved under 20 microseconds between neighbouring motes.

4. Compare RBS with sender-receiver synchronisation. In RBS a node broadcasts a beacon with no timestamp, and the receivers compare the times at which they received it, synchronising with each other. Because the broadcast leaves the sender once, its send time and access time are the same for all receivers and drop out of the error, leaving only differences in propagation and receive time. Sender-receiver protocols synchronise a receiver to a sender with a two-way exchange and must live with the sender's variable delays, unless, like TPSN, they time-stamp at the MAC layer.

5. Explain range-based and range-free localisation with one example of each. Range-based methods measure the distance or angle to anchors, using received signal strength, time of arrival, time difference of arrival or angle of arrival, and compute the position by trilateration. Range-free methods use only connectivity: in the centroid method a node places itself at the average position of the anchors it hears reliably; in DV-Hop a node counts hops to each anchor, converts hops to distances with an average hop size computed by the anchors, and trilaterates.

6. A node hears reference points at (0, 0), (10, 0) and (0, 10) but not the one at (10, 10). Where does the centroid method place it? At the average of the three: ((0 + 10 + 0) / 3, (0 + 0 + 10) / 3), about (3.33, 3.33).

Contents This chapter on its own page

munotes.in211

Chapter Thirty-Five

Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Ad-hoc networks in WSNs"

In one line

Proactive protocols keep a route to every node at all times by exchanging tables periodically, so a route is ready at once but the overhead never stops; reactive protocols find a route only when one is needed, so they are quiet when idle but make the first packet wait; hybrid protocols do the first near home and the second far away.

In the wording a student can write in an examination: ad hoc routing protocols fall into three classes. Proactive or table-driven protocols, such as DSDV and OLSR, maintain routes to all destinations continuously by exchanging routing information periodically and when the topology changes; routes are available immediately, but control traffic flows even when no data does. Reactive or on-demand protocols, such as AODV and DSR, discover a route only when a source needs one, by flooding a route request and receiving a route reply, and maintain it only while it is used; they have little overhead when traffic is light, at the cost of a route discovery delay. Hybrid protocols, such as ZRP, keep routes proactively within a local routing zone and discover routes beyond it reactively. The choice depends on how much traffic there is, how fast the topology changes, and how much delay the application can accept.

Why Internet routing does not simply carry over

Internet routing protocols were designed, in RFC 2501's words, for "the higher-speed, semi-static topology of the fixed Internet". An ad hoc network breaks both assumptions. Its links change as nodes move, fail or sleep, and every control message spends scarce bandwidth and battery. DSDV's authors started from the Internet's distance-vector routing (the Bellman-Ford method used by RIP) and set out to fix its "poor looping properties" when links break, which in a fixed network is a rare event and in an ad hoc network an everyday one.

RFC 2501 lists the properties a MANET routing protocol should have:

  1. Distributed operation: "This is an essential property, but it should be stated nonetheless."
  2. Loop-freedom, so that packets do not circulate.
  3. Demand-based operation: adapt to the traffic, not maintain routes between all nodes at all times, which uses energy and bandwidth more efficiently "at the cost of increased route discovery delay".
  4. Proactive operation: "The flip-side of demand-based operation", for when that delay "may be unacceptable".
  5. Security, since routing messages are easy to snoop, replay and forge over the air.
  6. "Sleep" period operation: coping with nodes that stop transmitting and receiving to save energy.
  7. Unidirectional link support.

Items 3 and 4 are the two families of this chapter; a hybrid protocol tries to have both.

munotes.in212

Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid

Proactive routing

DSDV: distance vector with sequence numbers

DSDV (Destination-Sequenced Distance-Vector) makes every mobile host a router "which periodically advertises its view of the interconnection topology". Each advertisement lists, for each destination, the number of hops to it and a sequence number "as originally stamped by the destination". The rules that make it work:

  1. Newer beats better. "Routes with more recent sequence numbers are always preferred"; "Of the paths with the same sequence number, those with the smallest metric will be used." A route learnt from a neighbour costs one more hop than the neighbour advertised.
  2. Destinations stamp even numbers. Each node advertises itself with a new, even sequence number each time it broadcasts.
  3. Breaks are stamped odd. When a link breaks, the node that notices gives every route through it a metric of infinity and a sequence number one higher, which is odd: "Building information to describe broken links is the only situation when the sequence number is generated by any Mobile Host other than the destination Mobile Host." The bad news therefore beats the old good news, and the destination's next even number beats the bad news.
  4. Full dumps and incrementals. A "full dump" carries the whole table and is sent relatively infrequently; an "incremental" carries only what changed since the last full dump and should fit in one packet.

DSDV, run. Six nodes: A, B, C, D, E and F, with links A-B, B-C, C-D, D-F, B-E and E-D, so B can reach D through C or through E. Every round, each node stamps itself with a new even number and broadcasts its whole table; after three rounds the link C-D breaks.

# DSDV, a proactive protocol: every node advertises its whole routing table
# every round, and the destination's sequence number says which news is newer.
INF = float("inf")
nodes = "ABCDEF"
links = {("A", "B"), ("B", "C"), ("C", "D"), ("D", "F"), ("B", "E"), ("E", "D")}

def neighbours(n):
    return sorted(b if a == n else a for a, b in links if n in (a, b))

table = {n: {n: (n, 0, 0)} for n in nodes}   # destination: (next hop, hops, seq)

def route(n, dest):
    if dest not in table[n]:
        return "%s to %s: unknown" % (n, dest)
    nxt, hops, seq = table[n][dest]
    return "%s to %s: via %s, %s hops, seq %d" % (
        n, dest, nxt, "inf" if hops == INF else hops, seq)

def advertise(rnd):
    for n in nodes:                           # each node stamps itself: new even seq
        table[n][n] = (n, 0, table[n][n][2] + 2)
    sent = {n: dict(table[n]) for n in nodes} # full dumps, all at once
    entries = sum(len(t) for t in sent.values())
    for n in nodes:
        for nb in neighbours(n):
            for dest, (_, hops, seq) in sent[nb].items():
                cur = table[n].get(dest)
                if cur is None or seq > cur[2] or (seq == cur[2] and hops + 1 < cur[1]):
                    table[n][dest] = (nb, hops + 1, seq)
    print("round %d  %-30s  %-30s  %2d entries broadcast" % (rnd, route("A", "D"),
                                                        route("B", "D"), entries))

def break_link(a, b):
    links.discard((a, b) if (a, b) in links else (b, a))
    for x, y in ((a, b), (b, a)):             # both ends mark routes through the other
        for dest, (nxt, hops, seq) in table[x].items():
            if nxt == y:
                table[x][dest] = (nxt, INF, seq + 1)    # odd seq: a broken route

for r in range(1, 4):
    advertise(r)
break_link("C", "D")
print("-- the link from C to D breaks --")
for r in range(4, 7):
    advertise(r)
munotes.in213

Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid

round 1  A to D: unknown                 B to D: unknown                  6 entries broadcast
round 2  A to D: unknown                 B to D: via C, 2 hops, seq 2    18 entries broadcast
round 3  A to D: via B, 3 hops, seq 2    B to D: via C, 2 hops, seq 4    30 entries broadcast
-- the link from C to D breaks --
round 4  A to D: via B, 3 hops, seq 4    B to D: via C, inf hops, seq 7  34 entries broadcast
round 5  A to D: via B, inf hops, seq 7  B to D: via E, 2 hops, seq 8    36 entries broadcast
round 6  A to D: via B, 3 hops, seq 8    B to D: via E, 2 hops, seq 10   36 entries broadcast

Reading the run.

  1. Routes spread one hop per round. B, two hops from D, learns its route in round 2; A, three hops away, in round 3. The sequence number A holds (2) is older than B's (4) because D's news takes longer to reach it.
  2. The break is stamped odd. When C loses its link to D, it marks D unreachable with sequence number 7, one more than the 6 it had. In round 4, B hears C's odd 7, which is newer than anything else it has, and marks D unreachable too, even though E still offers an older route.
  3. The destination's next even number repairs it. D stamped itself 8 in round 4; that reaches B through E in round 5, beating the odd 7, so B routes through E. A hears of the break in round 5 and of the repair in round 6. No route ever loops, because a stale route always carries an older number.
  4. The overhead never stops. Once every node knows every destination, 6 nodes broadcast 6 entries each, 36 entries every round, whether anything has changed or not. That standing cost is the price of having every route ready.
munotes.in214

Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid

OLSR: link state, flooded economically

OLSR (Optimized Link State Routing) is "a table driven, proactive protocol". A link-state protocol floods information about links through the whole network, which is expensive if every node rebroadcasts. OLSR's key idea is the multipoint relay (MPR): each node picks a subset of its neighbours such that every node two hops away can be reached through one of them, and "only nodes, selected as such MPRs, are responsible for forwarding control traffic". The RFC names three savings: fewer retransmissions when flooding, link-state messages generated only by MPRs, and MPRs reporting only the links to the nodes that chose them. It still "provides optimal routes (in terms of number of hops)", and is "particularly suitable for large and dense networks". OLSR picks MPRs only among neighbours with bidirectional links, which keeps its routes off one-way links. How flooding itself can be made cheaper is taken further in [Flooding, Gossiping and the Broadcast Storm].

Reactive routing

Reactive protocols keep nothing they are not using. AODV (Ad hoc On-Demand Distance Vector) "does not require nodes to maintain routes to destinations that are not in active communication"; DSR (Dynamic Source Routing) does not use "any periodic routing advertisement". Both have the same two mechanisms:

  1. Route discovery. A source with no route floods a route request (RREQ) through the network. When it reaches the destination, or a node that knows a fresh route to it, a route reply (RREP) comes back to the source along the path the request took.
  2. Route maintenance. While the route is in use, a node that finds the next link broken reports it (a route error), and the source finds another route or discovers again.

They differ in where the route lives. In AODV each node on the path keeps a table entry, the next hop towards the destination, and AODV borrows DSDV's destination sequence numbers to stay loop-free. In DSR the source puts the whole route in every packet's header, and nodes keep a route cache of routes they have learnt. AODV is traced step by step in [AODV: Route Discovery and Route Maintenance], and DSR, with the overhead of the two compared, in [DSR, and What Its Routes Cost Against AODV].

Hybrid routing: ZRP

The Zone Routing Protocol "limits the scope of the proactive procedure only to the node's local neighborhood", and searches beyond it reactively.

  • The routing zone. "A node's routing zone is defined as a collection of nodes whose minimum distance in hops from the node in question is no greater than a parameter referred to as the zone radius." Each node has its own zone, so neighbours' zones overlap.
  • IARP (Intrazone Routing Protocol) keeps routes inside the zone proactively, learning its neighbours from periodic "hello" beacons.
  • IERP (Interzone Routing Protocol) finds routes outside the zone reactively. Its queries are not flooded to everyone but bordercast: sent to the peripheral nodes, those exactly the zone radius away, which look up their own zones and bordercast further.
munotes.in215

Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid

Node A's routing zone of radius 2: nodes B to F inside, D, E and F on its edge, G outside

Figure 35.1 The ZRP draft's example: A's zone of radius 2

In the draft's own example, redrawn above, B and C are one hop from A; D and F are two hops away through C, and E two hops away through B (and three through C and F, but the shortest path counts). So B to F are in A's zone, D, E and F are its peripheral nodes, and G, three hops away, is outside.

The zone radius is the dial. "Large routing zones are preferred when demand for routes is high and/or the network consists of many slowly moving nodes", and for a fixed network the ideal radius "would be infinitely large", which is purely proactive routing. Smaller zones suit low route demand and fast-moving nodes, and with a radius of one hop "the ZRP defaults to a traditional reactive flooding protocol". The draft's claim for a well-chosen radius is that ZRP "can perform at least as well as (and often better than) its purely proactive and reactive constituent protocols".

Worked example: standing overhead against route delay

Take a network of 50 nodes and compare the control transmissions per hour. These are our own round numbers, chosen to show the shape of the trade.

Proactive. Every node broadcasts its table every 10 seconds (DSDV's authors suggest "once every few seconds"). That is 3,600 / 10 = 360 broadcasts per node per hour, and 50 × 360 = 18,000 broadcasts per hour for the network, whatever the traffic.

Reactive. Each route discovery floods one request, rebroadcast once by every node (50 transmissions), and returns one reply along a 5-hop path (5 transmissions): 55 transmissions per discovery.

  • With 20 discoveries an hour: 20 × 55 = 1,100 transmissions, about a sixteenth of the proactive cost.
  • With 400 discoveries an hour (many flows, or routes breaking often): 400 × 55 = 22,000, more than the proactive cost.
  • The two are equal at 18,000 / 55 discoveries per hour, about 327, roughly one every 11 seconds.

The delay. A proactive node sends its first packet at once. A reactive node first waits for the request to travel 5 hops out and the reply 5 hops back: 10 hop-times. At an assumed 10 ms per hop, that is 10 × 10 = 100 ms before the first data packet leaves, the "route acquisition time" RFC 2501 lists among the measures of a protocol.

munotes.in216

Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid

So: light, occasional traffic and a changing topology favour reactive routing; heavy traffic between many pairs, or an application that cannot wait, favours proactive routing; a hybrid sets the boundary between the two with its zone radius.

Where sensor networks fit

Most sensor network traffic flows from many nodes to one sink, not between arbitrary pairs. A sink that periodically sends a beacon down a collection tree is doing something proactive, but only towards one destination, which is far cheaper than keeping routes to everyone. That is why sensor networks mostly use their own routing protocols, from [Routing Challenges and Design Issues in WSNs] onwards, rather than DSDV or AODV as they are.

Distinctions

Proactive (table-driven)Reactive (on-demand)Hybrid
Routes keptTo every destination, alwaysOnly those in useProactive inside the zone, reactive beyond
Control trafficPeriodic and on change, even when idleOnly when a route is needed or breaksPeriodic inside zones; queries beyond
First packet delayNoneA route discoveryNone inside the zone; discovery beyond
SuitsHeavy traffic, low mobility, delay-sensitive useLight traffic, high mobilityNetworks that vary, tuned by zone radius
ExamplesDSDV, OLSRAODV, DSRZRP
DSDVAODVDSR
FamilyProactiveReactiveReactive
Route stored asNext hop in every node's tableNext hop in the nodes on the pathWhole route in the source's cache and each packet's header
Loop preventionDestination sequence numbersDestination sequence numbersThe source route itself
Periodic messagesTable broadcastsOptional hello messagesNone

What it does not mean

Reactive does not mean slow in general. Once a route is found, packets flow as fast as on a proactive route; only the first packet, or the first after a break, waits.

Proactive does not mean every route is correct. Tables lag behind the topology by the time news takes to spread, as A's table did in the run above.

A sequence number is not a timestamp. It orders pieces of news about one destination, set by that destination (or, for a break, one higher by the node that saw it); it says nothing about the time of day.

Hybrid is not automatically better. ZRP's advantage depends on choosing the zone radius to suit the network.

Quick revision

  • Internet routing assumes a "semi-static topology"; ad hoc links change, and control messages cost energy and bandwidth.
  • RFC 2501's properties: distributed operation, loop-freedom, demand-based operation, proactive operation, security, sleep period operation, unidirectional link support.
  • Proactive / table-driven: routes to all, always; DSDV (periodic tables, destination sequence numbers: newer wins, then fewer hops; even from the destination, odd with infinity for a break; full dumps and incrementals); OLSR (link state with MPRs that alone relay floods).
  • Reactive / on-demand: route discovery (flooded RREQ, RREP back) and route maintenance (route error); AODV (next hops, sequence numbers), DSR (source routes, route cache, no periodic messages).
  • Hybrid: ZRP, routing zone of radius in hops; IARP proactive inside, IERP reactive outside by bordercasting to peripheral nodes; radius 1 = reactive flooding; large radius = proactive.
  • Trade-off: standing overhead (proactive) against route discovery delay (reactive). In the example: 18,000 broadcasts an hour against 55 per discovery; equal at about 327 discoveries an hour.
munotes.in217

Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid

Test yourself

1. Classify ad hoc routing protocols with one example of each. Proactive or table-driven protocols maintain routes to all destinations at all times by exchanging routing information periodically (DSDV, OLSR). Reactive or on-demand protocols find a route only when a source needs one, by a route request and reply, and maintain it only while it is used (AODV, DSR). Hybrid protocols are proactive within a local zone and reactive beyond it (ZRP).

2. Explain DSDV. How does it avoid routing loops? Each node keeps a table with a next hop, a hop count and a sequence number for every destination, and broadcasts its table periodically (full dumps) and when routes change (incrementals). Each destination stamps its own entry with increasing even sequence numbers. A node prefers the route with the newest sequence number and, among equals, the fewest hops. When a link breaks, the node that detects it sets the routes through it to an infinite metric with an odd sequence number one higher, which supersedes the old routes until the destination's next even number arrives. Since stale routes always carry older numbers, they are never chosen over fresh ones, so loops do not form.

3. Compare proactive and reactive routing. Proactive routing has routes ready at once but sends control traffic continuously, even when there is no data, and its tables must track every change. Reactive routing sends control traffic only when a route is needed or breaks, which saves energy and bandwidth when traffic is light, but the first packet waits for a route discovery, and each discovery floods the network. Proactive suits heavy, delay-sensitive traffic with low mobility; reactive suits light traffic and high mobility.

4. Explain the Zone Routing Protocol. Each node has a routing zone, all nodes within a set number of hops (the zone radius). Inside the zone, routes are kept proactively by the Intrazone Routing Protocol. For destinations outside, the Interzone Routing Protocol sends route queries reactively, bordercasting them to the peripheral nodes at the edge of the zone, which consult their own zones. The radius tunes the protocol: large for high demand and slow nodes, small for low demand and fast nodes; a radius of one hop is plain reactive flooding.

munotes.in218

Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid

5. What are multipoint relays in OLSR? A subset of a node's bidirectional neighbours, chosen so that every node two hops away can be reached through at least one of them. Only MPRs rebroadcast flooded control messages and generate link-state information, which greatly reduces OLSR's control traffic compared with classic link-state flooding.

Contents This chapter on its own page

munotes.in219

Chapter Thirty-Six

AODV: Route Discovery and Route Maintenance

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Ad-hoc networks in WSNs" (and the paired practical, "Simulate the AODV routing protocol and analyze route discovery and route maintenance mechanisms")

In one line

AODV finds a route only when a source needs one, by flooding a route request that leaves a trail of reverse routes and returning a route reply along that trail, and it keeps the route only while it is used, repairing it with route errors, all numbered by destination sequence numbers so that no route loops.

In the wording a student can write in an examination: AODV (Ad hoc On-Demand Distance Vector) is a reactive routing protocol for MANETs, specified in RFC 3561. It uses three messages: Route Request (RREQ), Route Reply (RREP) and Route Error (RERR). In route discovery, a source that needs a route broadcasts an RREQ; each node that receives it for the first time records a reverse route to the source through the neighbour it heard it from, and rebroadcasts it; duplicates, recognised by the originator's address and RREQ ID, are discarded. When the RREQ reaches the destination, or an intermediate node with a fresh enough route, that node unicasts an RREP back along the reverse route, and each node on the way sets up a forward route to the destination. Each node keeps only the next hop for each destination. Destination sequence numbers, created by the destination, show which route information is newer and so guarantee loop freedom. In route maintenance, a node that detects a broken link marks the routes through it invalid, increments their sequence numbers and sends an RERR to the neighbours in its precursor list, and the source starts a new discovery if it still needs the route.

Why AODV, and what it keeps

RFC 3561's introduction states the aims: AODV "enables dynamic, self-starting, multihop routing", "allows mobile nodes to obtain routes quickly for new destinations, and does not require nodes to maintain routes to destinations that are not in active communication". Its operation "is loop-free, and by avoiding the Bellman-Ford 'counting to infinity' problem offers quick convergence when the ad hoc network topology changes".

And when routes work, AODV is silent: "As long as the endpoints of a communication connection have valid routes to each other, AODV does not play any role."

A route table entry holds: the destination's address, its destination sequence number and a flag saying whether that number is valid, state flags (valid, invalid, being repaired), the network interface, the hop count, the next hop, the list of precursors (the neighbours likely to use this node as their next hop to that destination), and a lifetime.

The three messages

RREQ (type 1) carries a hop count, an RREQ ID ("A sequence number uniquely identifying the particular RREQ when taken in conjunction with the originating node's IP address"), the destination's address and the last destination sequence number the originator knows, and the originator's address and its own sequence number. Five flag bits are defined: J and R (reserved for multicast), G (ask intermediate nodes to send a gratuitous RREP to the destination too), D (destination only: only the destination may reply) and U (unknown: the originator knows no sequence number for the destination). Six 32-bit words: 24 bytes.

munotes.in220

AODV: Route Discovery and Route Maintenance

RREP (type 2) carries the hop count, the destination's address and sequence number, the originator's address, and a lifetime for the route.

RERR (type 3) lists the unreachable destinations, each with its sequence number, and a count of them.

All three travel over UDP; broadcast messages use the limited broadcast address 255.255.255.255, and how far an RREQ spreads is set by the TTL in its IP header.

Route discovery

1. The source asks. "A node disseminates a RREQ when it determines that it needs a route to a destination and does not have one available." Before sending it, the source increments its own sequence number and its RREQ ID, sets the hop count to zero, and records the pair (its own address, the RREQ ID) so it will not process its own request when neighbours rebroadcast it.

2. Each node that hears it. A node receiving an RREQ first checks whether it has already seen one with the same originator address and RREQ ID; if so, it "silently discards the newly received RREQ". Otherwise it increments the hop count and creates or updates a reverse route to the originator, whose next hop is "the node from which the RREQ was received". That reverse route "will be needed if the node receives a RREP back to the node that originated the RREQ". If it cannot reply itself, and the TTL allows, it rebroadcasts the RREQ.

3. Who may reply. A node generates an RREP if either "(i) it is itself the destination", or (ii) it has an active route to the destination whose sequence number is valid and at least the one in the RREQ, and the destination-only flag is not set. Such a route is fresh enough: "a valid route entry for the destination whose associated sequence number is at least as great as that contained in the RREQ".

4. The reply comes back. The RREP is "unicast to the next hop toward the originator of the RREQ", along the reverse routes, its hop count incremented at each hop, so that at the originator it "represents the distance, in hops, of the destination from the originator". Each node that forwards it creates a forward route to the destination through the node it received the RREP from, and notes its precursors. When the RREP reaches the source, data can flow.

munotes.in221

AODV: Route Discovery and Route Maintenance

The destination's sequence number in the reply. A destination replying to an RREQ increments its own sequence number only "if the sequence number in the RREQ packet is equal to that incremented value"; otherwise it replies with its current number. That is how a new discovery after a break receives a number newer than the stale route.

Why sequence numbers prevent loops

"Managing the sequence number is crucial to avoiding routing loops." Each destination owns its number, and every node compares numbers before accepting route information: if the incoming number is older than the one stored, "the information related to that destination in the AODV message MUST be discarded, since that information is stale". A node replaces a route to a destination only with one carrying a newer sequence number, or the same number with fewer hops, unless its own route is already invalid. As in DSDV, stale information can never replace fresh information, so a node can never be persuaded to route through a neighbour whose route points back through it, which is how loops form in plain distance-vector routing. The comparison uses signed 32-bit arithmetic, so the numbers can wrap round past their maximum.

Route maintenance

Noticing a break. "Nodes monitor the link status of next hops in active routes." A node may learn of a broken link from its link layer, from failing to forward data, or from hello messages: a node can broadcast hellos every HELLO_INTERVAL (1,000 ms by default), and a node that hears nothing from a neighbour for more than ALLOWED_HELLO_LOSS × HELLO_INTERVAL = 2 × 1,000 = 2,000 ms assumes the link is lost.

Reporting it. A node starts route error processing in three cases: it detects a break on the next hop of an active route while sending data, it receives data for a destination it has no active route to, or it receives an RERR from a neighbour. For each unreachable destination it increments the destination sequence number (or copies it from the incoming RERR), marks the entry invalid, and keeps it for a while (DELETE_PERIOD) instead of deleting it at once. It sends the RERR to the neighbours in its precursor lists for those destinations: unicast if there is one, broadcast if there are several. The RERR travels back towards every source using the broken link, and each source that still needs the route discovers a new one.

Local repair. The node just upstream of a break "MAY choose to repair the link locally if the destination was no farther than MAX_REPAIR_TTL hops away", by broadcasting its own small RREQ for the destination instead of reporting the break. The source then often never notices.

munotes.in222

AODV: Route Discovery and Route Maintenance

AODV, traced

The network: S wants to send to D. S reaches D through A and C (S-A-C-D), and there is a longer way through A, E and F (S-A-E-F-D), and a spare node B linked to S and C.

# AODV route discovery, route error and rediscovery, traced on one network,
# following the processing rules of RFC 3561 (sections 6.3 to 6.11).
from collections import deque

links = {("S", "A"), ("S", "B"), ("A", "C"), ("B", "C"), ("C", "D"),
         ("A", "E"), ("E", "F"), ("F", "D")}
nodes = sorted({n for link in links for n in link})
seq = {n: 1 for n in nodes}               # each node's own sequence number
rreq_id = {n: 0 for n in nodes}
routes = {n: {} for n in nodes}           # dest: [next hop, hops, seq, valid]
precursors = {n: {} for n in nodes}       # dest: set of neighbours using us

def neighbours(n):
    return sorted(b if a == n else a for a, b in links if n in (a, b))

def discover(src, dst):
    seq[src] += 1                         # 6.1: increment before a discovery
    rreq_id[src] += 1
    known = routes[src].get(dst)
    want = known[2] if known else None    # last known destination seq, if any
    print("%s broadcasts RREQ %d for %s (own seq %d, dest seq %s)"
          % (src, rreq_id[src], dst, seq[src], want if want else "unknown"))
    seen, found, sent, dups = {src}, False, 1, 0
    queue = deque([(src, 0)])             # (node rebroadcasting, its hop count)
    while queue:
        sender, hops = queue.popleft()
        for n in neighbours(sender):
            if n in seen:                 # 6.5: same originator and RREQ ID
                dups += 1                 # a second copy: silently discarded
                continue
            seen.add(n)
            routes[n][src] = [sender, hops + 1, seq[src], True]   # reverse route
            if n == dst:
                if want is not None and want == seq[dst] + 1:
                    seq[dst] += 1         # 6.6.1
                print("  %s hears it from %s, %d hops: reverse route set; "
                      "%s will reply" % (n, sender, hops + 1, n))
                found = True
                continue                  # the destination does not rebroadcast
            print("  %s hears it from %s: reverse route to %s via %s; rebroadcasts"
                  % (n, sender, src, sender))
            queue.append((n, hops + 1))
            sent += 1
    print("  RREQ transmissions: %d; second copies discarded: %d" % (sent, dups))
    if found:
        reply(dst, src, seq[dst])

def reply(dst, src, dseq):
    node, hops = dst, 0
    while node != src:
        nxt = routes[node][src][0]        # follow the reverse route
        hops += 1
        routes[nxt][dst] = [node, hops, dseq, True]            # forward route
        precursors[node].setdefault(dst, set()).add(nxt)
        node = nxt
    print("  %s's RREP (dest seq %d) returns in %d hops: %s now routes to %s via %s"
          % (dst, dseq, hops, src, dst, routes[src][dst][0]))

def path(src, dst):
    p = [src]
    while p[-1] != dst:
        p.append(routes[p[-1]][dst][0])
    return " > ".join(p)

def link_breaks(a, b, dst):
    links.discard((a, b) if (a, b) in links else (b, a))
    print("link %s-%s breaks; %s notices while forwarding data to %s" % (a, b, a, dst))
    node = a
    while True:                           # 6.11: invalidate, bump seq, tell precursors
        r = routes[node][dst]
        r[2], r[3] = r[2] + 1 if node == a else r[2], False
        told = sorted(precursors[node].get(dst, ()))
        print("  %s marks %s invalid (dest seq %d)%s" % (node, dst, r[2],
              ", sends RERR to " + ", ".join(told) if told else ""))
        if not told:
            break
        for t in told:
            routes[t][dst][2] = r[2]      # copied from the RERR
        node = told[0]

discover("S", "D")
print("data path:", path("S", "D"))
link_breaks("C", "D", "D")
discover("S", "D")
print("data path:", path("S", "D"))
munotes.in223

AODV: Route Discovery and Route Maintenance

S broadcasts RREQ 1 for D (own seq 2, dest seq unknown)
  A hears it from S: reverse route to S via S; rebroadcasts
  B hears it from S: reverse route to S via S; rebroadcasts
  C hears it from A: reverse route to S via A; rebroadcasts
  E hears it from A: reverse route to S via A; rebroadcasts
  D hears it from C, 3 hops: reverse route set; D will reply
  F hears it from E: reverse route to S via E; rebroadcasts
  RREQ transmissions: 6; second copies discarded: 8
  D's RREP (dest seq 1) returns in 3 hops: S now routes to D via A
data path: S > A > C > D
link C-D breaks; C notices while forwarding data to D
  C marks D invalid (dest seq 2), sends RERR to A
  A marks D invalid (dest seq 2), sends RERR to S
  S marks D invalid (dest seq 2)
S broadcasts RREQ 2 for D (own seq 3, dest seq 2)
  A hears it from S: reverse route to S via S; rebroadcasts
  B hears it from S: reverse route to S via S; rebroadcasts
  C hears it from A: reverse route to S via A; rebroadcasts
  E hears it from A: reverse route to S via A; rebroadcasts
  F hears it from E: reverse route to S via E; rebroadcasts
  D hears it from F, 4 hops: reverse route set; D will reply
  RREQ transmissions: 6; second copies discarded: 7
  D's RREP (dest seq 2) returns in 4 hops: S now routes to D via A
data path: S > A > E > F > D
The seven-node network twice: the first route S, A, C, D with the reverse routes pointing back to S, and after the C-D link breaks, the repaired route S, A, E, F, D

Figure 36.1 The trace, drawn: reverse routes and the route found, before and after the break

Reading the first discovery.

  1. S increments its own sequence number to 2 and broadcasts RREQ 1. It knows nothing about D, so the destination sequence number is marked unknown (the U flag).
  2. A and B hear it directly and set reverse routes to S through S. C hears it first from A, so its reverse route to S goes through A; the copy that arrives later from B is one of the 8 second copies discarded. E also hears it from A.
  3. D hears it from C, three hops from S, and does not rebroadcast: it is the destination, so it replies. F still hears the RREQ from E and rebroadcasts it; nodes cannot know that the destination has already been found. In all, 6 nodes transmit the request (S, A, B, C, E and F).
  4. D's RREP travels D to C to A to S along the reverse routes, and each of them sets a forward route: C to D directly, A to D through C, S to D through A. The route is S > A > C > D, three hops, exactly the shortest path.
munotes.in224

AODV: Route Discovery and Route Maintenance

Reading the break and the rediscovery.

  1. The C-D link breaks, and C notices when data fails to get through. C marks its route to D invalid and increments the sequence number from 1 to 2. Its precursor for D is A, so it sends A an RERR; A marks its route invalid with the number from the RERR, and sends an RERR to its precursor, S.
  2. S, still needing D, discovers again: RREQ 2, with S's own number now 3, and the last destination number it knows, 2.
  3. This time D is reached only through E and F, four hops. Because the RREQ asks for number 2 and D's own number is 1, D increments to 2 before replying ("equal to that incremented value"). The RREP carries 2, at least as new as anything any node holds, so every node on the new path accepts it, and S now routes S > A > E > F > D.
  4. When RREQ 2 passed them, A and C held only invalid routes to D, numbered 2, so neither could answer it: an intermediate node may reply only from an active route. That is the rule that stops a stale route being handed back as an answer.

Worked example: the expanding ring search and its timeouts

Flooding every RREQ to the whole network is wasteful when the destination is near, so RFC 3561 says the originator "SHOULD use an expanding ring search". It starts with a small TTL, and each time the search times out, tries again further. With the RFC's default values (section 10):

ParameterDefault
NODE_TRAVERSAL_TIME40 ms
NET_DIAMETER35 hops
TTL_START1
TTL_INCREMENT2
TTL_THRESHOLD7
TIMEOUT_BUFFER2
RREQ_RETRIES2
munotes.in225

AODV: Route Discovery and Route Maintenance

  • The TTLs tried: 1, then 1 + 2 = 3, then 5, then 7, and beyond TTL_THRESHOLD the whole network, TTL = NET_DIAMETER = 35.
  • The timeout for each ring is RING_TRAVERSAL_TIME = 2 × NODE_TRAVERSAL_TIME × (TTL + TIMEOUT_BUFFER): for TTL 1, 2 × 40 × (1 + 2) = 240 ms; for TTL 3, 2 × 40 × 5 = 400 ms; for TTL 5, 2 × 40 × 7 = 560 ms; for TTL 7, 2 × 40 × 9 = 720 ms.
  • The whole-network timeout is NET_TRAVERSAL_TIME = 2 × NODE_TRAVERSAL_TIME × NET_DIAMETER = 2 × 40 × 35 = 2,800 ms.
  • Retries at full TTL use "a binary exponential backoff": after the first 2,800 ms wait, the next wait is 2 × 2,800 = 5,600 ms, and the one after that 2 × 5,600 = 11,200 ms, up to RREQ_RETRIES = 2 additional attempts.

So a source whose destination is unreachable waits 240 + 400 + 560 + 720 = 1,920 ms through the rings, then 2,800 + 5,600 + 11,200 = 19,600 ms at full TTL: 1,920 + 19,600 = 21,520 ms, about 21.5 seconds, before it drops the buffered packets and tells the application "Destination Unreachable". A destination two hops away, by contrast, answers within the second ring.

Distinctions

RREQRREPRERR
Sent byThe source (and rebroadcast)The destination, or a node with a fresh enough routeA node that detects a break or receives an RERR
HowBroadcastUnicast back along the reverse routeUnicast or broadcast to precursors
CreatesReverse routes to the sourceForward routes to the destinationInvalid entries, with bumped sequence numbers
Reverse routeForward route
Points towardsThe originator of the RREQThe destination
Set up byThe RREQ, at every node it reachesThe RREP, at every node it passes
Used forCarrying the RREP backCarrying the data
AODVDSDV
When routes are madeOn demandAll the time
Periodic trafficOptional hello messages onlyPeriodic table broadcasts
Loop freedomDestination sequence numbersDestination sequence numbers

What it does not mean

AODV does not keep a route to everyone. A node holds entries only for destinations in active communication, and for reverse routes created by requests it has seen.

The RREQ does not stop when the destination is found. Nodes that have not yet heard it keep rebroadcasting; only the duplicate rule limits the flood.

Any node does not answer from any route. An intermediate node replies only from an active route with a fresh enough sequence number, and not at all if the destination-only flag is set.

A route error is not a broadcast to everyone. It goes to the precursors, the neighbours actually using the broken route.

munotes.in226

AODV: Route Discovery and Route Maintenance

Quick revision

  • AODV (RFC 3561): reactive, loop-free, next-hop routing; silent while routes work.
  • Messages: RREQ (type 1; RREQ ID, destination and originator addresses and sequence numbers, hop count; flags J, R, G, D, U; 24 bytes), RREP (type 2; destination sequence number, hop count, lifetime), RERR (type 3; unreachable destinations with sequence numbers).
  • Discovery: broadcast RREQ; duplicates (same originator and RREQ ID) discarded; reverse route through the neighbour heard first; reply by the destination or a node with an active, fresh enough route; RREP unicast back; forward route at each hop.
  • Sequence numbers: owned by the destination; newer wins, then fewer hops; stale information discarded; loop freedom.
  • Maintenance: link breaks noticed by the link layer, failed forwarding or hello messages (every 1,000 ms; silence for 2,000 ms means lost); invalidate, increment the destination sequence number, RERR to precursors; source rediscovers; optional local repair.
  • Expanding ring: TTL 1, 3, 5, 7, then 35; ring timeouts 240, 400, 560, 720 ms; NET_TRAVERSAL_TIME 2,800 ms, then backoff 5,600 and 11,200 ms.
  • Our trace: first route S > A > C > D (3 hops, 6 RREQ transmissions); after the C-D break, S > A > E > F > D (4 hops), with D's number raised from 1 to 2.

Test yourself

1. Explain route discovery in AODV. A source that needs a route and has none increments its sequence number and RREQ ID and broadcasts a Route Request. Every node that receives it for the first time records a reverse route to the source through the neighbour it heard it from and, unless it can reply, rebroadcasts it; copies with the same originator and RREQ ID are discarded. The destination, or an intermediate node with an active route whose sequence number is at least the one requested, unicasts a Route Reply back along the reverse routes. Each node forwarding the reply sets a forward route to the destination, and when it reaches the source the route is ready.

2. How does AODV avoid routing loops? Every destination maintains its own sequence number, carried in RREQs, RREPs and RERRs. Nodes accept route information only if its sequence number is newer than what they hold, or equal with fewer hops, and discard stale information. When a link breaks, the sequence number of the lost destinations is incremented so that old routes cannot be revived. Because stale routes never replace fresh ones, a node can never route through a neighbour whose route leads back through itself.

3. Explain route maintenance in AODV. Nodes monitor the links to the next hops of active routes, through the link layer, failed transmissions or hello messages. When a link breaks, the upstream node invalidates the routes that used it, increments their destination sequence numbers, and sends a Route Error listing the unreachable destinations to the neighbours in its precursor lists; they do the same towards their precursors, until the sources are informed. A source that still needs the route starts a new route discovery; alternatively the node upstream of the break may try a local repair.

munotes.in227

AODV: Route Discovery and Route Maintenance

4. What is the difference between a reverse route and a forward route? A reverse route points towards the originator of a route request and is set up by every node the request reaches; it is used to carry the reply back. A forward route points towards the destination and is set up by every node the reply passes through; it carries the data.

5. With the default parameters, how long does the first ring of an expanding ring search wait for a reply? RING_TRAVERSAL_TIME = 2 × NODE_TRAVERSAL_TIME × (TTL + TIMEOUT_BUFFER) = 2 × 40 × (1 + 2) = 240 ms.

Contents This chapter on its own page

munotes.in228

Chapter Thirty-Seven

DSR, and What Its Routes Cost Against AODV

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Ad-hoc networks in WSNs" (and the paired practical, "Simulate the DSR protocol and compare routing overhead with AODV")

In one line

DSR is on-demand routing in which the source writes the whole route into every packet: discovery records the path as the request travels, the source and anyone who overhears keep routes in a cache, and the price of needing fewer route discoveries is a longer header on every data packet.

In the wording a student can write in an examination: DSR (Dynamic Source Routing), specified in RFC 4728, is a reactive routing protocol that uses source routing: each data packet carries in its header the complete, ordered list of nodes it must pass through. It has two mechanisms. Route Discovery: a source with no route broadcasts a Route Request containing a request identifier and a route record; each node appends its own address and rebroadcasts, discarding requests it has already seen or that already list it; the target returns a Route Reply carrying the recorded route. Route Maintenance: each node forwarding a packet confirms that the next link works (by a link-layer, passive or DSR acknowledgement); if it does not, the node returns a Route Error to the source, which uses another cached route or discovers a new one. Every node keeps a route cache of routes it has learnt, including routes overheard in other packets, and may reply from its cache. DSR has no periodic messages, can keep several routes per destination, and is loop-free because the route is written in the packet; its cost is the source route in every packet header.

Source routing

"The basic version of DSR uses explicit 'source routing', in which each data packet sent carries in its header the complete, ordered list of nodes through which the packet will pass." RFC 4728 names three benefits: the sender chooses and controls its routes, it can use several routes to one destination (to share the load), and "a simple guarantee that the routes used are loop-free". And a fourth follows from the headers themselves: "other nodes forwarding or overhearing any of these packets can also easily cache this routing information for future use".

DSR also has nothing periodic: no routing advertisements, no link sensing, no neighbour detection packets. When nobody is sending, DSR sends nothing.

Route discovery

The request. A source (the initiator) with a packet for a destination (the target) and no route in its cache broadcasts a Route Request. It carries the target's address, a request identification chosen by the initiator, and a route record: "a record listing the address of each intermediate node through which this particular copy of the Route Request has been forwarded".

Each node that hears it, if it is not the target:

  1. discards it if it has "recently seen another Route Request message from this initiator bearing this same request identification and target address", or if its own address is already in the route record;
  2. otherwise "appends its own address to the route record in the Route Request and propagates it by transmitting it as a local broadcast packet (with the same request identification)".
munotes.in229

DSR, and What Its Routes Cost Against AODV

The target checks nothing of the kind first: if the request is for it, it "SHOULD return a Route Reply to the initiator of this Route Request", for every copy that reaches it, each carrying the route that copy recorded. So one discovery can give the initiator several routes.

A DSR Route Request travelling from S through A and C to D, its route record growing at each hop, and the Route Reply returning over the reversed record

Figure 37.1 One copy of a Route Request, its record growing hop by hop, and the reply

The reply's own route. The target needs a route back to the initiator. It may use a route from its cache; it may start its own discovery, piggybacking the reply on it; or it may simply reverse the recorded route. On a MAC such as IEEE 802.11 that needs links to work both ways, the route "MUST be reversed in this way", which also tests that the route is bidirectional before the initiator uses it.

While waiting, the initiator keeps the packet in a send buffer, and limits repeated discoveries for the same target with an exponential back-off, "doubling the timeout between each successive discovery", so a partitioned network is not flooded with hopeless requests.

The route cache

The route cache is what makes DSR different in practice.

Caching what you hear. "A node forwarding or otherwise overhearing any packet SHOULD add all usable routing information from that packet to its own Route Cache." Which directions may be cached depends on the links: where links may be one-way, only the forward direction of a recorded route is cached; where the MAC needs two-way links anyway (as 802.11 does), both directions are.

Replying from the cache. A node that receives a Route Request and has a cached route to the target "generally returns a Route Reply to the initiator itself rather than forward the Route Request", joining the recorded route to its cached one. It must first check the joined route has no node twice; if it would, it must not reply and forwards the request normally. The RFC's reason is instructive: a node should only return routes that pass through itself, so that a future Route Error on that route will reach it and clean its cache.

Limiting the flood. A node may first send a non-propagating Route Request with a hop limit of 1, which only asks its neighbours: "an inexpensive method for determining if the target is currently a neighbor of the initiator or if a neighbor node has a route to the target cached". If that fails, it sends a propagating request. An expanding ring search is also allowed, doubling the hop limit each time.

munotes.in230

DSR, and What Its Routes Cost Against AODV

Route maintenance

Every hop confirms its own link. "Each node transmitting the packet is responsible for confirming that data can flow over the link from that node to the next hop." The confirmation can be a link-layer acknowledgement (as 802.11 provides), a passive acknowledgement (the sender overhears the next node forwarding the packet), or a DSR acknowledgement requested explicitly.

A broken link. After the maximum number of retransmissions without confirmation, the node treats the link as broken. It "SHOULD remove this link from its Route Cache and SHOULD return a 'Route Error' to each node that has sent a packet routed over that link". The source removes the link from its cache and, if it has another route, it can "send the packet using the new route immediately"; otherwise it starts a new discovery.

Salvaging. A node that cannot forward a packet because its next link broke, but has another route to the destination in its own cache, "SHOULD 'salvage' the packet rather than discard it", replacing the source route with its own. A 4-bit salvage count in the header stops a packet being salvaged endlessly.

Route shortening. A node that overhears a packet it is not the next hop for, but whose route names it further on, knows the nodes between are unnecessary. In the RFC's example, D overhears B sending to C on the route A-B-C-D-E, and returns a gratuitous Route Reply to A giving the shorter route A-B-D-E.

The headers

Every DSR packet carries a DSR Options header with a 4-byte fixed part, followed by options:

  • Source Route option (type 96): 4 bytes of fields, including Segments Left ("number of explicitly listed intermediate nodes still to be visited") and the Salvage count, then one 4-byte IPv4 address per intermediate node. Its length field must be set to 4n + 2, n being the number of addresses.
  • Route Request option (type 1): type, length, a 16-bit identification, the target address, and the route record so far.
  • Route Reply option: the route being returned.
  • Route Error, Acknowledgement Request and Acknowledgement options.

Worked example: what the source route costs. A route of h hops has h - 1 intermediate nodes, so a data packet carries 4 + 4 + 4 × (h - 1) bytes of DSR header.

  • For 3 hops: 4 + 4 + 4 × 2 = 16 bytes. For 10 hops: 4 + 4 + 4 × 9 = 44 bytes.
  • Against a 512-byte data packet, 44 bytes is 44 / 512, about 8.6 per cent. Against a 64-byte packet it is 44 / 64, about 69 per cent.
munotes.in231

DSR, and What Its Routes Cost Against AODV

That difference explains why studies disagreed: Das, Perkins and Royer, with 512-byte packets, "didn't find source routing overheads to be a very significant performance issue", while an earlier study using 64-byte packets blamed DSR's poorer results on exactly those overheads.

DSR against AODV, counted

The practical asks to "compare routing overhead with AODV". The program runs both on the network of [AODV: Route Discovery and Route Maintenance], through the same events: S sends 100 packets to D, and after the 50th the link C-D breaks. The packet in flight at the break is lost in both, so 99 are delivered.

# DSR and AODV on the same network and the same events, with every routing
# transmission counted. S sends 100 packets to D; the link C-D breaks after 50.
from collections import deque

LINKS = {("S", "A"), ("S", "B"), ("A", "C"), ("B", "C"), ("C", "D"),
         ("A", "E"), ("E", "F"), ("F", "D")}

def neighbours(links, n):
    return sorted(b if a == n else a for a, b in links if n in (a, b))

def flood(links, src, dst):
    """One route request flood. Returns the request transmissions and the
    route recorded by every copy that reaches dst, in order of arrival."""
    seen, queue, sent, arrived = {src}, deque([[src]]), 1, []
    while queue:
        record = queue.popleft()
        for n in neighbours(links, record[-1]):
            if n == dst:
                arrived.append(record + [dst])
            elif n not in seen:
                seen.add(n)
                queue.append(record + [n])
                sent += 1                 # n rebroadcasts once
    return sent, arrived

def header_bytes(route):                  # RFC 4728: 4-byte DSR Options header,
    return 4 + 4 + 4 * (len(route) - 2)   # source route option 4 + 4 per relay

def name(route):
    return "-".join(route)

# DSR: the target replies to every copy; the source caches every route.
rreq, arrived = flood(LINKS, "S", "D")
rrep = sum(len(r) - 1 for r in arrived)
cache = sorted(arrived, key=len)
print("DSR : %d RREQ transmissions; D replies to %d copies (%s), %d RREP hops"
      % (rreq, len(arrived), ", ".join(name(r) for r in arrived), rrep))
first = cache[0]
data_bytes = 50 * header_bytes(first) * (len(first) - 1)
rerr = first.index("C")                   # Route Error from C back to S
cache = [r for r in cache if "C-D" not in name(r)]
second = cache[0]
data_bytes += 49 * header_bytes(second) * (len(second) - 1)
dsr = rreq + rrep + rerr
print("      C-D breaks: Route Error, %d hops; S switches to its cached %s"
      % (rerr, name(second)))
print("      routing transmissions %d; source-route bytes sent in data %d" % (dsr, data_bytes))

# AODV: one reply, from the destination's first copy; rediscovery after the break.
rreq1, arrived = flood(LINKS, "S", "D")
rrep1 = len(arrived[0]) - 1
rerr = arrived[0].index("C")              # RERR along the precursors, C to S
rreq2, arrived2 = flood(LINKS - {("C", "D")}, "S", "D")
rrep2 = len(arrived2[0]) - 1
aodv = rreq1 + rrep1 + rerr + rreq2 + rrep2
print("AODV: %d RREQ + %d RREP; after the break %d RERR + %d RREQ + %d RREP (%s)"
      % (rreq1, rrep1, rerr, rreq2, rrep2, name(arrived2[0])))
print("      routing transmissions %d; extra bytes in data 0" % aodv)
print("routing load per delivered packet: DSR %.3f, AODV %.3f" % (dsr / 99, aodv / 99))
munotes.in232

DSR, and What Its Routes Cost Against AODV

DSR : 6 RREQ transmissions; D replies to 2 copies (S-A-C-D, S-A-E-F-D), 7 RREP hops
      C-D breaks: Route Error, 2 hops; S switches to its cached S-A-E-F-D
      routing transmissions 15; source-route bytes sent in data 6320
AODV: 6 RREQ + 3 RREP; after the break 2 RERR + 6 RREQ + 4 RREP (S-A-E-F-D)
      routing transmissions 21; extra bytes in data 0
routing load per delivered packet: DSR 0.152, AODV 0.212

Reading it.

  1. The discovery costs the same. Both floods take 6 request transmissions: every node except D rebroadcasts once.
  2. DSR gets two routes for one flood. D answers both copies that reach it, S-A-C-D and S-A-E-F-D, so DSR spends 3 + 4 = 7 reply transmissions to AODV's 3. Das, Perkins and Royer saw the same pattern at scale: DSR's routing load was "dominated by RREP packets, primarily due to multiple replies from destination".
  3. The break is where DSR saves. Both send a 2-hop route error back to S. AODV must then flood again (6 more requests and a 4-hop reply) before S can send, while DSR switches at once to the route already in its cache. Hence 15 routing transmissions for DSR against 21 for AODV; per delivered packet, 15 / 99, about 0.152, against 21 / 99, about 0.212.
  4. The price is in every data packet. DSR's 99 delivered packets carried 6,320 bytes of source routes across their hops: 50 packets × 16 bytes × 3 hops = 2,400, plus 49 × 20 × 4 = 3,920. AODV's data packets carried none.

What the large comparison found

Das, Perkins and Royer compared the two protocols in detailed simulations with 50 and 100 nodes, measuring the packet delivery fraction, the average end-to-end delay, and the normalized routing load, "the number of routing packets 'transmitted' per data packet 'delivered' at the destination", each hop-wise transmission counted once (the measure our last line computes).

  • Routing load. "DSR almost always has a lower routing load than AODV. The difference is often significant (by a factor of up to 5), if routing load is presented in terms of packet counts. Presenting routing loads in terms of bytes is, however, less impressive (at most about a factor of 2)."
  • Where the savings come from. AODV's routing load "was dominated by RREQ packets (often as much as 90% of all routing packets)", while "all the routing load savings for DSR came from a large saving on RREQs", thanks to its cache.
  • Delivery and delay. "DSR outperforms AODV in less 'stressful' situations, i.e., smaller number of nodes and lower load and/or mobility. AODV, however, outperforms DSR in more stressful situations, with widening performance gaps with increasing stress."
  • Why. DSR's cache is also its weakness: it has no "mechanism to expire stale routes or to determine the freshness of routes when multiple choices are available", so under heavy load and mobility it keeps trying routes that no longer work. AODV's sequence numbers exist to tell fresh routes from stale ones.
munotes.in233

DSR, and What Its Routes Cost Against AODV

Distinctions

DSRAODV
Where the route livesIn every packet's header (source route) and in route cachesIn each node's routing table, as a next hop
Routes per destinationSeveral, in the cacheOne
Discovery replyThe target replies to every request copyThe destination (or a fresh-enough node) replies once
Learning from overheard packetsYes, routes are cached from headersNo
Loop freedomThe route is written in the packetDestination sequence numbers
Periodic messagesNoneOptional hellos
Stale route handlingNo freshness measure in the basic cacheSequence numbers show freshness
Header cost per data packet8 bytes plus 4 per relayNone
Routing load (Das et al.)Lower: up to 5 times in packets, at most about 2 in bytesHigher, mostly route requests

What it does not mean

Source routing does not mean the source knows the whole network. It knows the routes it has discovered or overheard, nothing more.

More routes in the cache are not all good routes. A cached route may be stale; DSR learns this only when a Route Error comes back.

Lower routing load does not mean better delivery. DSR sent fewer routing packets in Das, Perkins and Royer's simulations yet delivered fewer data packets than AODV when the network was stressed.

DSR's header cost is not fixed. It grows by 4 bytes with every relay, and matters most when data packets are small and routes long.

Quick revision

  • DSR (RFC 4728): reactive source routing; the route is in every packet; no periodic messages; loop-free by construction.
  • Route Discovery: Route Request with request ID and route record; nodes append themselves and rebroadcast, discarding repeats or requests that already list them; the target replies to every copy; the reply returns by a cached route, a piggybacked discovery, or the reversed record (required over 802.11).
  • Route cache: routes from discoveries and overheard packets; replies from the cache (no repeated node allowed); non-propagating requests (hop limit 1) first; exponential back-off between discoveries.
  • Route Maintenance: each hop confirms its link (link-layer, passive or DSR acknowledgement); on failure, remove the link and send a Route Error to the senders; salvaging with a 4-bit count; route shortening by gratuitous replies.
  • Header: 4-byte DSR Options header; Source Route option 4 + 4n bytes (n relays): 16 bytes for 3 hops, 44 for 10.
  • Our count: DSR 15 routing transmissions, AODV 21; DSR carried 6,320 bytes of source routes; routing load per delivered packet 0.152 against 0.212.
  • Das, Perkins and Royer: DSR's routing load lower by up to 5 times in packets, at most 2 in bytes; DSR better when lightly stressed, AODV better under high load and mobility; DSR hurt by stale cached routes.
munotes.in234

DSR, and What Its Routes Cost Against AODV

Test yourself

1. Explain route discovery in DSR. A source with no cached route broadcasts a Route Request carrying the target's address, a request identifier and a route record. Each node that receives it discards it if it has already seen that request or is already in the record; otherwise it appends its address to the record and rebroadcasts. The target sends a Route Reply for each copy it receives, carrying the recorded route, back along the reversed route, a cached route or a piggybacked discovery. The source caches the routes and puts one in the header of each data packet.

2. What is the route cache in DSR, and how is it used? Each node keeps the routes it has learnt, from its own discoveries and from routes it forwards or overhears in other packets' headers. The source uses it to send without a new discovery; an intermediate node can answer a Route Request from it instead of forwarding the request; a node whose next link breaks can salvage a packet with another cached route. It cuts route discoveries sharply, but cached routes can become stale.

3. Explain route maintenance in DSR. Each node forwarding a packet is responsible for confirming that the link to the next hop works, through a link-layer acknowledgement, a passive acknowledgement (overhearing the next node forward it) or an explicit DSR acknowledgement. If the link fails after the allowed retransmissions, the node removes it from its cache and sends a Route Error to the sources that used it; the source then uses another cached route or starts a new discovery. The node may also salvage the packet with a route from its own cache.

4. Compare DSR and AODV. Both are reactive. DSR carries the whole route in each packet, caches several routes per destination (including overheard ones), lets the target reply to every request copy, and uses no periodic messages; AODV keeps one next-hop route per destination in each node's table and uses destination sequence numbers to keep routes fresh and loop-free. DSR generates fewer routing packets, mostly because it needs fewer route discoveries, but carries a header that grows with route length; AODV delivers better under high load and mobility because DSR's caches fill with stale routes.

munotes.in235

DSR, and What Its Routes Cost Against AODV

5. How many bytes of DSR header does a data packet carry on a 5-hop route? Four relays: 4 + 4 + 4 × 4 = 24 bytes.

Contents This chapter on its own page

munotes.in236

Chapter Thirty-Eight

Measuring a MANET Protocol: Throughput, Delivery Ratio and Delay

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Ad-hoc networks in WSNs" (and the paired practical, "Measure throughput, packet delivery ratio, and end-to-end delay for different MANET routing protocols")

In one line

A routing protocol is judged by what share of the packets arrive (delivery ratio), how long they take (end-to-end delay), how much data gets through per second (throughput), and how many routing packets it spends to achieve that (routing overhead), each measured over many identical scenarios, and none of them enough on its own.

In the wording a student can write in an examination: the main performance measures of a MANET routing protocol are: (1) packet delivery ratio (PDR), the number of data packets received by the destinations divided by the number sent by the sources; (2) average end-to-end delay, the mean time from a data packet's sending to its reception, including route discovery, queuing, retransmission and propagation delays; (3) throughput, the amount of data delivered per unit time, in bits per second; (4) routing overhead, the total number of routing packets transmitted, each hop counted as one transmission; and (5) normalized routing load, routing transmissions per data packet delivered. RFC 2501 adds route acquisition time, out-of-order delivery and efficiency ratios. For a fair comparison, every protocol must be run on the same scenarios (the same movement and the same traffic), across a range of contexts (network size, connectivity, rate of topology change, mobility, traffic load), with enough runs to average.

The measures, defined

Packet delivery ratio. Broch and colleagues define it as "The ratio between the number of packets" originated by the application's constant bit rate (CBR) sources and "the number of packets received by the CBR sink at the final destination". Das, Perkins and Royer call it the packet delivery fraction. As a formula, PDR = packets received / packets sent. It "describes the loss rate that will be seen by the transport protocols", and characterises "both the completeness and correctness of the routing protocol".

Average end-to-end delay. The mean, over delivered packets, of the time from sending to reception. Das, Perkins and Royer spell out what it includes: "all possible delays caused by buffering during route discovery latency, queuing at the interface queue, retransmission delays at the MAC, propagation and transfer times".

Throughput. The data delivered per unit of time: bits of data received, divided by the time over which they were received. RFC 2501 lists "End-to-end data throughput and delay" first among its quantitative metrics, and asks for "Statistical measures of data routing performance (e.g., means, variances, distributions)", not a single number.

Routing overhead. "The total number of routing packets transmitted during the simulation. For packets sent over multiple hops, each transmission of the packet (each hop) counts as one transmission." It matters because it "measures the scalability of a protocol, the degree to which it will function in congested or low-bandwidth environments, and its efficiency in terms of consuming node battery power".

munotes.in237

Measuring a MANET Protocol: Throughput, Delivery Ratio and Delay

Normalized routing load. Das, Perkins and Royer's version: "the number of routing packets 'transmitted' per data packet 'delivered' at the destination". Dividing by what was delivered makes protocols comparable across loads.

Path optimality (Broch and colleagues): the number of hops a packet actually took, minus the shortest path that existed when it was sent.

RFC 2501's further measures. Route acquisition time, "the time required to establish route(s) when requested", "of particular concern with 'on demand' routing algorithms"; the percentage of out-of-order delivery, which matters to TCP; and efficiency, the internal cost of the protocol: data bits transmitted per data bit delivered, control bits transmitted per data bit delivered, and packets transmitted per data packet delivered. On control bits it is strict: "anything that is not data is control overhead", including the headers of data packets.

The measures, computed from traces

Each line of a trace is a time in seconds, an event, and for data packets their number: send (the source hands a data packet to the network), recv (the destination receives it), and rreq, rrep or rerr (one hop-wise transmission of a routing packet). The two traces replay the scenario of the previous chapter, shortened to six data packets of 64 bytes, one every 0.25 s, with the link C-D breaking as packet 4 is forwarded.

The AODV run, aodv.trace:

0.000 send 1
0.000 rreq
0.003 rreq
0.004 rreq
0.007 rreq
0.008 rreq
0.011 rreq
0.012 rrep
0.016 rrep
0.020 rrep
0.041 recv 1
0.250 send 2
0.262 recv 2
0.500 send 3
0.512 recv 3
0.750 send 4
0.755 rerr
0.759 rerr
0.760 rreq
0.763 rreq
0.764 rreq
0.767 rreq
0.768 rreq
0.771 rreq
0.772 rrep
0.776 rrep
0.780 rrep
0.784 rrep
1.000 send 5
1.016 recv 5
1.250 send 6
1.266 recv 6

The DSR run, dsr.trace, identical for the data but with its own routing events: seven reply transmissions (two replies from D), and no second discovery:

0.000 send 1
0.000 rreq
0.003 rreq
0.004 rreq
0.007 rreq
0.008 rreq
0.011 rreq
0.012 rrep
0.013 rrep
0.016 rrep
0.017 rrep
0.020 rrep
0.021 rrep
0.024 rrep
0.041 recv 1
0.250 send 2
0.262 recv 2
0.500 send 3
0.512 recv 3
0.750 send 4
0.755 rerr
0.759 rerr
1.000 send 5
1.016 recv 5
1.250 send 6
1.266 recv 6

The program reads both and prints every measure:

# The measures a MANET comparison reports, computed from packet traces.
# A trace line is: time, event, and for data packets their number.
import statistics

BITS = 64 * 8                             # every data packet carries 64 bytes

def measure(name):
    sent, got, routing = {}, {}, 0
    for line in open(name):
        t, event, *pkt = line.split()
        if event == "send":
            sent[pkt[0]] = float(t)
        elif event == "recv":
            got[pkt[0]] = float(t)
        else:
            routing += 1                  # one hop-wise routing transmission
    delays = [1000 * (got[p] - sent[p]) for p in got]     # milliseconds
    span = max(got.values()) - min(sent.values())
    return (len(sent), len(got), len(got) / len(sent), statistics.mean(delays),
            statistics.median(delays), max(delays), len(got) * BITS / span,
            routing, routing / len(got))

print("delays in ms  sent  got  delivery   mean  median  worst   throughput  routing  load")
for proto in ("aodv", "dsr"):
    m = measure(proto + ".trace")
    print("%-12s %5d %4d %9.3f %6.1f %7.1f %6.1f %8.0f b/s %8d %5.2f"
          % ((proto.upper(),) + m))
munotes.in238

Measuring a MANET Protocol: Throughput, Delivery Ratio and Delay

delays in ms  sent  got  delivery   mean  median  worst   throughput  routing  load
AODV             6    5     0.833   19.4    16.0   41.0     2022 b/s       21  4.20
DSR              6    5     0.833   19.4    16.0   41.0     2022 b/s       15  3.00

Worked example: the numbers by hand

  • Delivery ratio: 5 of 6 packets arrived (packet 4 was lost at the break): 5 / 6, about 0.833.
  • Delays: 41, 12, 12, 16 and 16 ms. The mean is (41 + 12 + 12 + 16 + 16) / 5 = 97 / 5 = 19.4 ms; the median, the middle value in order, is 16 ms; the worst is 41 ms, packet 1, which waited for the first route discovery.
  • Throughput: 5 packets of 64 bytes are 5 × 64 × 8 = 2,560 bits, received between 0.000 s (the first send) and 1.266 s (the last reception): 2,560 / 1.266, about 2,022 bit/s.
  • Routing load: AODV 21 routing transmissions for 5 delivered packets, 21 / 5 = 4.2; DSR 15 / 5 = 3.0.

Control bits, as RFC 2501 counts them. The routing packets are not DSR's only overhead: its data packets carry source routes. Packets 1 to 3 crossed 3 hops with a 16-byte DSR header each, 3 × 16 × 3 = 144 bytes; packet 4 crossed 2 hops before it was lost, 16 × 2 = 32 bytes; packets 5 and 6 crossed 4 hops with 20-byte headers, 2 × 20 × 4 = 160 bytes. That is 144 + 32 + 160 = 336 bytes of header, against 5 × 64 = 320 bytes of data delivered. With packets this small, the count of routing packets favours DSR while the count of control bytes does not, the effect Das, Perkins and Royer found in the byte counts of their larger comparison.

The delay of each of the six data packets in the AODV trace, with packet 4 lost and the mean and median drawn across

Figure 38.1 The AODV trace's delays: one slow packet lifts the mean above four of the five

What each measure hides

  1. Delivery ratio hides time. Both protocols delivered 5 of 6. A packet that arrives a minute late counts the same as one that arrives at once.
  2. Average delay hides its spread, and its survivors. The mean of 19.4 ms is above four of the five delays, pulled up by one route discovery; the median (16 ms) and the worst (41 ms) tell more. And the lost packet is in no delay at all: a protocol that drops its slowest packets looks faster.
  3. Throughput hides the offered load. With 4 packets a second offered, no protocol can deliver more than 4 × 512 = 2,048 bit/s, so at light load throughput mostly measures the traffic generator. It separates protocols only when the network is loaded enough to lose.
  4. Routing overhead in packets hides bytes, and the MAC. DSR's headers count as data; long source routes do not show. And Das, Perkins and Royer found that route replies and errors cost the MAC more per packet than route requests, so "when the MAC overhead was factored in, DSR was found to generate about as much overall network load as AODV".
  5. Any one run hides chance. A different movement pattern would break a different link at a different moment. One trace proves nothing about a protocol; it shows how the measures are computed.
munotes.in239

Measuring a MANET Protocol: Throughput, Delivery Ratio and Delay

In these two traces, delivery, delay and throughput are identical, because the traces were built that way, and only the routing load separates the protocols. In a real comparison, all of them differ, and they can disagree.

Setting up a fair comparison

Broch and colleagues' comparison of DSDV, TORA, DSR and AODV is the standard model of method, and every choice in it has a reason:

  1. The same scenarios for every protocol. "Each run of the simulator accepts as input a scenario file that describes the exact motion of each node and the exact sequence of packets originated by each node". They "pre-generated 210 different scenario files" and ran all four protocols on each, so that "each protocol was challenged in an identical fashion".
  2. Many movement patterns. 50 nodes in a 1,500 m by 300 m area, moving by the random waypoint model for 900 simulated seconds, with pause times of 0, 30, 60, 120, 300, 600 and 900 seconds (0 is continuous motion, 900 is none), and "10 for each value of pause time": 7 × 10 = 70 movement patterns, at maximum speeds of 20 m/s and 1 m/s.
  3. Controlled traffic. Constant bit rate sources at 4 packets a second, with 10, 20 or 30 sources; 64-byte packets, chosen "to factor out congestive effects", because with 1,024-byte packets congestion swamped every protocol.
  4. A long rectangle. The 1,500 m by 300 m shape forces "the use of longer routes between nodes than would occur in a square space with equal node density", so that routing is really exercised.
  5. Discard the start. Camp, Boleng and Davies, who later studied the random waypoint model, recommend discarding "the initial 1000 s of simulation time", before the nodes' distribution settles.
munotes.in240

Measuring a MANET Protocol: Throughput, Delivery Ratio and Delay

RFC 2501 lists the contexts a comparison should vary: "Network size", "Network connectivity" (the average number of neighbours), "Topological rate of change", "Link capacity", "Fraction of unidirectional links", "Traffic patterns", "Mobility" and "Fraction and frequency of sleeping nodes". A result from one context does not carry over to another.

Distinctions

MeasureWhat it asksFormulaBetter is
Packet delivery ratioHow much arrives?Received / sentHigher
Average end-to-end delayHow fast?Mean of (receive time - send time), delivered packets onlyLower
ThroughputHow much per second?Bits received / timeHigher
Routing overheadHow much does routing cost?Routing transmissions, per hopLower
Normalized routing loadCost per delivered packetRouting transmissions / packets receivedLower
Path optimalityAre routes short?Hops taken - shortest hopsLower
External measures (RFC 2501)Internal measures (RFC 2501)
Seen fromThe applications using the networkThe protocol's own costs
ExamplesThroughput, delay, route acquisition time, out-of-order deliveryData and control bits per delivered bit; packets per delivered packet

What it does not mean

A higher delivery ratio does not make a protocol better everywhere. Das, Perkins and Royer found DSR ahead at light stress and AODV ahead at heavy stress: the ranking depends on the context.

A low routing overhead is not free. A protocol can save routing packets by caching routes that turn out stale, and pay in lost data.

Throughput is not the radio's rate. It is what reached the destinations, after losses and overheads, and at light load it simply equals what was offered.

Averages over one scenario are not results. They are one sample; the published comparisons average over tens of scenarios per setting.

Quick revision

  • PDR = data packets received / sent (Broch: "completeness and correctness"; Das: packet delivery fraction).
  • Average end-to-end delay: mean over delivered packets; includes route discovery, queuing, MAC retransmission, propagation.
  • Throughput: data bits delivered / time.
  • Routing overhead: routing transmissions, each hop counts once; normalized routing load = routing transmissions / data packets delivered.
  • Path optimality: hops taken - shortest hops.
  • RFC 2501 adds route acquisition time, out-of-order delivery, efficiency (bits and packets per delivered bit or packet; "anything that is not data is control overhead").
  • Our traces: PDR 0.833, mean delay 19.4 ms (median 16, worst 41), throughput about 2,022 bit/s, routing load 4.20 (AODV) against 3.00 (DSR); DSR's headers 336 bytes against 320 bytes of data.
  • Fair comparison (Broch): identical scenario files, 210 of them, 50 nodes, 1,500 m by 300 m, 900 s, 7 pause times × 10 patterns, 20 and 1 m/s, CBR 4 packets/s, 10 to 30 sources, 64-byte packets; discard the warm-up.
munotes.in241

Measuring a MANET Protocol: Throughput, Delivery Ratio and Delay

Test yourself

1. Define packet delivery ratio, end-to-end delay and throughput. Packet delivery ratio is the number of data packets received by the destinations divided by the number originated by the sources. Average end-to-end delay is the mean time between a data packet's origination and its reception, over the packets delivered, and includes route discovery, queuing, MAC retransmission and propagation delays. Throughput is the amount of data delivered per unit time, the bits received divided by the time over which they were received.

2. What is normalized routing load, and why normalize? The number of routing packet transmissions, each hop counted once, divided by the number of data packets delivered. Dividing by delivered packets makes the cost comparable between protocols and between loads: it says how much routing effort each successful delivery needed.

3. A source sends 200 packets; 180 arrive. The routing protocol makes 540 routing transmissions. Find the PDR and the normalized routing load. PDR = 180 / 200 = 0.9; normalized routing load = 540 / 180 = 3.

4. How should two MANET routing protocols be compared fairly? Run both on the same scenario files, the same node movements and the same traffic, generated in advance; use many scenarios for each setting and average the results; vary the contexts that matter (pause time or speed, number of sources, network size); discard the start of each run while the mobility model settles; and report several measures together, delivery ratio, delay, throughput and routing overhead, since each hides something the others show.

5. Why can an average delay make a protocol look better than it is? It is computed over delivered packets only, so a protocol that loses its slowest packets has a lower average; and a mean is pulled by a few long waits, so the median and the worst case should be reported beside it.

Contents This chapter on its own page

munotes.in242

Chapter Thirty-Nine

Energy Efficiency in Ad Hoc Networks: Where the Energy Goes

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Energy efficiency considerations in ad-hoc networks"

In one line

In an ad hoc network most energy is spent not sending but listening, with the radio on and nothing arriving, so saving energy means switching radios off: lowering transmit power helps little on a sensor radio, while letting redundant nodes sleep, scheduling sleep in the MAC, routing around weak batteries and sending less data help a great deal.

In the wording a student can write in an examination: energy is the scarcest resource of an ad hoc network, and RFC 2501 lists energy-constrained operation as a defining characteristic. Most of a node's energy goes to its radio, which has four states: transmit, receive, idle (listening) and sleep; idle listening costs nearly as much as receiving, and on sensor radios receiving costs as much as transmitting. The main sources of energy waste are collisions, overhearing (receiving packets meant for others), control packet overhead, and above all idle listening. Energy is saved by: (1) transmission power control, sending with just enough power; (2) topology control, keeping only enough nodes awake to preserve connectivity and letting redundant ones sleep (GAF, Span); (3) sleep scheduling or duty cycling in the MAC (S-MAC, B-MAC); (4) energy-aware routing, which chooses routes by energy rather than hop count; and (5) reducing traffic by in-network processing and aggregation.

Why energy shapes everything

RFC 2501 puts it among the four characteristics of a MANET: some or all nodes "may rely on batteries or other exhaustible means for their energy. For these nodes, the most important system design criteria for optimization may be energy conservation." It also asks every routing protocol to support "sleep" period operation: nodes "may stop transmitting and/or receiving (even receiving requires power) for arbitrary time periods", and the protocol should cope "without overly adverse consequences". In a sensor network, whose batteries are rarely replaced, both apply with full force.

Where the energy goes

The radio's states

A radio is in one of four states: transmitting, receiving, idle (on and listening, with nothing arriving), and asleep. The surprise, in every measurement, is how little idle saves over receiving. The S-MAC paper collects the numbers: "Many measurements have shown" that idle listening costs half to all of the energy needed to receive; Stemm and Katz measured idle, receive and send power in the ratios 1 : 1.05 : 1.4, and a 2 Mbit/s 802.11 module's specification gives 1 : 2 : 2.5. The GAF paper cites a third measurement, 1 : 1.2 : 1.7.

On a sensor radio the ratio tilts further: the CC2420 draws 18.8 mA to receive and 17.4 mA to transmit at full power, so receiving costs more than transmitting, and listening is receiving ([The Radio, the Sensors and the Power Supply of a Node]). A node whose radio is on and waiting spends almost as much as one that is busy.

munotes.in243

Energy Efficiency in Ad Hoc Networks: Where the Energy Goes

The four sources of waste

The S-MAC paper names the "major sources of energy waste":

  1. Collision: "When a transmitted packet is corrupted it has to be discarded, and the follow-on retransmissions increase energy consumption."
  2. Overhearing: "a node picks up packets that are destined to other nodes".
  3. Control packet overhead: "Sending and receiving control packets consumes energy too, and less useful data packets can be transmitted."
  4. Idle listening: "listening to receive possible traffic that is not sent". It is "especially true in many sensor network applications. If nothing is sensed, nodes are in idle mode for most of the time."

How the MAC layer attacks each of these is [MAC Protocols for Sensor Networks: The Job and Where the Energy Goes].

Idle listening swamps the routing protocol

The GAF paper's first figure makes the point with four routing protocols (AODV, DSR, DSDV and TORA) simulated with 50 nodes in a 1,500 m by 300 m area under the random waypoint model. Counting only the energy to send and receive packets, the on-demand protocols (AODV and DSR) used far less than DSDV, which spends energy building routes nobody uses. Counting idle listening as well, all four consumed roughly the same energy, within a few per cent: idle time dominated. The paper also found overhearing to be the largest extra cost of the on-demand protocols.

The lesson for design. Making packets cheaper, or fewer, barely changes the total while every radio listens all the time. The energy is saved by turning radios off, and every remedy below is, one way or another, a way to do that safely.

Remedy 1: transmission power control

A node that sends with only enough power to reach its receiver saves transmit energy and disturbs fewer neighbours, which also cuts collisions and overhearing. How much it saves depends on the radio. The CC2420's datasheet (Table 9) lists its output power settings with the current each draws:

Output power0 dBm-1 dBm-3 dBm-5 dBm-7 dBm-10 dBm-15 dBm-25 dBm
Current17.4 mA16.5 mA15.2 mA13.9 mA12.5 mA11.2 mA9.9 mA8.5 mA

Worked example. Going from 0 dBm to -25 dBm cuts the radiated power by 25 dB, a factor of 10 raised to 2.5, about 316. The transmitter's current falls only from 17.4 mA to 8.5 mA, a factor of 17.4 / 8.5, about 2. Most of the current runs the electronics, not the antenna.

Now suppose low power means the packet needs two hops instead of one. Counting the current of the transmitter and of the receiver for one packet time per hop:

munotes.in244

Energy Efficiency in Ad Hoc Networks: Where the Energy Goes

  • One hop at 0 dBm: 17.4 + 18.8 = 36.2 mA-packet-times.
  • Two hops at -25 dBm: 2 × (8.5 + 18.8) = 2 × 27.3 = 54.6 mA-packet-times.

The low-power route costs half as much again, because each extra hop brings another receiver, and the receiver costs 18.8 mA whatever the transmitter does. On a radio like this, turning the power down pays only when it does not add hops. This is [Single Hop or Multiple Hops: The Energy Argument Worked Out] seen from the datasheet: at short range the electronics dominate.

Remedy 2: topology control, letting redundant nodes sleep

When nodes are deployed densely, many of them are interchangeable for forwarding. Topology control keeps just enough awake to carry the traffic and puts the rest to sleep, taking turns.

GAF: a virtual grid of equivalent nodes

GAF (Geographical Adaptive Fidelity) conserves energy by finding nodes that are equivalent from a routing point of view and turning off the ones not needed, while keeping the network's ability to route, its "routing fidelity", unchanged.

The virtual grid. GAF divides the area into square cells, small enough that every node in a cell can talk to every node in each neighbouring cell (left, right, up or down). The worst case is two nodes at opposite far corners of two adjacent cells: they are r apart one way and 2r the other, so with a radio range R the paper requires r² + (2r)² to be at most R², that is, r at most R divided by the square root of 5. All nodes in one cell are then equivalent for forwarding, and only one per cell need be awake.

Three square cells of side r in a row; the diagonal from one corner of the middle cell to the far corner of the next is at most R; one node awake in each cell and two equivalent nodes asleep

Figure 39.1 GAF's virtual grid: one awake node per cell is enough

Three states. Each node is sleeping, in discovery (radio on, exchanging discovery messages with the other nodes of its cell, carrying its node and cell identifiers, its estimated active time and its state), or active (the cell's forwarding node for a while). Nodes move between them on timers and on hearing discovery messages from higher-ranked nodes in their cell, so that the active role rotates and the nodes share the cost.

The results reported. In the paper's analysis and simulations, GAF consumed 40 to 60 per cent less energy than an unmodified ad hoc routing protocol, and network lifetime grew with density: in one example, four times the node density gave three to six times the lifetime, depending on the mobility pattern.

GAF's grid, computed

The program places 50, 100, 200 and 400 nodes at random in a 100 m square with a radio range of 25 m, builds GAF's grid, keeps one node awake in each occupied cell, and checks that the awake nodes still form a connected network whenever the whole network is connected.

munotes.in245

Energy Efficiency in Ad Hoc Networks: Where the Energy Goes

# GAF's virtual grid (Xu, Heidemann and Estrin 2001). Cells have side
# r = R / sqrt(5), so any node in a cell can reach any node in the next cell
# across or down. One node per occupied cell stays awake; the others sleep.
import math
import random
from collections import deque

SIDE, R = 100.0, 25.0                    # field side and radio range, in metres
r = R / math.sqrt(5)
print("range %.0f m: cell side %.2f m; adjacent cells' farthest points %.2f m apart"
      % (R, r, math.hypot(r, 2 * r)))

def connected(points):
    seen, todo = {0}, deque([0])
    while todo:
        a = todo.popleft()
        for b in range(len(points)):
            if b not in seen and math.dist(points[a], points[b]) <= R:
                seen.add(b)
                todo.append(b)
    return len(seen) == len(points)

print(" nodes  cells  awake  nodes per cell  all connected  awake connected")
for n in (50, 100, 200, 400):
    rnd = random.Random(n)
    pts = [(rnd.uniform(0, SIDE), rnd.uniform(0, SIDE)) for _ in range(n)]
    cells = {}
    for x, y in pts:
        cells.setdefault((int(x // r), int(y // r)), []).append((x, y))
    awake = [members[0] for members in cells.values()]
    print("%6d %6d %5.0f%% %15.1f %14s %16s" % (n, len(cells), 100 * len(cells) / n,
          n / len(cells), connected(pts), connected(awake)))
range 25 m: cell side 11.18 m; adjacent cells' farthest points 25.00 m apart
 nodes  cells  awake  nodes per cell  all connected  awake connected
    50     41    82%             1.2          False            False
   100     60    60%             1.7           True             True
   200     76    38%             2.6           True             True
   400     80    20%             5.0           True             True

Reading it.

  1. The rule holds exactly. With R = 25 m the cells are 11.18 m wide, and the farthest points of two adjacent cells are exactly 25.00 m apart: the square root of 11.18² + 22.36², which is R.
  2. Density is what GAF sells. With 50 nodes, 41 of the 81 cells are occupied and 82 per cent of the nodes must stay awake: almost nothing sleeps. With 400 nodes, 80 cells are occupied and only 20 per cent are awake at any moment; each cell's duty is shared by 5 nodes on average.
  3. Connectivity is kept. Whenever the whole network was connected, the awake nodes alone were connected too. (With 50 nodes the network was not connected to begin with.)
  4. Sharing is not the same as lifetime. Five nodes per cell could, taking turns, make each cell's service last up to about five times longer. But a cell with a single node gets no rest, and when it dies it may cut the network as surely as before; the paper's measured gain (three to six times for four times the density) is the realistic figure.
munotes.in246

Energy Efficiency in Ad Hoc Networks: Where the Energy Goes

Span: a backbone of coordinators

Span, from MIT, takes a similar idea without positions. Its starting point is that when a region has enough nodes, only a few need be awake to forward traffic. Each node decides locally whether to sleep or to join a forwarding backbone as a coordinator, based on an estimate of how many of its neighbours would benefit from its being awake and on how much energy it has left, and coordinators take turns over time. The paper reports that the gain grows with density and with the ratio of idle to sleep power, and that in its simulations an 802.11 network in power-saving mode lasted about twice as long with Span as without.

Remedy 3: sleep scheduling in the MAC

Topology control decides which nodes may sleep for long periods; sleep scheduling makes every node's radio sleep most of the time, waking briefly to talk. S-MAC puts neighbours on common listen and sleep schedules so they wake together; B-MAC and X-MAC let each node wake briefly to sample the channel and use long or strobed preambles to catch it. Their costs and designs are [Duty Cycling: Preamble Sampling, B-MAC and X-MAC] and the three S-MAC chapters.

Remedy 4: energy-aware routing and doing less

Energy-aware routing chooses routes by energy (the energy a route costs, or the battery its nodes have left) rather than by hop count, so that no node is drained by carrying everyone's traffic; that is the next chapter, [Energy-aware Routing]. And the cheapest packet is the one never sent: in-network processing and aggregation combine readings on the way to the sink ([Design Principles: Distributed Organisation and In-network Processing]).

Distinctions

TechniqueWhat it switches off or reducesWorks best whenExample
Transmission power controlTransmit powerTransmit power dominates the radio's currentSetting the CC2420's PA_LEVEL
Topology controlWhole nodes, for long periodsNodes are dense and redundantGAF, Span
Sleep schedulingEvery radio, most of the timeTraffic is lightS-MAC, B-MAC, X-MAC
Energy-aware routingThe load on weak nodesTraffic concentrates on some nodesNext chapter
In-network processingThe data itselfReadings are redundantAggregation
GAFSpan
Needs positionsYes (GPS or localisation)No
Who stays awakeOne node per virtual grid cellCoordinators forming a backbone
How chosenDiscovery messages and ranks within a cellLocal estimates of benefit and remaining energy
Reported gain40 to 60 per cent less energy; lifetime grows with densityAbout twice the lifetime with 802.11 power saving

What it does not mean

Lower transmit power does not always save energy. On the CC2420 it saves at most about half the transmitter's current, and if it adds hops it costs more.

munotes.in247

Energy Efficiency in Ad Hoc Networks: Where the Energy Goes

An on-demand routing protocol does not save much energy by itself. It sends fewer packets, but while radios listen all the time, idle listening dominates and the totals come out much the same.

Sleeping nodes are not lost nodes. In GAF and Span they are the reserve that makes the network last; the protocol keeps enough awake to preserve routing.

Density is not waste. Extra nodes that would otherwise overhear and collide become, with topology control, extra lifetime.

Quick revision

  • RFC 2501: energy-constrained operation; "sleep" period operation.
  • Radio states: transmit, receive, idle, sleep; idle : receive : transmit measured at 1 : 1.05 : 1.4 (Stemm and Katz), 1 : 2 : 2.5, 1 : 1.2 : 1.7; CC2420 receives at 18.8 mA, transmits at 17.4 mA.
  • Waste (S-MAC): collision, overhearing, control packet overhead, idle listening.
  • GAF's figure: counting idle energy, AODV, DSR, DSDV and TORA use about the same energy; idle dominates.
  • Transmission power control: CC2420 from 0 to -25 dBm cuts radiated power about 316 times but current only from 17.4 to 8.5 mA; one hop at 0 dBm (36.2) beats two at -25 dBm (54.6).
  • GAF: virtual grid, cell side r at most R divided by the square root of 5; one awake node per cell; states sleeping, discovery, active; 40 to 60 per cent less energy; lifetime grows with density.
  • Span: coordinators form a backbone, chosen by local benefit and energy; about 2 times the lifetime with 802.11 power saving.
  • Our grid: 400 nodes, 20 per cent awake, 5 nodes per cell; the awake set stays connected.
  • Also: sleep scheduling (S-MAC, B-MAC, X-MAC), energy-aware routing, aggregation.

Test yourself

1. What are the major sources of energy waste in an ad hoc or sensor network? Collisions, which force retransmissions; overhearing, receiving packets meant for other nodes; control packet overhead, energy spent on packets that carry no data; and idle listening, keeping the radio on to receive traffic that never comes, which in sensor networks is usually the largest.

2. Why does idle listening dominate the energy of an ad hoc network? Because an idle radio draws nearly as much power as a receiving one (measured ratios such as 1 : 1.05 : 1.4 for idle, receive and transmit), and radios spend most of their time idle, waiting for traffic. Simulations of four routing protocols found that, once idle energy was counted, all consumed about the same, so energy is saved mainly by turning radios off.

3. Explain GAF. Geographical Adaptive Fidelity divides the area into a virtual grid of square cells so small that any node in a cell can communicate with any node in an adjacent cell (side r with r² + (2r)² at most R², so r is at most R divided by the square root of 5). All nodes in a cell are then equivalent for routing, so one stays active while the others sleep, rotating through sleeping, discovery and active states so that they share the load. It saved 40 to 60 per cent of the energy in the paper's evaluation, and lifetime grows with node density.

munotes.in248

Energy Efficiency in Ad Hoc Networks: Where the Energy Goes

4. Does reducing transmission power always save energy? Use the CC2420 to explain. No. Cutting the CC2420's output from 0 dBm to -25 dBm reduces its current only from 17.4 mA to 8.5 mA, because most of the current runs the electronics, and the receiver draws 18.8 mA regardless. If the lower power means two hops instead of one, the cost per packet rises from 17.4 + 18.8 = 36.2 to 2 × (8.5 + 18.8) = 54.6 in the same units.

5. List the techniques for energy efficiency in ad hoc networks. Transmission power control; topology control that lets redundant nodes sleep (GAF, Span); sleep scheduling or duty cycling in the MAC (S-MAC, B-MAC, X-MAC); energy-aware routing; and reducing the data sent by in-network processing and aggregation.

Contents This chapter on its own page

munotes.in249

Chapter Forty

Energy-aware Routing

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Energy efficiency considerations in ad-hoc networks"

In one line

Energy-aware routing chooses routes by energy instead of hops: either the route that spends the least energy, or the route that spares the nodes with the least battery left, or, best, the cheapest route among those whose nodes are all still healthy, because a network dies when its first critical node dies, not when its total energy runs out.

In the wording a student can write in an examination: shortest-path routing sends traffic along the fewest hops, so the same nodes, especially those near the sink, relay again and again, drain first, and cut the network off while other nodes still have energy: the energy hole or hot spot problem. Energy-aware routing uses energy in the route metric. Minimum Total Transmission Power Routing (MTPR) chooses the route that needs the least total transmission power, which minimises energy per packet but can overuse particular nodes. Minimum Battery Cost Routing (MBCR) gives each node a cost that rises as its remaining battery falls (for example 1 / remaining capacity) and chooses the route with the smallest sum, but can still pick a route through one nearly empty node. Min-Max Battery Cost Routing (MMBCR) chooses the route whose weakest node is strongest, using batteries fairly at the price of more total energy. Conditional Max-Min Battery Capacity Routing (CMMBCR) uses MTPR among routes whose nodes all have more than a threshold of battery left, and MMBCR when no such route exists. The goal is to maximise network lifetime, not just to minimise energy per packet.

Why the fewest hops is not enough

The LEACH authors describe what happens when every node forwards along the minimum-energy path towards a base station: "the nodes closest to the base station will be used to route a large number of data messages to the base station. Thus these nodes will die out quickly, causing the energy required to get the remaining data to the base station to increase and more nodes to die. This will create a cascading effect that will shorten system lifetime."

Karl and Willig list the same problem among a sensor network's design issues: nonhomogeneous energy consumption, "the forming of 'hotspots'". And the LEACH authors summarise the answer that power-aware routing gives: "Routes that are longer, but which use nodes with more energy than the nodes along the shorter routes, are favored, helping avoid 'hot spots' in the network."

Two different goals hide under "save energy":

  1. Minimise the energy per packet, the sum over the route.
  2. Maximise the network's lifetime, however it is defined: until the first node dies, until the network partitions, or until some fraction of nodes is lost.

Toh puts the conflict plainly for the first goal: total transmission power "does not reflect directly on the lifetime of each host. If the minimum total transmission power routes are via a specific host, the battery of this host will be exhausted quickly".

munotes.in250

Energy-aware Routing

Toh's four metrics

1. Minimum Total Transmission Power Routing (MTPR). Each link's cost is the transmission power needed to reach the next node with an acceptable signal-to-noise ratio, which grows with distance (Toh uses a 1 / d to the power n roll-off, with n = 2 for short and n = 4 for longer distances). The route with the smallest total is found by an ordinary shortest-path algorithm such as Dijkstra's or Bellman-Ford. Because power grows faster than distance, MTPR tends to choose many short hops, which add delay and instability; adding each receiver's power to the cost, as one refinement does, pulls it back towards fewer hops.

2. Minimum Battery Cost Routing (MBCR). Let c be a node's remaining battery capacity, from 0 to 100. Its cost is f(c) = 1 / c, so "The less capacity it has, the more reluctant it is". A route's cost is the sum of its nodes' costs, and the cheapest route is chosen. The flaw: "because only the summation of values of battery cost functions is considered, a route containing nodes with little remaining battery capacity may still be selected".

3. Min-Max Battery Cost Routing (MMBCR). A route's cost is the largest f(c) among its nodes, that is, the cost of its weakest node, and the route whose weakest node is strongest is chosen. Batteries are used "more fairly", but "since there is no guarantee that minimum total transmission power paths will be selected under all circumstances, it can consume more power".

4. Conditional Max-Min Battery Capacity Routing (CMMBCR). Toh's own proposal, with a threshold γ between 0 and 100. "The basic idea behind CMMBCR is that when all nodes in some possible routes between a source and a destination have sufficient remaining battery capacity (i.e., above a threshold), a route with minimum total transmission power among these routes is chosen." When every route contains a node below γ, it chooses the route with the largest minimum capacity, as MMBCR would. With γ = 0 it is MTPR; with γ = 100 it is MMBCR. γ is "a protection margin".

Worked example: two routes, four metrics

A source has two routes to the sink. Route 1 has 2 relays, with 90 and 10 per cent of their batteries left; route 2 has 4 relays, each with 30 per cent. Route 1, with fewer hops, needs less total transmission power.

Two routes from S to the sink: route 1 through relays with 90 and 10 per cent battery, route 2 through four relays with 30 per cent each

Figure 40.1 The worked example: MBCR takes the short route through the weak relay, MMBCR the long one

munotes.in251

Energy-aware Routing

MetricRoute 1 (relays at 90, 10)Route 2 (relays at 30, 30, 30, 30)Chooses
Fewest hops, or MTPR2 hops, less power4 hops, more powerRoute 1
MBCR: sum of 1 / c1/90 + 1/10 = 10/90 = 1/9, about 0.1114 × 1/30 = 2/15, about 0.133Route 1
MMBCR: largest 1 / c1/10 = 0.11/30, about 0.033Route 2
CMMBCR, γ = 20Relay at 10 is below γAll relays at or above γRoute 2
CMMBCR, γ = 5All relays at or above γAll relays at or above γRoute 1 (less power)

MBCR still sends traffic through the relay with 10 per cent left, because the other relay's 90 per cent keeps the sum low: exactly the flaw Toh describes. MMBCR protects it. CMMBCR protects it only when it is actually in danger (below γ), and otherwise takes the cheaper route.

The four policies, run to the first death

The program places a sink and 12 nodes on a grid (columns at 20, 40, 60 and 80 m from the sink, rows at 20, 50 and 80 m), with a radio range of 40 m. Every round, every node sends one 2,000-bit packet to the sink along the route its policy chooses. Energy follows the first-order radio model of the LEACH papers: 50 nJ per bit for the electronics and 100 pJ per bit per square metre for the amplifier, with 0.5 J in each battery. The run stops when the first node dies.

# Four ways to choose routes, on one network, until the first node dies.
# Radio model: the first-order model of the LEACH papers (50 nJ/bit for the
# electronics, 100 pJ/bit/m^2 for the amplifier), 2,000-bit packets, 0.5 J each.
import heapq
import math

SINK = (0, 50)
NODES = [(x, y) for x in (20, 40, 60, 80) for y in (20, 50, 80)]
RANGE, BITS, START = 40.0, 2000, 0.5

def tx(d):                                 # joules to send one packet over d metres
    return BITS * (50e-9 + 100e-12 * d * d)

RX = BITS * 50e-9                          # joules to receive one packet

places = [SINK] + NODES                    # index 0 is the sink
links = {i: [j for j in range(len(places)) if j != i
             and math.dist(places[i], places[j]) <= RANGE] for i in range(len(places))}

def link_energy(i, j):                     # sending i to j, plus j receiving
    return tx(math.dist(places[i], places[j])) + (RX if j else 0)

def route(src, energy, policy, gamma=0.25):
    """The route from src to the sink (index 0) that a policy chooses."""
    if policy == "fewest hops":
        return dijkstra(src, energy, lambda i, j: 1)
    if policy == "least energy":           # MTPR
        return dijkstra(src, energy, link_energy)
    if policy == "strongest weakest":      # MMBCR
        return widest(src, energy)
    healthy = dijkstra(src, energy, link_energy,
                       allowed=lambda n: energy[n] >= gamma * START)
    return healthy or widest(src, energy)  # CMMBCR, threshold gamma

def dijkstra(src, energy, w, allowed=lambda n: True):
    best, heap = {src: 0.0}, [(0.0, src, [src])]
    while heap:
        c, n, path = heapq.heappop(heap)
        if n == 0:
            return path
        if c > best.get(n, math.inf):
            continue
        for m in links[n]:
            if m and (energy[m] <= 0 or not allowed(m)):
                continue
            nc = c + w(n, m)
            if nc < best.get(m, math.inf):
                best[m] = nc
                heapq.heappush(heap, (nc, m, path + [m]))
    return None

def widest(src, energy):
    """The route whose weakest relay has the most energy left."""
    best, heap = {src: math.inf}, [(-math.inf, src, [src])]
    while heap:
        neg, n, path = heapq.heappop(heap)
        if n == 0:
            return path
        for m in links[n]:
            if m and energy[m] <= 0:
                continue
            width = min(-neg, energy[m] if m else math.inf)
            if width > best.get(m, -1):
                best[m] = width
                heapq.heappush(heap, (-width, m, path + [m]))
    return None

for policy in ("fewest hops", "least energy", "strongest weakest", "conditional"):
    energy = [math.inf] + [START] * len(NODES)
    rounds, used = 0, 0.0
    while all(e > 0 for e in energy[1:]):
        for src in range(1, len(places)):    # every node reports once a round
            path = route(src, energy, policy)
            for a, b in zip(path, path[1:]):
                energy[a] -= tx(math.dist(places[a], places[b]))
                used += link_energy(a, b)
                if b:
                    energy[b] -= RX
        rounds += 1
    dead = [NODES[i - 1] for i in range(1, len(places)) if energy[i] <= 0]
    left = sum(energy[1:]) / (START * len(NODES))
    print("%-18s %.2f mJ a round; first death after %3d rounds, at %s; %2.0f%% unused"
          % (policy, 1000 * used / rounds, rounds, dead[0], 100 * left))
munotes.in252

Energy-aware Routing

fewest hops        8.32 mJ a round; first death after 288 rounds, at (20, 20); 60% unused
least energy       8.32 mJ a round; first death after 142 rounds, at (40, 50); 80% unused
strongest weakest  10.46 mJ a round; first death after 447 rounds, at (20, 20); 22% unused
conditional        8.67 mJ a round; first death after 404 rounds, at (40, 50); 42% unused

Reading it.

  1. The same energy, twice the lifetime difference. Fewest hops and least energy spend exactly the same energy per round, 8.32 mJ, yet the first node dies after 288 rounds in one case and after 142 in the other. The total is the same; where it is spent is not. Under least energy, the central node at (40, 50) spent its battery fastest and died after 142 rounds, about half as long.
  2. The energy hole. When the first node dies under least energy, 80 per cent of the network's energy is still unused; under fewest hops, 60 per cent. The network is broken, or about to be, while most of its batteries are full. That is the cascade the LEACH authors describe, caught at its first step.
  3. Sparing the weakest. "Strongest weakest" (MMBCR) spends more per round, 10.46 mJ, because it takes longer routes to avoid tired nodes, and yet lasts 447 rounds, and uses the batteries far more evenly: only 22 per cent is left when the first node dies. It is the most "fair", as Toh says, and here also the longest-lived.
  4. The compromise. The conditional policy (CMMBCR with γ at 25 per cent) spends almost as little as least energy, 8.67 mJ, while the nodes are healthy, and switches to protecting the weak when they are not: 404 rounds.
munotes.in253

Energy-aware Routing

A different network or radio model would change the numbers, and in a network where one node is the only way to the sink no metric can save it. What the run shows reliably is the shape: minimising energy per packet and maximising lifetime are different goals, and the metrics that watch the batteries serve the second.

The energy hole, and what else is done about it

Near the sink, every packet converges on the same few nodes: in a network where all traffic flows to one sink, the nodes one hop away relay everyone's data. Energy-aware routing spreads the load among them, but cannot abolish the funnel. Two other remedies appear in the sources held for this book:

  • Rotating the heavy role. LEACH rotates the cluster-head role among nodes at random, which its authors describe as achieving "the same goal" as power-aware routing ([LEACH: Clusters That Take Turns]).
  • Sending less. Aggregating data on the way to the sink cuts what the last hops must carry ([Design Principles: Distributed Organisation and In-network Processing]).

Distinctions

MetricRoute costFavoursWeakness
Minimum hopNumber of hopsShort, fast routesIgnores energy; overuses central nodes
MTPRSum of transmission powerLeast energy per packetOveruses some nodes; many short hops
MBCRSum of 1 / remaining capacityRoutes with plenty of battery overallCan still use a nearly empty node
MMBCRLargest 1 / remaining capacityRoutes whose weakest node is strongestMay spend more total energy
CMMBCRMTPR among routes above γ, else MMBCRLow energy while all are healthyDepends on choosing γ
Energy per packetNetwork lifetime
What it measuresThe cost of one deliveryHow long the network keeps working
Minimised or maximised byMTPRMMBCR, CMMBCR
Depends onThe route's linksHow evenly the load falls on the nodes

What it does not mean

Energy-aware does not mean shortest. The routes energy-aware metrics choose are often longer, deliberately.

Least energy per packet does not mean longest life. In the run above it gave the shortest lifetime of the four.

munotes.in254

Energy-aware Routing

A network with energy left is not necessarily alive. When the nodes around the sink die, the rest cannot reach it, however full their batteries.

Battery-aware routing is not free. Nodes must learn their neighbours' remaining energy, which costs messages, and the extra route length costs energy every packet.

Quick revision

  • Energy hole / hot spot: nodes near the sink relay everyone's traffic and die first, in a cascade (LEACH on MTE), leaving most energy unused.
  • Goals: minimum energy per packet against maximum network lifetime (first death, partition).
  • MTPR: minimise total transmission power; Dijkstra or Bellman-Ford; many short hops; overuses some nodes.
  • MBCR: minimise the sum of f(c) = 1 / c; can still choose a nearly empty node.
  • MMBCR: minimise the largest 1 / c (strongest weakest node); fair, may cost more power.
  • CMMBCR: threshold γ; MTPR among routes whose nodes are all above γ, otherwise MMBCR; γ = 0 is MTPR, γ = 100 is MMBCR.
  • Worked example: MBCR picks the route through the relay at 10 per cent (1/9 against 2/15); MMBCR avoids it.
  • Our run: fewest hops 288 rounds, least energy 142, MMBCR 447, CMMBCR 404; least energy leaves 80 per cent unused at the first death.

Test yourself

1. What is energy-aware routing? Why is it needed? Routing that uses energy, either the energy a route consumes or the remaining battery of its nodes, as the route selection metric instead of hop count. It is needed because shortest-path routing makes the same nodes, especially those near the sink, relay most of the traffic, so they die early and cut off the network while other nodes still have energy.

2. Explain MTPR, MBCR and MMBCR. MTPR chooses the route with the minimum total transmission power, minimising energy per packet but possibly overusing some nodes. MBCR gives each node a cost that grows as its battery falls, such as 1 / remaining capacity, and chooses the route with the smallest total cost; it can still select a route through a nearly exhausted node if the others are well charged. MMBCR chooses the route whose most expensive node (the one with the least battery) is cheapest, so the weakest node is protected and batteries are used fairly, at the cost of sometimes using more total power.

3. Explain CMMBCR. Conditional max-min battery capacity routing uses a threshold γ. If there are routes in which every node has more than γ battery left, it chooses among them the one with minimum total transmission power. If every route contains a node below γ, it chooses the route whose weakest node has the most battery, as MMBCR does. With γ = 0 it behaves like MTPR and with γ = 100 like MMBCR.

munotes.in255

Energy-aware Routing

4. What is the energy hole problem? In a network where all data flows to a sink, the nodes nearest the sink relay the traffic of all the others, so they use energy fastest and die first. Their death disconnects the rest of the network, or forces longer, costlier routes that kill more nodes, even though most of the network's energy is still unused.

5. A route has relays with 50 and 25 per cent battery; another has three relays at 40 per cent. Which does MBCR choose, and which MMBCR? MBCR: 1/50 + 1/25 = 3/50 = 0.06 against 3 × 1/40 = 3/40 = 0.075, so the first route. MMBCR: the first route's weakest relay costs 1/25 = 0.04, the second's 1/40 = 0.025, so the second route.

Contents This chapter on its own page

munotes.in256

Chapter Forty-One

Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Security and privacy issues in ad-hoc networks"

In one line

A sensor network must keep its data secret, genuine, unaltered and fresh, and keep working under attack, with nodes too small for ordinary cryptography, lying unattended in the open, talking over a radio anyone can hear, and each relaying for the others; so attacks come at every layer, from jamming the radio to forging routes and desynchronising connections.

In the wording a student can write in an examination: the security goals of a sensor network are data confidentiality (readings are not disclosed to outsiders), data authentication (a receiver can verify who sent a message), data integrity (a message is not altered in transit), data freshness (a message is recent, not a replay), and availability (the network keeps working in spite of denial-of-service attacks). Security is harder than in ordinary networks because nodes have little memory, processing and energy (public-key cryptography does not fit), communication is wireless and easily overheard or jammed, nodes are unattended and can be captured and tampered with, and every node is a router. Attackers may be mote-class or laptop-class, and outsiders or insiders holding captured keys. Denial-of-service attacks occur at every layer: physical (jamming, tampering), link (collision, exhaustion, unfairness), network (neglect and greed, homing, misdirection, black holes) and transport (flooding, desynchronisation), each with its defences.

What must be protected

SPINS, a set of security protocols built for sensor networks, sets out the requirements:

  1. Data confidentiality. "A sensor network should not leak sensor readings to neighboring networks." Readings and keys are kept secret by encrypting them "with a secret key that only intended receivers possess".
  2. Data authentication. Because "an adversary can easily inject messages, the receiver needs to ensure that data used in any decision-making process originates from a trusted source". Between two nodes a shared key and a message authentication code (MAC) suffice. For a broadcast they do not: with a shared MAC key, "any one of the receivers knows the MAC key, and hence, could impersonate the sender".
  3. Data integrity: "the received data is not altered in transit by an adversary". SPINS obtains it from authentication, "which is a stronger property".
  4. Data freshness: "the data is recent, and it ensures that no adversary replayed old messages". SPINS separates weak freshness, a partial ordering good enough for sensor readings, from strong freshness, a total order on a request and its response, needed for time synchronisation.

To these add availability. Wood and Stankovic define a denial-of-service attack broadly: "any event that diminishes or eliminates a network's capacity to perform its expected function". For a network that raises alarms, a network that is silent under attack is a failed network, whatever its encryption.

Why it matters. Wood and Stankovic's examples go beyond the battlefield: protecting "the location and status of casualties" after a disaster; in public safety, false alarms "could cause panic or disregard for warning systems"; in home healthcare, privacy is "paramount". And their advice on method: "Attempts to add security afterwards usually prove unsuccessful."

munotes.in257

Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks

Why a sensor network is harder to protect

The node is too small for ordinary cryptography. SPINS was built for motes with an 8-bit processor at 4 MHz, 8 kilobytes of instruction flash, 512 bytes of RAM and a 10 kbit/s radio, on which TinyOS took about 3,500 bytes of the flash, leaving 4,500 bytes for security and the application. On such a node the working memory "is not sufficient to even hold the variables for asymmetric cryptographic algorithms", such as RSA with 1,024-bit keys, "let alone perform operations with them". Digital signatures would also add 50 to 1,000 bytes to every packet, by SPINS's count.

Energy is the budget for security too. "Wireless communication is the most energy-consuming function performed by these devices, so we need to minimize communications overhead." Every byte a security protocol adds is paid for in lifetime.

The medium is open. RFC 2501 notes that mobile wireless networks "are generally more prone to physical security threats than are fixed-cable nets", with "eavesdropping, spoofing, and denial-of-service attacks". Dargie and Poellabauer add that "wireless communications make it easy for an adversary to eavesdrop on sensor transmissions".

The nodes are unattended. "The remote and unattended operation of sensor nodes increases their exposure to malicious intrusions and attacks." Wood and Stankovic: "we cannot expect to control access to hundreds of nodes spread over several kilometers"; an attacker can capture a node and "extract sensitive material such as cryptographic keys".

Every node is a router. "Since every node is potentially a router, this adds new vulnerabilities to the network-layer problems experienced on the Internet." A captured node can lie about routes, and all traffic through it is at its mercy; that is the next chapter.

Worked example: what a signature costs on the radio. SPINS's motes send at 10 kbit/s. A signature of 50 bytes is 50 × 8 = 400 bits, which takes 400 / 10,000 = 0.04 s, 40 ms, to send; one of 1,000 bytes is 8,000 bits, 0.8 s. A typical sensor reading fits in a few bytes, so the signature would take tens to hundreds of times longer to send than the data it protects, on the most energy-hungry part of the node. That is why sensor security is built from symmetric primitives, as SPINS is ([Keys and Link Security: Key Predistribution, SPINS and 802.15.4]).

Who the attacker is

Karlof and Wagner draw two distinctions:

  1. Mote-class and laptop-class attackers. A mote-class attacker "has access to a few sensor nodes with similar capabilities to our own". A laptop-class attacker may have "greater battery power, a more capable CPU, a high-power radio transmitter, or a sensitive antenna": where an ordinary node "might only be able to jam the radio link in its immediate vicinity", a laptop-class attacker "might be able to jam the entire sensor network", or eavesdrop on all of it.
  2. Outsiders and insiders. An outsider "has no special access to the sensor network". An insider is "an authorized participant in the sensor network" that "has gone bad", such as a captured node with its keys. Encryption and authentication stop outsiders; they do not stop an insider who holds valid keys.
munotes.in258

Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks

Denial of service, layer by layer

Wood and Stankovic's Table 1 is the map examiners expect:

LayerAttackDefences
PhysicalJammingSpread spectrum, priority messages, lower duty cycle, region mapping, mode change
PhysicalTamperingTamper-proofing, hiding
LinkCollisionError-correcting code
LinkExhaustionRate limitation
LinkUnfairnessSmall frames
Network and routingNeglect and greedRedundancy, probing
Network and routingHomingEncryption
Network and routingMisdirectionEgress filtering, authorisation, monitoring
Network and routingBlack holesAuthorisation, monitoring, redundancy
TransportFloodingClient puzzles
TransportDesynchronisationAuthentication

Physical layer

Jamming "interferes with the radio frequencies a network's nodes are using". A few jammers can silence many nodes, and for single-frequency networks "this attack is simple and effective". The standard defence is spread spectrum, which low-cost sensor radios often lack. Otherwise, nodes can lower their duty cycle and wait it out ("By spending energy frugally, the nodes may be able to outlive an adversary, who must continue to jam at greater expense"), send short high-priority reports in gaps, have the nodes around a jammed region map its boundary and report it, or switch mode to another medium such as infrared or optical.

Tampering: physical capture, damage or analysis of a node, including extracting its keys. Defences are tamper-proofing the package, reacting "in a fail-complete manner" (for example erasing keys and programs), and hiding or camouflaging nodes.

Worked example: outlasting a jammer. Suppose a mote-class jammer transmits continuously with a CC2420 at full power, 17.4 mA, from the same two AA cells this book allows a node, 2,500 mAh. It lasts 2,500 / 17.4 hours, about 144 hours, a little under 6 days. A node that answers the jamming by dropping to a 1 per cent duty cycle draws, by the program in [How Long a Node Lasts: The Energy Budget Worked Out], about 0.215 mA on average, which would last about 1.3 years. The jammer runs out long before the node: Wood and Stankovic's strategy, in numbers. A laptop-class jammer with a car battery is another matter.

munotes.in259

Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks

Link layer

Collision. An attacker "may only need to induce a collision in one octet of a transmission to disrupt an entire packet", at almost no energy cost to itself. Error-correcting codes tolerate some corruption, but a malicious node can always corrupt more than the code corrects.

Exhaustion. Repeatedly provoking retransmissions, or, with RTS and CTS, sending request after request so the victim keeps answering, until both batteries run down. The defence is rate limitation: the MAC ignores excessive requests.

Unfairness. Using these tricks intermittently, or abusing a priority scheme, so that others get the channel less often and miss deadlines. Small frames limit how long any one node can hold the channel.

Network and routing layer

Neglect and greed. A node that drops messages "on a random or arbitrary basis" is neglectful; if it also favours its own traffic, it is greedy. Redundancy (several paths or duplicate messages) and probing limit the damage.

Homing. A passive attacker watches traffic to find the important nodes (cluster heads, key managers, the base station) and then attacks them. Encrypting headers hop by hop hides where traffic goes.

Misdirection. Forwarding messages "along wrong paths, perhaps by fabricating malicious route advertisements", which can aim a flood of traffic at a chosen victim. Egress filtering (a parent checks that packets from below really came from its children), authorisation of routing messages, and monitoring defend.

Black holes. In distance-vector routing, a node advertises "zero-cost routes to every other node", so traffic flows towards it and is lost, and its neighbours are exhausted by the contention. Authorisation (only authorised nodes may send routing information), monitoring (a watchdog checks that the next hop forwards what it was given) and redundancy defend. The routing attacks that grow out of this are [Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood].

Transport layer

Flooding. Sending many connection requests so the victim fills its memory with half-open connections. Client puzzles make each requester prove it has spent computation before the victim commits resources.

Desynchronisation. Forging messages with sequence numbers or control flags that make two end points keep asking for retransmissions, "causing them to waste energy in an endless synchronization-recovery protocol". Authenticating every packet, including the transport header's control fields, lets the end points ignore forgeries.

What protection costs, and what capture buys

The constraints above are not rhetorical. The program prices them: what a message authentication code costs in the frame, what key predistribution buys and what it gives away when nodes are captured, and why an attacker with a screwdriver never bothers with the cipher.

# What protection costs a node, and what an attacker gets for a captured one.
import math

FRAME_ABOVE_MAC = 127 - 25            # octets left after the 802.15.4 MAC header
CCM_MAX = 21                          # AES-CCM-128 at its largest, RFC 4944's figure
RATE, I_TX, V = 250e3, 17.4e-3, 3.0   # CC2420: 250 kbit/s, 17.4 mA at 0 dBm, two cells

print("What link security costs, in the only budget a node has:")
for mic, label in ((4, "a 4 octet check"), (8, "an 8 octet check"),
                   (16, "a 16 octet check"), (CCM_MAX, "the largest CCM overhead")):
    share = 100.0 * mic / FRAME_ABOVE_MAC
    energy = mic * 8 / RATE * I_TX * V
    print("  %-26s %2d octets of %d, %4.1f%% of the frame, %6.2f uJ to send"
          % (label, mic, FRAME_ABOVE_MAC, share, energy * 1e6))
print("  the cost is not the arithmetic, which a mote does in microseconds. It is the")
print("  octets: every check byte is a byte of reading that did not fit in the frame.")

# Key predistribution: the chance two nodes share a key, from a pool of P with k
# keys each (Eschenauer and Gligor). 1 minus the chance their sets are disjoint.
def share_probability(pool, k):
    p = 1.0
    for i in range(k):
        p *= float(pool - k - i) / (pool - i)
    return 1 - p

print()
print("Key predistribution: a pool of keys, a few given to each node before it is")
print("deployed, and two neighbours can talk only if they happen to hold one in common.")
print("  pool" + "".join("%12s" % ("%d keys" % k) for k in (10, 25, 50, 100, 200)))
for pool in (1000, 5000, 10000, 100000):
    print("  %6d" % pool
          + "".join("%11.1f%%" % (100 * share_probability(pool, k))
                    for k in (10, 25, 50, 100, 200)))
print("  a node holding 200 keys out of 10,000 shares one with a neighbour %.0f%% of the"
      % (100 * share_probability(10000, 200)))
print("  time, on 200 keys of storage: that is the whole trick, and it is why the scheme")
print("  exists at all on a part with kilobytes of memory.")

# What capture costs. Each captured node hands over its k keys.
print()
print("And what the same scheme costs when nodes are captured, pool of 10,000:")
print("  keys a node holds" + "".join("%14s" % ("%d captured" % c) for c in (1, 5, 20, 50)))
for k in (25, 50, 100, 200):
    row = []
    for c in (1, 5, 20, 50):
        exposed = 1 - (1 - float(k) / 10000) ** c
        row.append("%13.1f%%" % (100 * exposed))
    print("  %10d keys   " % k + "".join(row))
print("  the same number that makes two neighbours likely to share a key makes a handful")
print("  of captured nodes likely to hold yours. Connectivity and resilience are the")
print("  same dial turned in opposite directions, which is the finding of the scheme and")
print("  not a flaw in it.")

# Why the attacker does not need to break the cryptography.
print()
print("Why the attacker does not attack the cipher:")
print("  a 128 bit key has 2^128 values, so a machine trying a million million a second")
years = 2.0 ** 128 / 1e12 / (365.25 * 24 * 3600)
print("  needs about %.1e years, which is %.1e times the age of the universe."
      % (years, years / 1.38e10))
print("  Walking to the node and reading the key out of its memory takes an afternoon.")
print("  That is why the chapter's threat model starts with physical capture and not")
print("  with cryptanalysis.")
munotes.in260

Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks

What link security costs, in the only budget a node has:
  a 4 octet check             4 octets of 102,  3.9% of the frame,   6.68 uJ to send
  an 8 octet check            8 octets of 102,  7.8% of the frame,  13.36 uJ to send
  a 16 octet check           16 octets of 102, 15.7% of the frame,  26.73 uJ to send
  the largest CCM overhead   21 octets of 102, 20.6% of the frame,  35.08 uJ to send
  the cost is not the arithmetic, which a mote does in microseconds. It is the
  octets: every check byte is a byte of reading that did not fit in the frame.

Key predistribution: a pool of keys, a few given to each node before it is
deployed, and two neighbours can talk only if they happen to hold one in common.
  pool     10 keys     25 keys     50 keys    100 keys    200 keys
    1000        9.6%       47.3%       92.8%      100.0%      100.0%
    5000        2.0%       11.8%       39.6%       87.0%      100.0%
   10000        1.0%        6.1%       22.2%       63.6%       98.3%
  100000        0.1%        0.6%        2.5%        9.5%       33.0%
  a node holding 200 keys out of 10,000 shares one with a neighbour 98% of the
  time, on 200 keys of storage: that is the whole trick, and it is why the scheme
  exists at all on a part with kilobytes of memory.

And what the same scheme costs when nodes are captured, pool of 10,000:
  keys a node holds    1 captured    5 captured   20 captured   50 captured
          25 keys             0.2%          1.2%          4.9%         11.8%
          50 keys             0.5%          2.5%          9.5%         22.2%
         100 keys             1.0%          4.9%         18.2%         39.5%
         200 keys             2.0%          9.6%         33.2%         63.6%
  the same number that makes two neighbours likely to share a key makes a handful
  of captured nodes likely to hold yours. Connectivity and resilience are the
  same dial turned in opposite directions, which is the finding of the scheme and
  not a flaw in it.

Why the attacker does not attack the cipher:
  a 128 bit key has 2^128 values, so a machine trying a million million a second
  needs about 1.1e+19 years, which is 7.8e+08 times the age of the universe.
  Walking to the node and reading the key out of its memory takes an afternoon.
  That is why the chapter's threat model starts with physical capture and not
  with cryptanalysis.
munotes.in261

Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks

Distinctions

GoalProtects againstTypical mechanism
ConfidentialityEavesdroppingEncryption with a shared secret key
AuthenticationInjected or impersonated messagesMessage authentication code
IntegrityAltered messagesObtained through authentication
FreshnessReplayed messagesCounters or nonces (weak or strong freshness)
AvailabilityDenial of serviceThe layer-by-layer defences above
munotes.in262

Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks

Mote-class attackerLaptop-class attacker
EquipmentA few nodes like the network's ownMore energy, CPU, transmitter power, antenna
ReachJams or overhears nearbyMay jam or overhear the whole network
OutsiderInsider
AccessNo special accessAn authorised node gone bad, with valid keys
Stopped by encryption and authenticationLargely yesNo

What it does not mean

Encryption is not security. It gives confidentiality against outsiders; it does nothing against jamming, a captured node with its keys, or a node that drops packets.

A MAC is not a signature. A message authentication code proves a message came from someone holding the shared key; in a broadcast, every receiver holds it and could forge.

Denial of service is not only flooding. It is anything that reduces the network's ability to do its job, including collisions of a single octet and quiet, selective dropping.

A failed node and an attacked node can look alike. A neglectful relay and a broken one both lose packets, which is why prevention is safer than detection.

Quick revision

  • Goals (SPINS): confidentiality, authentication, integrity, freshness (weak and strong); plus availability (Wood and Stankovic: DoS is "any event that diminishes or eliminates a network's capacity to perform its expected function").
  • Harder because: tiny nodes (SPINS mote: 4 MHz 8-bit CPU, 8 KB flash, 512 B RAM, 10 kbit/s; RSA-1024's variables do not fit; signatures 50 to 1,000 bytes a packet), energy, open radio, unattended nodes that can be captured, every node a router.
  • Attackers (Karlof and Wagner): mote-class or laptop-class; outsider or insider.
  • DoS by layer (Wood and Stankovic): physical jamming (spread spectrum, lower duty cycle, priority, region mapping, mode change), tampering (tamper-proofing, hiding); link collision (error-correcting codes), exhaustion (rate limitation), unfairness (small frames); network neglect and greed (redundancy, probing), homing (encryption), misdirection (egress filtering, authorisation, monitoring), black holes (authorisation, monitoring, redundancy); transport flooding (client puzzles), desynchronisation (authentication).
  • Worked: a signature of 50 to 1,000 bytes takes 40 ms to 0.8 s at 10 kbit/s; a continuous CC2420 jammer on 2,500 mAh lasts about 6 days, a node at 1 per cent duty cycle about 1.3 years.
munotes.in263

Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks

Test yourself

1. What are the security requirements of a wireless sensor network? Data confidentiality, so that readings and keys are not disclosed; data authentication, so that a receiver can verify the sender; data integrity, so that messages are not altered in transit; data freshness, so that old messages cannot be replayed; and availability, so that the network keeps functioning in spite of denial-of-service attacks.

2. Why is security in sensor networks more difficult than in conventional networks? The nodes have little memory, processing power and energy, so public-key cryptography and long signatures are impractical; communication is wireless and can be overheard, injected or jammed by anyone in range; nodes are unattended and can be captured and their keys extracted; every node acts as a router, so a compromised node can disrupt routing; and security overhead is paid for directly in network lifetime.

3. Explain denial-of-service attacks at the physical and link layers, with defences. At the physical layer, jamming interferes with the radio frequencies (defences: spread spectrum, lower duty cycles to outlast the jammer, high-priority messages in gaps, mapping the jammed region, switching communication mode), and tampering physically compromises nodes (tamper-proofing, erasing keys when tampered with, hiding nodes). At the link layer, collision corrupts packets deliberately (error-correcting codes), exhaustion provokes endless retransmissions or requests to drain batteries (rate limitation), and unfairness denies others fair channel access (small frames).

4. Distinguish mote-class from laptop-class, and outsider from insider attackers. A mote-class attacker has a few devices like the network's own nodes and can affect only its surroundings; a laptop-class attacker has more powerful devices, with more energy, a stronger transmitter and a better antenna, and may jam or overhear the whole network. An outsider has no authorised access to the network; an insider is an authorised node, typically captured, that holds valid keys, which encryption and authentication cannot stop.

5. What is a black hole attack, and how is it countered? A malicious node advertises very cheap routes (in distance-vector routing, zero-cost routes to everyone), so neighbouring nodes route traffic through it, and it discards the traffic; the contention around it also exhausts its neighbours. It is countered by allowing only authorised nodes to send routing information, by monitoring neighbours' forwarding with watchdogs, and by sending over redundant paths.

Contents This chapter on its own page

munotes.in264

Chapter Forty-Two

Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Security and privacy issues in ad-hoc networks"

In one line

In a sensor network every node routes for the others and all traffic heads for one base station, so an attacker who can make itself look like the best way there, by lying about routes, pretending to be many nodes, tunnelling packets across the network or shouting louder than everyone, can draw in the traffic of a whole region and then drop or alter it at will.

In the wording a student can write in an examination: Karlof and Wagner list the attacks on sensor network routing as: (1) spoofed, altered or replayed routing information, which can create loops, attract or repel traffic and partition the network; (2) selective forwarding, in which a malicious node drops some packets (a black hole drops all); (3) sinkhole attacks, in which a compromised node makes itself look especially attractive, for example by advertising a high-quality route to the base station, so that traffic from a whole area passes through it; (4) the Sybil attack, in which one node presents several identities, defeating multipath routing, distributed storage and geographic routing; (5) wormholes, in which colluding attackers tunnel packets over a fast out-of-band link and replay them elsewhere, making distant nodes appear close; (6) the HELLO flood, in which a laptop-class attacker broadcasts with high power so that every node believes the attacker is its neighbour; and (7) acknowledgement spoofing, forging link-layer acknowledgements so that weak or dead links look good. Countermeasures include link-layer encryption and authentication (against outsiders), identity verification through the base station (against Sybil), bidirectional link verification (against the HELLO flood), multipath routing (against selective forwarding), authenticated broadcast, and routing designs such as geographic routing in which sinkholes and wormholes are meaningless.

What secure routing should achieve

Karlof and Wagner set the ideal: "a secure routing protocol should guarantee the integrity, authenticity, and availability of messages in the presence of adversaries of arbitrary power." Secrecy of the data itself they leave to the application, but not eavesdropping that the routing protocol makes possible: "Eavesdropping achieved by the cloning or rerouting of a data flow should be prevented."

Against insiders, perfection is out of reach, so they ask for graceful degradation: the protocol's effectiveness "should degrade no faster than a rate approximately proportional to the ratio of compromised nodes to total nodes in the network". Capture one node in a hundred and lose about one hundredth of the network, not all of it.

Why sensor networks are easy prey. Many of their routing protocols "are quite simple, and for this reason are sometimes even more susceptible to attacks". And their traffic pattern helps the attacker: "Since all packets share the same ultimate destination (in networks with only one base station), a compromised node needs only to provide a single high quality route to the base station in order to influence a potentially large number of nodes."

munotes.in265

Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood

The seven attacks

1. Spoofed, altered or replayed routing information. The most direct attack: by forging, changing or replaying routing messages, an attacker "may be able to create routing loops, attract or repel network traffic, extend or shorten source routes, generate false error messages, partition the network, increase end-to-end latency".

2. Selective forwarding. "Malicious nodes may refuse to forward certain messages and simply drop them." Dropping everything makes the node a black hole, which neighbours may notice and route around; dropping only some packets, for example everything from a few chosen nodes, attracts less suspicion. It works best when the attacker is on the path of the flow, which is what the next two attacks achieve.

3. Sinkhole attacks. "The adversary's goal is to lure nearly all the traffic from a particular area through a compromised node, creating a metaphorical sinkhole with the adversary at the center." The usual trick is to advertise "an extremely high quality route to a base station"; each neighbour then forwards through the attacker and passes the attractive route on to its own neighbours, so the attacker gains a large "sphere of influence". A sinkhole "makes selective forwarding trivial".

4. The Sybil attack. "A single node presents multiple identities to other nodes in the network." Schemes that rely on several independent nodes (replicas, distributed storage, disjoint routes) may in fact be relying on one attacker. It is especially dangerous to geographic routing: a node normally accepts one position from each neighbour, but with many identities the attacker can "be in more than one place at once".

5. Wormholes. "An adversary tunnels messages received in one part of the network over a low latency link and replays them in a different part." Usually two colluding nodes, far apart, relay packets over a channel only they have. Near a base station a wormhole is devastating: nodes many hops away are convinced they are "only one or two hops away via the wormhole", which creates a sinkhole at the far end. A wormhole can also simply convince two distant nodes that they are neighbours.

6. The HELLO flood. Many protocols have nodes broadcast HELLO packets, and a node hearing one "may assume that it is within (normal) radio range of the sender". A laptop-class attacker "with large enough transmission power could convince every node in the network that the adversary is its neighbor". Nodes far away then send their packets towards a neighbour that cannot hear them, "into oblivion". The attacker need not even forge anything: it can "simply re-broadcast overhead packets with enough power to be received by every node in the network". Despite the name, it is not a multi-hop flood but "a single hop broadcast".

munotes.in266

Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood

7. Acknowledgement spoofing. Over a broadcast medium, an attacker can forge link-layer acknowledgements for packets meant for its neighbours, "convincing the sender that a weak link is strong or that a dead or disabled node is alive". Packets then go down links that lose them, a quiet form of selective forwarding.

The attacks, run against TinyOS beaconing

Karlof and Wagner analyse the simplest collection protocol, TinyOS beaconing, which "constructs a breadth first spanning tree rooted at a base station": the base station broadcasts a route update, and every node takes as its parent "the first node from which it hears a routing update" and rebroadcasts it. The program builds that tree on 100 nodes in a 10 by 10 grid, 10 m apart, with the base station in one corner and a radio range of 12 m, so that each node hears only its grid neighbours. Then it attacks it four ways. For each node except the base station and the attackers, it follows the parents to see where the node's data really ends up: at the base station, at an attacker, or nowhere ("stranded", when a node's parent is too far away to hear it), and whether it passes through a wormhole.

# TinyOS beaconing (Karlof and Wagner, section VII-A): the base station floods a
# route update and every node takes as its parent the first node it hears it
# from. One tree, then four attacks on it. 100 nodes on a 10 by 10 grid, 10 m
# apart, with a radio range of 12 m, so each node hears only its grid neighbours.
import heapq
import math

N, GAP, RANGE = 10, 10.0, 12.0
nodes = [(x * GAP, y * GAP) for y in range(N) for x in range(N)]
at = {p: i for i, p in enumerate(nodes)}
BASE = at[(0.0, 0.0)]

def build(seeds, loud=None, check=False):
    """seeds: {node: (time it sends the update, its own parent)}. A loud node's
    update reaches every node (a HELLO flood). With check=True a node accepts a
    parent only over a link that works both ways. Returns each node's parent."""
    parent, heap = {}, [(t, n, p) for n, (t, p) in seeds.items()]
    heapq.heapify(heap)
    while heap:
        t, n, p = heapq.heappop(heap)
        if n in parent:
            continue                          # it already heard the update
        parent[n] = p
        for m in range(len(nodes)):
            if m in parent:
                continue
            close = math.dist(nodes[n], nodes[m]) <= RANGE
            if close or (n == loud and not check):
                heapq.heappush(heap, (t + 1, m, n))
    return parent

def fate(parent, n, tunnel=()):
    """Follow n's parents. Where does its data end up, and does it pass the tunnel?"""
    through = False
    while parent[n] is not None:
        p = parent[n]
        if (n, p) in tunnel:
            through = True
        elif math.dist(nodes[n], nodes[p]) > RANGE:
            return "stranded", through        # the parent cannot hear it
        n = p
    return n, through

def report(name, parent, attackers=(), tunnel=()):
    fates = [fate(parent, n, tunnel) for n in range(len(nodes))
             if n != BASE and n not in attackers]
    ends = [f for f, _ in fates]
    print("%-36s base %2d  attacker %2d  stranded %2d  through the tunnel %2d"
          % (name, ends.count(BASE), sum(e in attackers for e in ends),
             ends.count("stranded"), sum(t for _, t in fates)))

report("no attack", build({BASE: (0, None)}))
fake = at[(90.0, 90.0)]                       # a mote that claims to be a base station
report("spoofed base station at (90, 90)",
       build({BASE: (0, None), fake: (0, None)}), {fake})
loud = at[(50.0, 50.0)]                       # a laptop-class attacker in the middle
report("HELLO flood from (50, 50)",
       build({BASE: (0, None), loud: (0.5, None)}, loud=loud), {loud})
report("the same, links checked both ways",
       build({BASE: (0, None), loud: (0.5, None)}, loud=loud, check=True), {loud})
near, far = at[(10.0, 0.0)], at[(70.0, 70.0)] # two colluding nodes and their tunnel
report("wormhole from (10, 0) to (70, 70)",
       build({BASE: (0, None), far: (1, near)}), {near, far}, {(far, near)})
munotes.in267

Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood

no attack                            base 99  attacker  0  stranded  0  through the tunnel  0
spoofed base station at (90, 90)     base 54  attacker 44  stranded  0  through the tunnel  0
HELLO flood from (50, 50)            base  2  attacker  4  stranded 92  through the tunnel  0
the same, links checked both ways    base 28  attacker 70  stranded  0  through the tunnel  0
wormhole from (10, 0) to (70, 70)    base 97  attacker  0  stranded  0  through the tunnel 59

Reading it, line by line.

  1. No attack. All 99 nodes reach the base station. Each has as parent a neighbour one hop closer, which is what a breadth-first tree is.
  2. A spoofed base station. Routing updates in TinyOS beaconing are not authenticated, so "it is possible for any node to claim to be a base station and become the destination of all traffic in the network". A single ordinary mote in the far corner, claiming to be a base station, becomes the parent of every node nearer to it than to the real one: 44 of 99 nodes now send their data to the attacker. The attacker is a mote-class sinkhole, and it did nothing but send one forged update.
  3. A HELLO flood. A laptop-class attacker in the middle hears the base station's update at once (with a sensitive receiver) and rebroadcasts it loudly enough for every node to hear. Every node that has not yet heard the update from a real neighbour, 96 of them, takes the attacker as its parent. The 4 that are its real neighbours send to it; the other 92 are stranded, "sending packets into oblivion" towards a parent that cannot hear them. Only the base station's own 2 neighbours escape.
  4. The two-way check, and what it does not fix. Karlof and Wagner's simplest defence is "to verify the bidirectionality of a link before taking meaningful action based on a message received over that link". With it, only the attacker's real neighbours accept it as parent, and nobody is stranded. But the attacker still had the update first, in the middle of the network, so its neighbours rebroadcast it before the real wavefront arrives: 70 nodes end up routing through the attacker. The check has turned a HELLO flood into a sinkhole. Stopping that needs updates that cannot be replayed early, or a routing design in which being first means nothing.
  5. A wormhole. Two colluding nodes, one beside the base station at (10, 0) and one far away at (70, 70), tunnel the update between them. The far end rebroadcasts it as if it were one hop from the base station, and 59 of the 97 other nodes route through the tunnel. Every one of those packets is at the attackers' mercy, yet the links the nodes see are all real: no link check catches this.
munotes.in268

Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood

The 10 by 10 grid with the base station in one corner, the wormhole's two ends joined by a dashed line, and the 59 nodes that route through it shaded

Figure 42.1 The wormhole run: every shaded node's data passes through the attackers' tunnel

The grid is idealised, but the lesson is Karlof and Wagner's: in a protocol that builds its tree simply by "the reception of a packet", "sinkholes are easy to create because there is no information for a defender to verify".

Countermeasures

Against outsiders: link-layer security. "The majority of outsider attacks against sensor network routing protocols can be prevented by simple link layer encryption and authentication using a globally shared key." An outsider cannot then join the topology, so the Sybil attack, most sinkholes and selective forwarding, and acknowledgement spoofing are stopped. Two attacks are not: wormholes and HELLO floods, which only relay or amplify legitimate packets, need no key.

Against the Sybil attack: verify identities. With a globally shared key, an insider can "masquerade as any (possibly even nonexistent) node". Karlof and Wagner propose that every node share a unique symmetric key with the trusted base station; two neighbours then verify each other's identity through a Needham-Schroeder-like protocol and set up a shared key. The base station can also limit how many verified neighbours a node may have, so a captured node cannot befriend the whole network.

Against the HELLO flood: check links both ways. A node acts on a message only after confirming the link works in both directions, as in the run above. The identity verification protocol does this too.

munotes.in269

Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood

Against wormholes and sinkholes: design them out. They "are very difficult to defend against, especially when the two are used in combination", and "it is extremely difficult to retrofit existing protocols", so "the best solution is to carefully design routing protocols in which wormholes and sinkholes are meaningless". Geographic routing is their example: traffic flows towards the base station's physical position, so it is hard to attract elsewhere, and a false link through a wormhole is exposed because the "neighbours" find they are "well beyond normal radio range" of each other ([Geographic Routing: Greedy Forwarding and GPSR]). Positions must still be trusted, though: a node that lies about its location can put itself on a path.

Against selective forwarding: many paths. Sending over multiple disjoint paths protects against a limited number of compromised nodes; braided paths, which share nodes but not links, and choosing the next hop at random from several candidates give probabilistic protection.

Authenticated broadcast. No node should be able to forge a base station's broadcast, yet every node must be able to check it. That asymmetry, built from symmetric keys, is SPINS's microTESLA, in [Keys and Link Security: Key Predistribution, SPINS and 802.15.4].

Distinctions

AttackWhat the attacker doesMain effectCountermeasure
Spoofed, altered, replayed routing informationForges or replays routing messagesLoops, attracted or repelled traffic, partitionsAuthentication
Selective forwardingDrops some or all packets it should forwardLost dataMultipath and braided routes
SinkholeAdvertises a very attractive routeA region's traffic passes through itHard; design routing where it is meaningless
SybilPresents many identitiesDefeats multipath and geographic routingIdentity verification via the base station
WormholeTunnels packets out of bandDistant nodes seem close; sinkholesHard; geographic routing exposes false links
HELLO floodBroadcasts with very high powerNodes adopt an unreachable neighbourBidirectional link verification
Acknowledgement spoofingForges link-layer acknowledgementsWeak or dead links seem goodAuthenticated acknowledgements
SinkholeBlack hole
AimAttract trafficDiscard traffic
MethodAdvertise an attractive routeDrop everything received
RelationA sinkhole makes a black hole, or selective forwarding, effectiveThe simplest selective forwarding

What it does not mean

Link-layer encryption does not stop wormholes or HELLO floods. Both work by relaying legitimate packets, which need no key.

A two-way link check does not stop sinkholes. In the run it ended the stranding but left 70 nodes routing through the attacker.

A wormhole is not a broken link. Every link the nodes see is real; the deception is in the tunnel they cannot see.

munotes.in270

Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood

Insiders are not stopped by keys. A captured node holds valid keys; the aim against insiders is graceful degradation, not immunity.

Quick revision

  • Karlof and Wagner's goals: integrity, authenticity, availability; against insiders, graceful degradation in proportion to the fraction compromised.
  • Seven attacks: spoofed, altered or replayed routing information; selective forwarding (black hole); sinkhole; Sybil; wormhole; HELLO flood; acknowledgement spoofing.
  • Sensor networks are prone to sinkholes because all traffic goes to one base station.
  • Sybil: one node, many identities; defeats multipath, storage and geographic routing.
  • Wormhole: tunnel over an out-of-band link; creates sinkholes; invisible to link checks.
  • HELLO flood: high-power single-hop broadcast; nodes adopt an out-of-range parent.
  • Countermeasures: link-layer encryption and authentication (outsiders; not wormholes or HELLO floods); identity verification with the base station (Sybil); bidirectional verification (HELLO flood); multipath and braided routes (selective forwarding); geographic routing (sinkholes and wormholes); authenticated broadcast.
  • Our run on 99 nodes: spoofed base station 44 captured; HELLO flood 92 stranded; with the two-way check 70 in a sinkhole; wormhole 59 through the tunnel.

Test yourself

1. Explain the sinkhole attack. Why are sensor networks particularly vulnerable to it? A compromised node makes itself look especially attractive to the routing protocol, for example by advertising a very high-quality route to the base station, so that its neighbours route through it and pass the attractive route on, and the traffic of a whole area converges on the attacker, which can then drop or alter it. Sensor networks are vulnerable because all traffic goes to the same base station, so a single attractive route to it influences a large number of nodes.

2. What is a Sybil attack? An attack in which one node presents multiple identities to the other nodes. It defeats schemes that rely on several independent nodes, such as multipath routing, distributed storage and topology maintenance, and it lets an attacker appear in several places at once in geographic routing. It is countered by verifying identities, for example through keys each node shares with a trusted base station.

3. Explain the wormhole attack. Two colluding attackers, in different parts of the network, tunnel packets between themselves over a fast out-of-band link and replay them at the other end. Nodes near one end then believe they are only a hop or two from nodes, or a base station, near the other, so routes are drawn through the tunnel, creating a sinkhole that the attackers control. Because the nodes see only real links, it is very hard to detect.

4. What is a HELLO flood attack, and how is it countered? A laptop-class attacker broadcasts a HELLO or routing message with enough power to reach every node, so that every node believes the attacker is its neighbour and may choose it as parent; distant nodes then send packets to a node that cannot hear them. It is countered by verifying that a link works in both directions before acting on a message received over it.

munotes.in271

Routing Attacks: Sinkhole, Sybil, Wormhole and HELLO Flood

5. Which attacks does link-layer encryption with a shared key stop, and which does it not? It stops most outsider attacks: an outsider cannot join the topology, so Sybil attacks, most sinkholes and selective forwarding, and spoofed acknowledgements are prevented. It does not stop wormholes or HELLO floods, which only relay or amplify legitimate packets, and it is useless against insiders who hold the shared key.

Contents This chapter on its own page

munotes.in272

Chapter Forty-Four

Privacy in Ad Hoc and Sensor Networks

Syllabus topic Module 1, "WSN Operating Systems and Ad-hoc Networks: Security and privacy issues in ad-hoc networks"

In one line

Encryption hides what a sensor network says but not where or when it says it: an eavesdropper who follows the radio traffic back hop by hop can find the animal, the soldier or the patient that caused it, so privacy needs routes that wander, fake traffic, delays, and aggregation that never reveals any single reading.

In the wording a student can write in an examination: privacy threats in sensor networks are of two kinds. Content privacy concerns the data in the packets, and is protected by encryption. Contextual privacy concerns what can be learnt without reading the data: the location of the source (where an event happened), the location of the sink or base station, and the time at which events happened. In the panda-hunter game, sensors report the location of pandas to a sink, and a hunter eavesdrops on the radio and traces the packets back, hop by hop, to the source and so to the panda. Defences include phantom routing (each packet first takes a random walk to a random "phantom" source before going to the sink), fake sources that send decoy traffic, and delays that hide timing. Their success is measured by the safety period, the time or number of packets before the adversary finds the source, and paid for in energy and latency. Privacy-preserving data aggregation computes sums and averages without revealing individual readings, for example by slicing each reading into random pieces shared among neighbours (SMART), or by additively homomorphic encryption, whose ciphertexts can be added without being decrypted.

Content and context

Encryption with the keys of [Keys and Link Security: Key Predistribution, SPINS and 802.15.4] hides what a packet says. But a sensor network leaks through its traffic itself. A node transmits because it sensed something; the transmission can be located with a directional antenna; and each hop can be traced back to the one before. Three kinds of context are at risk:

  1. Source location: where the event was, and so where the thing being sensed is.
  2. Sink location: where the base station is. Destroy it, and the whole network is silenced ([Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks]).
  3. Time: when an event happened, which, combined with location, says what was happening.

The panda and the hunter

"The source location privacy problem was introduced in [6] using the panda hunter game", Mutalemwa and Shin write: a network deployed "to continuously monitor activities and location of the pandas in their habitat". A node that senses a panda becomes a source and reports to the sink. That work "showed that using a directional antenna, an adversary could monitor the pattern of broadcasts between sensor nodes, identify the location of immediate sender node of the packet and, using this information, it could trace back the packet route to find the ultimate source of the packets and thus the panda". The same threat applies on a battlefield, where "soldiers may wear sensor nodes", and in any monitoring of something valuable.

munotes.in281

Privacy in Ad Hoc and Sensor Networks

Two kinds of local hunter. A patient adversary "patiently waits at a node until it hears a packet and moves to the immediate sender node of the packet", repeating until it reaches the source. A cautious adversary also sets a timer: if nothing arrives in time, it "will roll back to its previous node", and it remembers where it has been so as not to go round in circles. A global adversary, which can watch the whole network's traffic at once, is harder still.

The safety period is the measure of privacy: "the time required for an adversary to back trace and capture the asset", or, in the other definition, the longest time the asset stays in one place before moving on. The longer the adversary needs, the better the protection.

The defences

Shortest-path routing gives the least privacy. Every packet from the source takes the same route, so each packet moves the hunter one hop closer. Its routes "provide a short safety period and very poor privacy while consuming very low energy".

Fake sources. Other nodes act as sources too, sending "fake packets" that "are of the same length as the real packets, and they are encrypted", to lead the adversary away. Their "most significant limitation" is "very high energy consumption due to the high volume of packets" they need.

Phantom routing. Introduced to cut that cost: "Phantom routing is a two-phase routing scheme where packets are first forwarded to a random phantom source through random walk, and then, a succeeding flooding or single-path routing is used to forward the packets to sink node." Each packet seems to come from a different place, so the hunter's trail keeps leading to phantoms.

The game, run

The program places sensors on a 21 by 21 grid with the sink in the middle, at (10, 10), and a panda at (2, 2), 16 hops away. One packet a time step reports the panda. A patient hunter starts at the sink: whenever a packet passes through the node where it waits, it steps to the node that sent it. With shortest-path routing, and then with phantom routing whose random walk (never revisiting a node) lasts 5, 10 or 20 hops, the program counts how many packets the panda sends before the hunter reaches it, over 20 runs each, and how many hops each packet costs.

munotes.in282

Privacy in Ad Hoc and Sensor Networks

# The panda-hunter game. Sensors on a 21 by 21 grid report a panda at (2, 2)
# to the sink at (10, 10), one packet per time step. A patient hunter starts at
# the sink; whenever it overhears a packet, it steps to the node that sent it.
# Phantom routing sends each packet on a random walk of h hops first.
import random

N, SINK, PANDA = 21, (10, 10), (2, 2)

def neighbours(p):
    x, y = p
    return [(a, b) for a, b in ((x + 1, y), (x - 1, y), (x, y + 1), (x, y - 1))
            if 0 <= a < N and 0 <= b < N]

def to_sink(p):                          # the next hop on a shortest path to the sink
    x, y = p
    if x != SINK[0]:
        return (x + (1 if SINK[0] > x else -1), y)
    return (x, y + (1 if SINK[1] > y else -1))

def packet_path(h, rnd):
    path = [PANDA]
    while len(path) <= h:                # a random walk that does not revisit a node
        options = [q for q in neighbours(path[-1]) if q not in path]
        if not options:
            break
        path.append(rnd.choice(options))
    while path[-1] != SINK:
        path.append(to_sink(path[-1]))
    return path

def safety_period(h, seed, limit=2000):
    """Packets sent before the hunter reaches the panda, and their mean hops."""
    rnd, hunter, hops = random.Random(seed), SINK, 0
    for sent in range(1, limit + 1):
        path = packet_path(h, rnd)
        hops += len(path) - 1
        if hunter in path[1:]:
            hunter = path[path.index(hunter) - 1]   # step to the node it heard
        if hunter == PANDA:
            break
    return sent, hops / sent

for h in (0, 5, 10, 20):
    runs = [safety_period(h, seed) for seed in range(20)]
    periods = [p for p, _ in runs]
    print("walk %2d hops: the hunter needs %5.1f packets on average (%3d to %3d); "
          "%.1f hops a packet" % (h, sum(periods) / len(periods), min(periods),
                                 max(periods), sum(c for _, c in runs) / len(runs)))
walk  0 hops: the hunter needs  16.0 packets on average ( 16 to  16); 16.0 hops a packet
walk  5 hops: the hunter needs  70.1 packets on average ( 38 to 145); 20.5 hops a packet
walk 10 hops: the hunter needs  98.4 packets on average ( 32 to 306); 23.5 hops a packet
walk 20 hops: the hunter needs 150.2 packets on average ( 50 to 519); 30.2 hops a packet
The 21 by 21 grid with the sink in the middle and the panda near a corner: one solid shortest route, and three dashed routes that wander near the panda before joining it

Figure 44.1 The program's routes: every packet on the solid path, or each on its own after a random walk

Reading it.

  1. Shortest path: 16 packets, every time. The route never changes, so each packet brings the hunter exactly one hop closer, and after 16 packets it is at the panda.
  2. A walk of 5 hops: about 70 packets. Each packet enters the shortest-path tree at a different point, so the hunter often waits at nodes the next packets do not pass, or follows a trail that leads to a phantom. The safety period grows more than fourfold.
  3. Longer walks, longer safety, but not without limit. 10 hops give about 98 packets and 20 about 150, while the worst runs (32, 50) show a lucky hunter can still be quick: random routes delay the hunter; they do not guarantee anything.
  4. The price. Every packet travels 16 hops on the shortest path, but 20.5, 23.5 and 30.2 hops on average with the walks: nearly twice the energy and delay for the longest walk. Privacy in sensor networks is always bought with energy, as the fake-source schemes show more starkly.
munotes.in283

Privacy in Ad Hoc and Sensor Networks

Sink location and time

The sink. All traffic converges on the sink, so the pattern of traffic points to it, and a hunter who wants to disable the network rather than find one panda follows the traffic forwards. Abuzneid, Sobh and Faezipour name the goal alongside source location privacy (their "BLP", base station location privacy) and defend it, as for the source, with anonymity and traffic that does not give the destination away.

Time. Networks "could suffer from time correlation attacks" mounted "by observing the time between correlative packets sent and received in a certain neighborhood", which let an adversary "trace forward and backward the messages until they reach to the BS or to the source". "Hence, hiding temporal information is crucial for both anonymity and location privacy." The defences, they note, use "either using delays or fake messages": holding packets for random times before forwarding them, so that when a packet arrives says little about when its event happened, at the cost of latency.

Privacy in data aggregation

A network that averages readings on the way to the sink ([Design Principles: Distributed Organisation and In-network Processing]) exposes each reading to the aggregating node, unless the scheme prevents it. He and colleagues give the applications that make this matter: sensors in homes measuring water and electricity, whose readings "could reveal daily activities of a household, such as when all family members are gone or when someone is taking a shower", and health studies in which "individual's health data should be kept private".

CPDA (Cluster-based Private Data Aggregation) forms clusters and uses the algebra of polynomials so that a cluster can compute its sum while "no individual node can know the data values of other nodes", in Bista and Chang's summary.

SMART (Slice-Mix-AggRegaTe): "Each node i ... slices its private data d_i randomly into J pieces", keeps one, and sends the other J - 1, encrypted, to randomly chosen neighbours; every node then adds up what it holds and the partial sums are aggregated towards the sink as usual. The sum is unchanged, but no neighbour sees more than a random piece.

munotes.in284

Privacy in Ad Hoc and Sensor Networks

Additively homomorphic encryption. Castelluccia, Mykletun and Tsudik replace the exclusive-or of a stream cipher with addition: a reading m is encrypted as c = m + k (mod M), with a keystream k that each node shares only with the sink. Adding ciphertexts adds the readings: "c1 + c2 = Enc(m1 + m2, k1 + k2, M)". The aggregating nodes add ciphertexts they cannot read, and the sink subtracts the sum of the keys. M must exceed the largest possible sum, and the paper's rule is M = 2 to the power of the ceiling of log2(p × n), with p the largest reading and n the number of readings.

Both, run on four readings.

# Two ways to aggregate readings without revealing any single one.
import math
import random

readings = {"A": 23, "B": 31, "C": 19, "D": 27}  # four homes' private readings
print("true sum:", sum(readings.values()))

# 1. SMART (He and colleagues 2007): each node slices its reading into J random
#    pieces, keeps one, and sends the others to neighbours; every node then adds
#    what it holds. No single message, and no neighbour, sees a whole reading.
rnd, J = random.Random(4), 3
held = {n: 0 for n in readings}
for n, value in readings.items():
    others = rnd.sample([m for m in readings if m != n], J - 1)
    pieces = [rnd.randint(-50, 50) for _ in range(J - 1)]
    held[n] += value - sum(pieces)                   # the piece it keeps
    for m, piece in zip(others, pieces):
        held[m] += piece
        print("  %s sends %s a slice of %4d" % (n, m, piece))
print("SMART: the nodes' partial sums", held, "add up to", sum(held.values()))

# 2. Castelluccia, Mykletun and Tsudik (2005): c = m + k (mod M). Ciphertexts add
#    up to an encryption of the sum; the sink, knowing every key, subtracts them.
M = 2 ** math.ceil(math.log2(max(readings.values()) * len(readings)))  # the paper's rule
keys = {n: rnd.randrange(M) for n in readings}       # each shared with the sink only
cipher = {n: (readings[n] + keys[n]) % M for n in readings}
total = sum(cipher.values()) % M                     # what the aggregators compute
print("CMT: ciphertexts", cipher, "sum to", total,
      "; the sink decrypts", (total - sum(keys.values())) % M)
true sum: 100
  A sends B a slice of  -37
  A sends C a slice of   42
  B sends C a slice of  -31
  B sends D a slice of  -39
  C sends A a slice of    1
  C sends D a slice of   20
  D sends B a slice of  -22
  D sends A a slice of   16
SMART: the nodes' partial sums {'A': 35, 'B': 42, 'C': 9, 'D': 14} add up to 100
CMT: ciphertexts {'A': 115, 'B': 101, 'C': 63, 'D': 54} sum to 77 ; the sink decrypts 100
munotes.in285

Privacy in Ad Hoc and Sensor Networks

Reading it. In SMART, A's reading of 23 leaves A only as two random slices, -37 and 42, and the part it keeps; B, C and D each hold a jumble of their own and others' pieces, and their partial sums, 35, 42, 9 and 14, reveal no single reading, yet add to the true 100. In the encrypted scheme, the largest reading is 31 and there are 4, so p × n = 124 and M is 2 to the power 7, that is 128; the ciphertexts 115, 101, 63 and 54 sum, modulo 128, to 77, and the sink, subtracting the four keys, recovers 100. Neither scheme told any aggregating node what any home used.

Distinctions

Content privacyContextual privacy
What is protectedThe data in the packetsWhere and when packets come from and go
ThreatEavesdropping on the payloadTraffic analysis: tracing hops, timing, rates
DefenceEncryptionPhantom routing, fake sources, delays, anonymity
CostComputation and a few bytesEnergy and latency, often much more
Patient adversaryCautious adversary
BehaviourWaits at a node until it hears a packet, then moves to its senderWaits with a timer; rolls back if nothing comes; remembers visited nodes
WeaknessCan wait a long time at a node that random routes rarely useCan be led round by very random routes and roll back often
SMARTAdditively homomorphic encryption
IdeaSplit each reading into random slices shared with neighboursEncrypt with c = m + k (mod M); add ciphertexts
CostMore messages (the slices)Almost none beyond normal aggregation
Who can see a readingNobody without colluding with the neighboursOnly the sink, which holds the keys

What it does not mean

Encryption is not privacy. An encrypted packet still reveals that someone sent it, from where and when.

Random routes are not a guarantee. In the run, a lucky hunter still reached the panda in 32 packets against a 10-hop walk; phantom routing raises the average, not the minimum.

Privacy is not free. The longest walk nearly doubled the hops per packet, and fake sources cost far more.

Aggregation is not anonymity. An aggregate hides individual readings only if the scheme is designed to; in plain aggregation, every aggregating node reads its children's values.

Quick revision

  • Content privacy (encryption) against contextual privacy: source location, sink location, time.
  • Panda-hunter game: sensors report pandas; a hunter with a directional antenna traces packets back hop by hop to the source.
  • Adversaries: patient (wait, then step to the sender), cautious (timer, roll back, remember), global (sees all traffic).
  • Safety period: the time for the adversary to find the asset.
  • Defences: fake sources (decoy packets; costly), phantom routing (random walk to a phantom source, then flooding or single-path to the sink), delays and fake messages for temporal privacy.
  • Our run: shortest path 16 packets; phantom walk of 5 hops about 70, 10 about 98, 20 about 150; hops per packet from 16 to 30.2.
  • Private aggregation: CPDA (clusters, polynomials), SMART (slice into J pieces, keep one, send J - 1), additively homomorphic encryption c = m + k (mod M), M = 2 to the power of the ceiling of log2(p × n).
munotes.in286

Privacy in Ad Hoc and Sensor Networks

Test yourself

1. Distinguish content privacy and contextual privacy in sensor networks. Content privacy protects the data carried in packets, and is provided by encryption. Contextual privacy protects information that can be inferred without reading the packets, such as the location of the source of an event, the location of the sink, and the time at which events occur, from the pattern and timing of transmissions; encryption does not protect it.

2. Explain the panda-hunter game and phantom routing. Sensors monitor pandas and report sightings to a sink. A hunter eavesdrops with a directional antenna: each time it hears a packet it moves to the node that sent it, tracing the route back hop by hop until it finds the source and the panda. Phantom routing defends by sending each packet first on a random walk of several hops to a random phantom source, and only then along flooding or a single path to the sink, so that successive packets come from different places and the hunter's trail keeps leading away from the real source.

3. What is the safety period? The time, or number of packets, that an adversary needs to trace the traffic back and locate the asset; equivalently, the longest time the asset can stay in one place safely. It measures how much privacy a routing scheme provides.

4. How can data be aggregated without revealing individual readings? In SMART each node slices its reading into random pieces, keeps one and sends the others to neighbours; every node adds what it holds, and the partial sums are aggregated as usual, so the total is correct but no node sees another's reading. With additively homomorphic encryption, each reading is encrypted as the reading plus a key modulo M; aggregators add the ciphertexts without decrypting them, and the sink, which knows all the keys, subtracts their sum to obtain the total.

munotes.in287

Privacy in Ad Hoc and Sensor Networks

5. With readings of at most 40 from 10 nodes, what modulus does Castelluccia, Mykletun and Tsudik's rule give? p × n = 40 × 10 = 400, and log2(400) is between 8 and 9, so its ceiling is 9 and M is 2 to the power 9, that is 512.

Contents This chapter on its own page

munotes.in288

Chapter Forty-Five

MAC Protocols for Sensor Networks: The Job and Where the Energy Goes

Syllabus topic Module 1, "Medium Access Control (MAC) in WSN: Fundamentals of MAC protocols for sensor networks"

In one line

A MAC protocol decides when each node may use the radio channel it shares with its neighbours; in a sensor network its first duty is to save energy, because a radio left on to listen for traffic that never comes spends almost the whole battery doing nothing.

In the wording a student can write in an examination: the medium access control (MAC) protocol, part of the data link layer, controls access to the shared wireless channel so that two interfering nodes do not transmit at the same time. In a sensor network it has two goals: to create the network infrastructure (the links over which data travels hop by hop) and to share the channel fairly and efficiently among the nodes. Traditional wireless MACs aim at fairness, low latency, throughput and bandwidth utilisation; a sensor network's MAC puts energy efficiency first, then scalability and adaptability to changes in network size, density and topology. Energy is wasted in four ways: collisions (corrupted frames must be sent again), overhearing (receiving frames meant for other nodes), control packet overhead (energy spent on frames that carry no data) and idle listening (keeping the receiver on for traffic that is never sent), usually the largest. MAC protocols are contention-based (nodes compete for the channel: ALOHA, CSMA, S-MAC, B-MAC), schedule-based or contention-free (each node transmits in its own time slot, frequency or code: TDMA, FDMA, CDMA), or hybrid. Sensor MACs save energy above all by duty cycling: keeping the radio asleep most of the time.

What the MAC layer decides

The data link layer, in Akyildiz and colleagues' survey, "is responsible for the multiplexing of data streams, data frame detection, medium access and error control." The MAC is the medium access part, and its question is simple to state: when may this node transmit? Every node within range of a receiver shares the same channel, and Ye, Heidemann and Estrin name the consequence: "One fundamental task of the MAC protocol is to avoid collisions so that two interfering nodes do not transmit at the same time."

Four properties of radio make the question harder than on a wire, and each returns later in the block.

The channel is a broadcast one. A frame is heard by every node in range, wanted or not. That is how overhearing wastes energy, and how one node's transmission can collide with another's at a receiver both can reach.

A radio sends or receives, never both. The CC2420 takes 192 microseconds to turn around from receiving to transmitting ([The Radio, the Sensors and the Power Supply of a Node]). While it transmits it hears nothing, so a sender cannot notice a collision as it happens, as a wired Ethernet station can; it finds out only when no acknowledgement comes back.

munotes.in289

MAC Protocols for Sensor Networks: The Job and Where the Energy Goes

Collisions happen at the receiver. What matters is what the receiver hears, not what the sender hears. Two senders out of each other's range can both find the channel quiet and still collide at a receiver between them: the hidden terminal of [Hidden and Exposed Terminals, and RTS and CTS].

A sleeping radio cannot receive. A sensor MAC saves energy by switching the radio off, but a frame sent to a sleeping node is lost. The MAC must arrange that the receiver is awake when the sender sends, either by agreeing a schedule or by making the sender wait: the subject of chapters 50 to 53.

What a sensor network asks of its MAC

Ye, Heidemann and Estrin rank the attributes of a MAC protocol for a sensor network. The first is energy efficiency: nodes run on batteries that are "often very difficult to change or recharge". The second is scalability, meaning adaptability "to the change in network size, node density and topology", as nodes die, join and move. The rest are the goals of every other network: fairness, latency, throughput and bandwidth utilisation. "These attributes are generally the primary concerns in traditional wireless voice and data networks, but in sensor networks they are secondary."

Their reasons are the network's own. Fairness matters less because "all nodes cooperate for a single common task": a node with much to send should get the channel, even if a neighbour with little waits. Latency matters less while nothing is happening: "Sub-second latency is not important, and we can trade it off for energy savings."

Akyildiz and colleagues state the job as two goals. "The first is the creation of the network infrastructure": with thousands of nodes scattered in a field, the MAC must establish the links over which data travels hop by hop, which gives the network its self-organising ability. "The second objective is to fairly and efficiently share communication resources between sensor nodes."

Polastre, Hill and Culler, designing B-MAC from what deployments needed, list seven goals:

  • Low Power Operation
  • Effective Collision Avoidance
  • Simple Implementation, Small Code and RAM Size
  • Efficient Channel Utilization at Low and High Data Rates
  • Reconfigurable by Network Protocols
  • Tolerant to Changing RF/Networking Conditions
  • Scalable to Large Numbers of Nodes

The third and fifth are a sensor node's own: a MAC must fit in a few kilobytes of memory ([Inside a Sensor Node: The Five Units]), and a routing or application layer that knows the traffic should be able to tune it.

Why the existing MACs will not do. Akyildiz and colleagues take the obvious candidates in turn.

  • Cellular networks have base stations on a wired backbone with unlimited power, and each phone is one hop from one. Their MAC aims at quality of service and bandwidth efficiency, "Power conservation assumes only secondary importance", and access is a dedicated assignment made by the base station. A sensor network has no such central controller.
  • Bluetooth is a star: a master with up to seven slaves in a piconet, on a centrally assigned TDMA schedule with frequency hopping, at about 20 dBm and tens of metres.
  • MANETs must build and maintain their infrastructure under mobility, so their MAC aims at quality of service when nodes move; their batteries can be replaced by the user.
munotes.in290

MAC Protocols for Sensor Networks: The Job and Where the Energy Goes

A sensor network has many more nodes, transmits at about 0 dBm over shorter ranges, and changes its topology more often, through failures as much as movement. Because power conservation comes first, "none of the existing Bluetooth or MANET MAC protocols can be directly used."

The classes of MAC protocol

Contention-based (random access). A node transmits when it has something to send, and nodes compete for the channel. Collisions can happen, so the protocol needs rules to avoid and recover from them: carrier sensing, random backoff, a handshake, acknowledgements. ALOHA and CSMA ([Contention: ALOHA and CSMA]), IEEE 802.11 and 802.15.4's CSMA-CA, and the sensor MACs S-MAC and B-MAC are contention-based. Such protocols are simple and adapt easily when nodes come and go, but Dargie and Poellabauer name their energy problem: besides collisions and recovery, "sensor nodes may have to listen to the medium at all times to ensure that no transmissions will be missed."

Schedule-based (contention-free). Access is assigned: each node has its own time slot (TDMA), frequency (FDMA) or code (CDMA), so there are no collisions, and a node can switch its radio off outside its slots. Ye, Heidemann and Estrin: "TDMA protocols have a natural advantage of energy conservation compared to contention protocols, because the duty cycle of the radio is reduced and there is no contention-introduced overhead and collisions." The price is organisation: TDMA "usually requires the nodes to form real communication clusters", and when a cluster's membership changes it is hard to change the frame length and slot assignment, "So its scalability is normally not as good as that of a contention-based protocol." LEACH's clusters use TDMA inside each cluster ([LEACH: Clusters That Take Turns]); [TDMA and Schedule-based MAC] works a schedule out.

Demand assignment. Nodes ask for capacity (by polling or reservation) and are granted it. Akyildiz and colleagues set it aside: "Demand-based MAC schemes may be unsuitable for sensor networks due to their large messaging overhead and link setup delay."

Hybrid. Most real sensor MACs mix the two. S-MAC puts contention inside a shared schedule of listen and sleep, "a combined scheduling and contention scheme"; 802.15.4's superframe has a contention period and a period of guaranteed slots ([The 802.15.4 Superframe and Guaranteed Time Slots]).

munotes.in291

MAC Protocols for Sensor Networks: The Job and Where the Energy Goes

Protocols that sleep to save energy divide again. Synchronised protocols such as S-MAC and T-MAC "negotiate a schedule that specifies when nodes are awake and asleep within a frame", so neighbours wake together. Asynchronous protocols such as B-MAC keep no common schedule; each node wakes briefly to sample the channel, and a sender with data transmits a preamble long enough for the receiver to catch it. In the X-MAC paper's words, "Idle listening is reduced in asynchronous protocols by shifting the burden of synchronization to the sender."

The four sources of wasted energy

Ye, Heidemann and Estrin's list has become the standard answer. For each, what it is, and where the block meets its cure.

Four panels: in the first, A and C both send to B, which receives neither; in the second, A sends to B and C, nearby, receives the frame too; in the third, RTS, CTS and ACK frames surround one DATA frame; in the fourth, a long bar of radio-on time with nothing arriving

Figure 45.1 The four sources of wasted energy, in Ye, Heidemann and Estrin's order

1. Collision. "When a transmitted packet is corrupted it has to be discarded, and the follow-on retransmissions increase energy consumption. Collision increases latency as well." Both the lost frame and its repeat cost energy at both ends. Cures: listening before sending and random backoff ([Contention: ALOHA and CSMA], [CSMA/CA Worked Step by Step]), the RTS and CTS handshake ([Hidden and Exposed Terminals, and RTS and CTS]), or schedules with no contention at all.

2. Overhearing, "meaning that a node picks up packets that are destined to other nodes." In a dense network every frame reaches many nodes that only discard it. Cure: sleep while a neighbour talks to someone else, which S-MAC takes from PAMAS ([S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing]). The long preamble of low-power listening makes it worse, which X-MAC repairs ([Duty Cycling: Preamble Sampling, B-MAC and X-MAC]).

3. Control packet overhead. "Sending and receiving control packets consumes energy too, and less useful data packets can be transmitted." RTS, CTS, acknowledgements and schedule messages carry no data. Cures: few and short control frames, and one handshake for a whole burst of fragments (S-MAC's message passing).

4. Idle listening, "listening to receive possible traffic that is not sent." It is the largest because a sensor network is quiet most of the time, and an idle receiver costs half to all as much as a busy one: Stemm and Katz measured idle, receive and send power in the ratios 1 : 1.05 : 1.4 ([Energy Efficiency in Ad Hoc Networks: Where the Energy Goes]). Cure: switch the radio off. Every sensor MAC in this block is built round this one, by sleeping on a schedule ([S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step]), by sampling the channel briefly ([Duty Cycling: Preamble Sampling, B-MAC and X-MAC]), or by owning a slot ([TDMA and Schedule-based MAC]).

munotes.in292

MAC Protocols for Sensor Networks: The Job and Where the Energy Goes

One node's hour, charged

How large is each waste? The program follows one relay node for an hour. Its radio is a CC2420: 18.8 mA to receive or listen, 17.4 mA to transmit, 0.02 mA powered down, 250 kbit/s on air. The traffic is illustrative: once a minute the node sends three data frames (its own reading and its two children's) and receives two (from the children), each acknowledged, and hears ten exchanges between neighbours that are not for it. One frame in 20 that it sends collides and is sent again, after listening for an acknowledgement that does not come. A data frame occupies 50 bytes on air and an acknowledgement 11.

Every second the radio is on and doing none of these things is idle listening. The program charges each category with current × time, in milliampere-seconds (mA-s), three ways: quiet, with the radio always on; during an event with sixty times the traffic, still always on; and quiet with the radio on only 1 per cent of the time, if the neighbours keep the same schedule, so that everything still happens while the node is awake.

# Where one node's radio energy goes in an hour. The currents are the CC2420's;
# the traffic is illustrative: a relay sends its own reading and its two
# children's once a minute, receives the children's, and hears ten exchanges a
# minute between neighbours that are not for it. Every frame is acknowledged.
RX, TX, SLEEP = 18.8, 17.4, 0.02     # mA: receive (or listen), transmit, power down
RATE = 250_000                       # bits per second
DATA, ACK = 50, 11                   # bytes on air: a data frame, an acknowledgement
LOST = 0.05                          # one frame in 20 collides and is sent again

def air(nbytes):                     # seconds a frame occupies the channel
    return nbytes * 8 / RATE

def budget(per_minute, awake):
    """Charge in mA-s over one hour, by where it goes. per_minute scales the
    traffic; awake is the fraction of the hour the radio is on."""
    sent, received = 3 * per_minute * 60, 2 * per_minute * 60
    overheard = 10 * per_minute * 60
    time = {                         # seconds (transmitting, receiving)
        "data": (sent * air(DATA), received * air(DATA)),
        "control": (received * air(ACK), sent * air(ACK)),
        "collisions": (LOST * sent * air(DATA), LOST * sent * air(ACK)),
        "overhearing": (0, overheard * (air(DATA) + air(ACK))),
    }
    on = 3600 * awake
    busy = sum(t + r for t, r in time.values())
    time["idle listening"] = (0, on - busy)
    charge = {k: t * TX + r * RX for k, (t, r) in time.items()}
    if awake < 1:
        charge["sleep"] = (3600 - on) * SLEEP
    return charge

for name, per_minute, awake in (("quiet, radio always on", 1, 1.0),
                                ("an event, sixty times the traffic, always on", 60, 1.0),
                                ("quiet, radio on 1 per cent of the time", 1, 0.01)):
    charge = budget(per_minute, awake)
    total = sum(charge.values())
    print(name)
    for k, q in sorted(charge.items(), key=lambda kv: -kv[1]):
        print("  %-15s %9.1f mA-s %7.2f %% %9.2f times the data"
              % (k, q, 100 * q / total, q / charge["data"]))
    print("  total %.1f mA-s; average %.3f mA; 2,500 mAh lasts %.1f days"
          % (total, total / 3600, 2500 / (total / 3600) / 24))
munotes.in293

MAC Protocols for Sensor Networks: The Job and Where the Energy Goes

quiet, radio always on
  idle listening    67646.6 mA-s   99.95 %   7846.91 times the data
  overhearing          22.0 mA-s    0.03 %      2.55 times the data
  data                  8.6 mA-s    0.01 %      1.00 times the data
  control               1.9 mA-s    0.00 %      0.22 times the data
  collisions            0.3 mA-s    0.00 %      0.04 times the data
  total 67679.5 mA-s; average 18.800 mA; 2,500 mAh lasts 5.5 days
an event, sixty times the traffic, always on
  idle listening    65678.5 mA-s   97.08 %    126.98 times the data
  overhearing        1321.1 mA-s    1.95 %      2.55 times the data
  data                517.2 mA-s    0.76 %      1.00 times the data
  control             115.6 mA-s    0.17 %      0.22 times the data
  collisions           18.6 mA-s    0.03 %      0.04 times the data
  total 67651.1 mA-s; average 18.792 mA; 2,500 mAh lasts 5.5 days
quiet, radio on 1 per cent of the time
  idle listening      643.4 mA-s   86.07 %     74.64 times the data
  sleep                71.3 mA-s    9.53 %      8.27 times the data
  overhearing          22.0 mA-s    2.95 %      2.55 times the data
  data                  8.6 mA-s    1.15 %      1.00 times the data
  control               1.9 mA-s    0.26 %      0.22 times the data
  collisions            0.3 mA-s    0.04 %      0.04 times the data
  total 747.6 mA-s; average 0.208 mA; 2,500 mAh lasts 501.6 days

Reading it. Three findings, each of which the rest of the block builds on.

With the radio always on, idle listening is nearly everything. In the quiet hour it takes 99.95 per cent of the charge, about 7,847 times what moving the data costs. The node averages 18.800 mA and a 2,500 mAh battery lasts 5.5 days, whatever the MAC does about collisions or control frames: those two together cost about a quarter of what the useful data costs (0.22 and 0.04 times it). (With the processor's 0.5 mA added, [How Long a Node Lasts: The Energy Budget Worked Out] found 5.4 days.)

With the radio always on, traffic hardly changes the bill. Sixty times the traffic cost 67,651.1 mA-s against 67,679.5, slightly less, because the CC2420 draws less to transmit (17.4 mA) than to listen (18.8 mA). Energy is set by how long the radio is on, not by how much it carries. Even during the event, idle listening is 97.08 per cent.

munotes.in294

MAC Protocols for Sensor Networks: The Job and Where the Energy Goes

Once the radio sleeps, the next waste shows. On 1 per cent of the time, the hour costs 747.6 mA-s, about 90 times less, and the battery lasts 501.6 days. Idle listening is still the largest item, but overhearing, at 22.0 mA-s, now costs 2.55 times the useful data, and the sleep current itself is 9.53 per cent. That is the order in which S-MAC attacks the problem: sleep first, then sleep through the neighbours' conversations ([S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing]).

What the table does not show is the price of sleeping: a frame that arrives while its receiver sleeps must wait for the next wake-up, at every hop. [S-MAC: Latency, Adaptive Listening and the Energy Saved] counts it.

Distinctions

Contention-basedSchedule-based
AccessCompete when there is dataOwn slot, frequency or code
CollisionsPossible; avoided and recovered fromNone within the schedule
ListeningMay have to listen all the timeRadio off outside own slots
Changes in the networkAdapts easilyFrame and slots must be reassigned
ExamplesALOHA, CSMA, 802.11, S-MAC, B-MACTDMA in LEACH's clusters, Bluetooth's piconet
Idle listeningOverhearing
On the airNothingA frame for another node
The radioOn, receiving silenceOn, receiving and then discarding
CureSleep when no traffic is expectedSleep while neighbours talk to others
Synchronised duty cyclingAsynchronous duty cycling
IdeaNeighbours agree when to be awakeEach node samples the channel briefly on its own
Who paysEveryone, to keep schedules in stepThe sender, with a long preamble
ExamplesS-MAC, T-MACB-MAC, X-MAC

What it does not mean

Energy first does not mean fairness and latency do not matter. They are traded, not abandoned; Ye, Heidemann and Estrin argue that a loss per hop need not mean a loss end to end, and the application still sets a limit on delay.

Less traffic does not mean less energy. With the receiver always on, the quiet hour cost slightly more than the busy one. Only turning the radio off saves.

Schedule-based is not free of overhead. It has no collisions, but keeping clocks and slots in step costs messages and guard times ([TDMA and Schedule-based MAC]).

Idle listening is not the same as receiving nothing useful. Overhearing receives a real frame that is not for the node; idle listening receives nothing at all.

Quick revision

  • The MAC decides when a node may transmit on a shared channel; its basic task is to avoid collisions between interfering nodes.
  • Radio makes it hard: a broadcast channel, a half-duplex radio (192 microseconds to turn around on the CC2420), collisions at the receiver, and sleeping receivers.
  • Sensor MAC priorities (Ye, Heidemann and Estrin): energy efficiency, then scalability and adaptability; fairness, latency, throughput and bandwidth utilisation secondary.
  • Two goals (Akyildiz and colleagues): create the network infrastructure; share the channel fairly and efficiently.
  • B-MAC's seven goals: low power, collision avoidance, simple and small, efficient at low and high rates, reconfigurable, tolerant of changing conditions, scalable.
  • Classes: contention-based, schedule-based (TDMA, FDMA, CDMA), demand assignment (unsuitable: overhead and setup delay), hybrid; duty cycling synchronised (S-MAC, T-MAC) or asynchronous (B-MAC, X-MAC).
  • Four wastes: collision, overhearing, control packet overhead, idle listening.
  • The program: always on, idle listening 99.95 per cent, 2,500 mAh lasts 5.5 days; sixty times the traffic costs no more; awake 1 per cent, about 90 times less and 501.6 days, with overhearing then 2.55 times the useful data.
munotes.in295

MAC Protocols for Sensor Networks: The Job and Where the Energy Goes

Test yourself

1. What is the job of a MAC protocol, and what are its two goals in a sensor network? The MAC protocol controls access to the shared wireless channel, deciding when each node may transmit, so that interfering nodes do not transmit at the same time. In a sensor network it must create the network infrastructure, establishing the links over which data travels hop by hop, and share the communication resources fairly and efficiently among the nodes.

2. How do the requirements on a sensor network MAC differ from those on a traditional wireless MAC? A traditional MAC aims at fairness, low latency, high throughput and bandwidth utilisation. A sensor network MAC puts energy efficiency first, because batteries are hard to change, then scalability and adaptability to changes in network size, density and topology. Fairness matters less because all nodes serve one application, and latency can often be traded for energy while nothing is happening.

3. List and explain the four sources of energy waste at the MAC layer. Collision: corrupted frames are discarded and sent again, costing energy and delay. Overhearing: a node receives frames addressed to other nodes. Control packet overhead: energy spent sending and receiving frames such as RTS, CTS and acknowledgements that carry no data. Idle listening: the receiver is kept on to listen for traffic that is not sent; because networks are mostly quiet and listening costs nearly as much as receiving, it is usually the largest.

4. Compare contention-based and schedule-based MAC protocols. In contention-based protocols nodes transmit when they have data and compete for the channel; they are simple and adapt to changes, but suffer collisions and may have to listen all the time. In schedule-based protocols each node has its own slot, frequency or code; there are no collisions and radios can sleep outside their slots, but the schedule must be set up and kept, clusters must be formed, and scalability is poorer.

munotes.in296

MAC Protocols for Sensor Networks: The Job and Where the Energy Goes

5. A radio draws 18.8 mA listening and 0.02 mA asleep. If it listens 1 per cent of an hour and sleeps the rest, what charge does it use, ignoring its traffic? Awake 36 s: 36 × 18.8 = 676.8 mA-s. Asleep 3,564 s: 3,564 × 0.02 = 71.28 mA-s. Total 676.8 + 71.28 = 748.08 mA-s, against 3,600 × 18.8 = 67,680 mA-s always on, about 90 times less.

6. Why can a wireless sender not detect a collision while transmitting? Because a radio either transmits or receives, not both at once: while it sends it cannot hear another signal, and the collision happens at the receiver, which may be in range of a sender it cannot hear. So wireless MACs avoid collisions before sending and detect failure afterwards, by a missing acknowledgement.

Contents This chapter on its own page

munotes.in297

Chapter Forty-Six

Contention: ALOHA and CSMA

Syllabus topic Module 1, "Medium Access Control (MAC) in WSN: Fundamentals of MAC protocols for sensor networks"

In one line

In ALOHA a node sends whenever it has a frame and sends again if no acknowledgement comes, which wastes at least four-fifths of the channel; carrier sensing, listening before sending, recovers most of it, provided the delay before a transmission can be heard is short compared with a frame.

In the wording a student can write in an examination: in pure ALOHA a station transmits a frame as soon as it has one; if two frames overlap, both are lost, and each sender retransmits after a random delay when no acknowledgement arrives. A frame is vulnerable to any other frame started within one frame time before or after it, so with an offered traffic of G frames per frame time the throughput is S = G e to the power -2G, at most 1/(2e), about 0.184, at G = 0.5. Slotted ALOHA divides time into slots of one frame and lets stations start only at the beginning of a slot, which halves the vulnerable period: S = G e to the power -G, at most 1/e, about 0.368, at G = 1. CSMA (carrier sense multiple access) listens to the channel before sending, and is defined by what a station does when the channel is busy: 1-persistent waits until it is idle and then sends at once; non-persistent waits a random time and senses again; p-persistent (in slots) sends with probability p when the channel is idle and otherwise waits a slot. With a propagation delay of a hundredth of a frame, CSMA's capacity rises to 0.529 (1-persistent) and 0.815 (non-persistent). In energy terms ALOHA wastes whole frames in collisions, while CSMA spends a short listen before each attempt; non-persistent sensing lets the radio sleep while it waits.

ALOHA: send, and send again if nobody answers

The first packet radio network was built at the University of Hawaii, where Norman Abramson's 1970 report describes consoles sending to the central computer, the MENEHUNE, on one shared radio channel of 24,000 baud. Each packet was 704 bits (80 characters, 32 bits of identification and control and 32 parity bits) and lasted 29 milliseconds, about 34 with the receiver's synchronisation.

The access rule was as simple as a rule can be. "Each user at a console transmits packets to the MENEHUNE over the same high data rate channel in a completely unsynchronized (from one user to another) manner." The central station answered: "If and only if a packet is received without error it is acknowledged by the MENEHUNE." A console that heard no acknowledgement in time sent the packet again. Kleinrock and Tobagi add the essential detail: a station that retransmitted at once would collide again with the station it had collided with, so each waits a random time before retrying.

munotes.in298

Contention: ALOHA and CSMA

The vulnerable period. A frame lasting one frame time, starting at time t, is destroyed by any other frame that starts in the frame time before t (whose tail overlaps it) or in the frame time after t (whose head overlaps it). The vulnerable period is therefore two frame times.

The formula. Let G be the offered traffic: the average number of frames, new and retransmitted, that stations start per frame time. The analyses assume the start times form a Poisson stream, whose defining property is that the probability of no start at all in an interval of length x frame times is e to the power -Gx. A frame survives if nothing else starts in its two-frame window, with probability e to the power -2G, and the throughput S (frames delivered per frame time) is the traffic times the fraction that survives:

S = G e to the power -2G.

This is Abramson's equation 6, written in his report with the packet duration as the time unit. It rises, peaks and falls. Setting its slope to zero gives G = 0.5, where S = 0.5 × e to the power -1 = 1/(2e), about 0.184: at best, 18.4 per cent of the channel carries frames that arrive. In Abramson's words, the channel capacity "is reduced to roughly one sixth of its value if we were able to fill the channel with a continuous stream of uninterrupted data."

Why it collapses. Past G = 0.5, more attempts deliver fewer frames. The frames that fail are retried, raising G further, which delivers still fewer: Abramson notes that at this point the channel becomes unstable and the number of retransmissions grows without limit. An ALOHA channel must be run well below its peak.

A slip in the original. Abramson's report gives the maximum, 1/2e, as 0.186, in the text and on the axis of its figure 3 (page 13, read from the page image). The value of 1/(2e) is 0.184 (e is about 2.718, so 2e is about 5.437, and 1 / 5.437 is about 0.184), which is what Kleinrock and Tobagi and every later analysis give. A book that says 0.186 has copied the slip.

Worked example: Abramson's own sum. How many active consoles could the channel carry? If each sends a packet every 30 seconds on average, and a packet occupies 0.034 s, the traffic per console is 0.034 / 30, about 0.001133 packet times per second of channel. The channel carries at most 1/(2e) of its time, so the number of consoles is 1 / (2e × 0.001133), and 2e × 0.001133 is about 0.00616; 1 / 0.00616 is about 162, the figure the report gives. Consoles that are idle use no capacity at all, so many more than 162 could be connected.

munotes.in299

Contention: ALOHA and CSMA

Slotted ALOHA

The improvement came from L. G. Roberts, whose result Kleinrock and Tobagi report: divide time into slots exactly one frame long, and let a station start only at the beginning of a slot. Two frames that collide then overlap completely, and a frame is vulnerable only to others started in its own slot: one frame time instead of two. The survival probability becomes e to the power -G, and

S = G e to the power -G, at most 1/e, about 0.368, at G = 1.

Slotting doubles the capacity. Its price is a common clock: every station must know where the slots begin, which Bonaventure notes "can be imposed by a single clock that is received by all terminals."

Slotted ALOHA is not history. The first step of getting onto a cellular network is a random access of this kind, and the specifications of Module II say so: GPRS, "The access to the GPRS uplink uses a Slotted-Aloha based reservation protocol"; UMTS, "The random-access transmission is based on a Slotted ALOHA approach with fast acquisition indication"; TETRA, "The random access protocol is based on slotted ALOHA procedures". A phone asking for a channel sends a short burst in a random access slot and hopes it is alone.

CSMA: listen before sending

In ALOHA a station sends without regard to whether anyone else is sending. Carrier sense multiple access listens first. In Bonaventure's words, "CSMA requires all nodes to listen to the transmission channel to verify that it is free before transmitting a frame". A station that hears another's carrier holds back, and collisions can happen only when two stations start so close together that neither has yet heard the other.

That closeness is measured by a: the delay before a transmission can be sensed by the others, as a fraction of a frame's transmission time. Kleinrock and Tobagi take it as the propagation delay, the same for every pair of stations. A frame started at time t can be hit only by a frame another station starts before it has heard this one, in the window from t to t + a (or one started up to a before t, which this station could not yet hear). The vulnerable period shrinks from two frame times to about a, and that is the whole of CSMA's gain.

What a station does when it hears the channel busy defines three protocols, which Kleinrock and Tobagi set out and analyse:

RuleChannel idleChannel busy
1-persistentTransmitKeep listening until it goes idle, then transmit at once (with probability one)
Non-persistentTransmitGive up for now; sense again after a random delay
p-persistent (slotted, slot = a)Transmit with probability p; with 1 - p wait one slot and sense againWait until idle, then apply the rule
munotes.in300

Contention: ALOHA and CSMA

1-persistent never lets the channel go idle while someone is ready, which suits light traffic. Its flaw is certain collision under load: every station that became ready during a transmission waits for its end, and they all start together. Non-persistent breaks that synchrony by making each waiting station return at its own random time, at the price of idle gaps the channel could have used. p-persistent sits between them: the stations waiting at the end of a transmission each start with only a small probability p per slot, so they are spread out; Kleinrock and Tobagi found the best p for a = 0.01 to be about 0.03.

Kleinrock and Tobagi derive the throughput of each. For non-persistent CSMA it is short enough to write:

S = G e to the power -aG / (G(1 + 2a) + e to the power -aG).

With a = 0 it becomes G / (1 + G), which approaches 1 as G grows: with instant sensing, a perfectly used channel. The 1-persistent formula is longer, and is in the program below.

The curves, computed

The program evaluates all four formulas over a range of offered traffic, for a = 0.01, and then finds each one's maximum, the capacity, at a = 0.01 and at a = 0.12. (Why 0.12 is the section after.)

# Throughput S (frames delivered per frame time) against offered traffic G
# (frames attempted per frame time), from the formulas of Abramson (pure ALOHA),
# Roberts (slotted ALOHA) and Kleinrock and Tobagi (CSMA). a is the delay before
# a transmission can be sensed, as a fraction of a frame's transmission time.
from math import exp

def pure(G, a):     return G * exp(-2 * G)
def slotted(G, a):  return G * exp(-G)
def nonpersistent(G, a):
    return G * exp(-a * G) / (G * (1 + 2 * a) + exp(-a * G))
def one_persistent(G, a):
    top = G * (1 + G + a * G * (1 + G + a * G / 2)) * exp(-G * (1 + 2 * a))
    return top / (G * (1 + 2 * a) - (1 - exp(-a * G)) + (1 + a * G) * exp(-G * (1 + a)))

MODES = [("pure ALOHA", pure), ("slotted ALOHA", slotted),
         ("1-persistent CSMA", one_persistent), ("non-persistent CSMA", nonpersistent)]

print("a = 0.01")
print("%5s" % "G" + "".join("%9s" % h for h in ("pure", "slotted", "1-pers", "non-pers")))
for G in (0.25, 0.5, 1, 2, 5, 10, 20):
    print("%5g" % G + "".join("%9.3f" % f(G, 0.01) for _, f in MODES))

loads = [10 ** (k / 1000) for k in range(-2000, 3001)]    # G from 0.01 to 1,000
for a in (0.01, 0.12):
    print("capacity, a = %g" % a)
    for name, f in MODES:
        S, G = max((f(G, a), G) for G in loads)
        print("  %-20s S = %.3f at G = %.2f" % (name, S, G))
munotes.in301

Contention: ALOHA and CSMA

a = 0.01
    G     pure  slotted   1-pers non-pers
 0.25    0.152    0.195    0.235    0.199
  0.5    0.184    0.303    0.407    0.331
    1    0.135    0.368    0.529    0.493
    2    0.037    0.271    0.369    0.649
    5    0.000    0.034    0.038    0.786
   10    0.000    0.000    0.000    0.815
   20    0.000    0.000    0.000    0.772
capacity, a = 0.01
  pure ALOHA           S = 0.184 at G = 0.50
  slotted ALOHA        S = 0.368 at G = 1.00
  1-persistent CSMA    S = 0.529 at G = 1.02
  non-persistent CSMA  S = 0.815 at G = 9.44
capacity, a = 0.12
  pure ALOHA           S = 0.184 at G = 0.50
  slotted ALOHA        S = 0.368 at G = 1.00
  1-persistent CSMA    S = 0.439 at G = 0.90
  non-persistent CSMA  S = 0.483 at G = 2.26
Throughput against offered traffic on a logarithmic axis from 0.1 to 100: pure ALOHA peaks at 0.184, slotted ALOHA at 0.368, 1-persistent CSMA at 0.529 near G of 1, and non-persistent CSMA rises slowly to 0.815 near G of 10

Figure 46.1 The program's formulas drawn, for a = 0.01

Reading it. The capacities at a = 0.01 are Kleinrock and Tobagi's: pure ALOHA 0.184 at G = 0.5, slotted ALOHA 0.368 at G = 1, 1-persistent CSMA 0.529 and non-persistent CSMA 0.815, the same four figures Bonaventure quotes from them.

No protocol wins at every load. At G = 0.5 the 1-persistent channel delivers 0.407 against the non-persistent 0.331: when traffic is light, waiting a random time wastes channel that 1-persistent would have used. From G = 2 upwards the order reverses (0.369 against 0.649), because 1-persistent's waiting stations collide at the end of every transmission. Non-persistent CSMA reaches its capacity only at an offered traffic near 9.4, where most attempts hear the channel busy and back off.

Carrier sensing depends on a; ALOHA does not. At a = 0.12 the two ALOHA capacities are unchanged and the CSMA ones fall to 0.439 and 0.483, not far above slotted ALOHA. Kleinrock and Tobagi make the same point, and add that for large enough a slotted ALOHA beats every CSMA mode, because decisions based on partly out-of-date knowledge of the channel do harm.

A sensor radio's a

For the packet radios of the 1970s, a was the propagation delay, and small: Kleinrock and Tobagi estimate about 0.005 of a packet's time for their system. Radio travels 300 metres in a microsecond, so across a sensor network's tens of metres it is smaller still.

munotes.in302

Contention: ALOHA and CSMA

For a sensor radio the delay that matters is the radio's own. After a clear channel assessment says the channel is free, the radio must switch from receiving to transmitting before its frame goes out, and until it does, nobody can hear it. IEEE 802.15.4 allows up to aTurnaroundTime, 12 symbol periods, for that switch; at 62.5 ksymbol/s a symbol lasts 16 microseconds, so the turnaround is up to 12 × 16 = 192 microseconds, the CC2420's figure. The assessment itself listens for 8 symbol periods, 128 microseconds. A neighbour that assesses the channel inside that window hears nothing and sends too.

Worked example. The 50-byte frame of [MAC Protocols for Sensor Networks: The Job and Where the Energy Goes] takes 400 bits / 250,000 bits per second = 1.6 ms. Counting the turnaround alone, a = 0.192 / 1.6 = 0.12, which is why the program computed the capacity at 0.12. Counting the assessment too, a = (0.128 + 0.192) / 1.6 = 0.2. A short frame on a sensor radio has an a ten or twenty times that of the 1970s packet radios, and CSMA's advantage over ALOHA shrinks accordingly. Longer frames help: the same delays are a smaller fraction of a longer frame.

The channel simulated, and the energy counted

A formula rests on its assumptions, so the second program builds the channel itself. Attempts arrive at random, as a Poisson stream of G per frame time. For pure ALOHA a frame succeeds if no other starts within one frame time either side; for slotted ALOHA, if it is alone in its slot. For non-persistent CSMA it follows the busy periods: an attempt within a of the start of a transmission cannot hear it, sends and collides; one later in the busy period hears it and gives up; one after the channel is heard idle starts a new busy period, which delivers a frame only if nobody else joins it.

It also counts, per delivered frame, the two things a radio spends energy on here: transmissions (each a frame's worth of transmit current) and channel senses (each a short listen, 128 microseconds on 802.15.4).

# The same channels simulated. Attempts arrive at random (a Poisson stream of G
# per frame time, new frames and retries together, as in the analyses); every
# frame lasts 1. Counted per delivered frame: transmissions (the energy spent
# sending) and channel senses (the short listens before sending).
import random
from math import exp

def aloha(G, T, rnd, slotted):
    t, starts = 0.0, []
    while t < T:
        t += rnd.expovariate(G)
        starts.append(int(t) + 1 if slotted else t)   # slotted: wait for the next slot
    ok = sum(1 for i in range(1, len(starts) - 1)
             if starts[i] - starts[i - 1] >= 1 and starts[i + 1] - starts[i] >= 1)
    return ok, len(starts), 0

def nonpersistent(G, a, T, rnd):
    t, ok, sent, senses = 0.0, 0, 0, 0
    first = last = -9.0          # the current busy period's first and last starts
    n = 0                        # and how many transmissions it holds
    while t < T:
        t += rnd.expovariate(G)
        senses += 1
        if t < first + a:        # not heard yet: this station sends, and collides
            last, n, sent = t, n + 1, sent + 1
        elif t < last + 1 + a:   # heard busy: give up, retry later
            continue
        else:                    # heard idle: a new busy period begins
            if n == 1:
                ok += 1          # the previous one held a single frame: delivered
            first, last, n, sent = t, t, 1, sent + 1
    return ok, sent, senses

def formula(mode, G, a):
    if mode == "pure":
        return G * exp(-2 * G)
    if mode == "slotted":
        return G * exp(-G)
    return G * exp(-a * G) / (G * (1 + 2 * a) + exp(-a * G))

rnd, T = random.Random(46), 100_000
print("%-21s %5s %6s %8s %6s %7s" % ("", "G", "S sim", "formula", "sent", "senses"))
for mode, G, a in (("pure", 0.05, 0), ("pure", 0.5, 0), ("slotted", 1.0, 0),
                   ("non-persistent", 1.0, 0.01), ("non-persistent", 9.44, 0.01),
                   ("non-persistent", 2.26, 0.12)):
    if mode == "non-persistent":
        ok, sent, senses = nonpersistent(G, a, T, rnd)
    else:
        ok, sent, senses = aloha(G, T, rnd, mode == "slotted")
    label = mode + (" a=%g" % a if a else "")
    print("%-21s %5.2f %6.3f %8.3f %6.2f %7.2f"
          % (label, G, ok / T, formula(mode, G, a), sent / ok, senses / ok))
print("(sent and senses are counted per delivered frame)")
munotes.in303

Contention: ALOHA and CSMA

                          G  S sim  formula   sent  senses
pure                   0.05  0.046    0.045   1.10    0.00
pure                   0.50  0.184    0.184   2.72    0.00
slotted                1.00  0.368    0.368   2.71    0.00
non-persistent a=0.01  1.00  0.491    0.493   1.02    2.03
non-persistent a=0.01  9.44  0.815    0.815   1.20   11.58
non-persistent a=0.12  2.26  0.483    0.483   1.67    4.67
(sent and senses are counted per delivered frame)

Reading it. Over 100,000 frame times the simulated throughput agrees with every formula to within a few thousandths, so the formulas can be trusted within their assumptions.

ALOHA's price is paid in whole frames. At its best load each delivered frame costs about 2.7 transmissions, e (about 2.718) in theory, since G / S = e to the power 2G = e at G = 0.5, and likewise for slotted ALOHA at G = 1. The extra 1.7 frames of transmit energy per delivered frame are pure waste.

CSMA trades frames for listens. At its best load, non-persistent CSMA with a = 0.01 sends about 1.2 frames per delivered frame and listens about 11.6 times. A listen costs 128 microseconds of receive current against 1.6 ms of transmit current for a 50-byte frame, so ten listens cost about as much as one frame: 10 × 0.128 = 1.28 ms at 18.8 mA against 1.6 ms at 17.4 mA. With the sensor radio's a of 0.12, collisions return (1.67 frames per delivered frame) and CSMA's saving over ALOHA's 2.7 narrows.

munotes.in304

Contention: ALOHA and CSMA

At light load, nothing much happens. At G = 0.05, a sensor network's usual state, even pure ALOHA delivers a frame for every 1.1 or so transmissions (e to the power 2 × 0.05 = e to the power 0.1, about 1.105). The protocols differ when many nodes send at once, which in a sensor network is exactly when an event happens and the data matters most.

What each rule costs a sensor node

The simulation counts transmissions and short listens. What it cannot count is the listening a rule demands while the node waits, and there the rules differ sharply.

ALOHA never listens before sending, but its node must still have its receiver on to hear the acknowledgement, and to receive anything at all; and every collision wastes a full frame.

1-persistent CSMA is the hungriest. A station that finds the channel busy waits for it to go idle, and to know when that happens it must keep its receiver on for the rest of the other station's frame. Bonaventure's description of persistent CSMA says it plainly: "the terminal will continuously listen to the channel and transmit its frame as soon as the channel becomes free."

Non-persistent CSMA suits a battery. "A non-persistent CSMA node does not continuously listen to the channel to determine when it becomes free": between one short sense and the next, its radio can sleep.

p-persistent CSMA listens slot by slot until it sends or the channel is taken.

Sensor MACs are built on the non-persistent pattern. B-MAC waits an initial random backoff, then assesses the channel; "If the channel is not clear, an event signals the service for a congestion backoff time", and the node tries again later. IEEE 802.15.4's CSMA-CA does the same with a backoff that grows after each busy assessment ([CSMA/CA Worked Step by Step], [CSMA-CA, Data Transfer and Frames in 802.15.4]). The history agrees: the DARPA packet radio network ran slotted ALOHA and CSMA, and Polastre, Hill and Culler record that "TDMA and slotted ALOHA solutions in PRNET were ultimately dismissed due to their inability to scale."

Carrier sensing has one more weakness, which no persistence rule cures: the sender listens where it is, but collisions happen where the receiver is. [Hidden and Exposed Terminals, and RTS and CTS] takes that up.

munotes.in305

Contention: ALOHA and CSMA

Distinctions

Pure ALOHASlotted ALOHA
When to sendAt onceAt the start of the next slot
Vulnerable period2 frame times1 frame time (the slot)
ThroughputG e to the power -2GG e to the power -G
Capacity1/(2e), about 0.184, at G = 0.51/e, about 0.368, at G = 1
NeedsNothingA common clock for the slots
1-persistentNon-persistentp-persistent
Busy channelListen until idle, then sendBack off a random time, sense againWait until idle, then send with probability p per slot
Light loadBest: no idle gapsWastes idle timeIn between
Heavy loadWaiting stations collideBestGood, with a small p
Energy while waitingReceiver on throughoutRadio can sleepListens each slot
Capacity, a = 0.010.5290.815Highest, near p = 0.03
ALOHACSMA
Listens before sendingNoYes
CollisionsAny overlapOnly within a of each other
Depends on aNoYes: its capacity falls as a grows

What it does not mean

18.4 per cent is not what ALOHA delivers in practice. It is the ceiling under a model; a real channel run near it goes unstable.

G is not the traffic the stations generate. G counts every attempt, retries included; S counts what gets through. In a stable channel the new traffic equals S.

Non-persistent is not simply better. It wins under heavy load and loses under light load, and it waits longer before sending.

A small propagation delay does not make a small a. On a sensor radio the turnaround and the channel assessment, hundreds of microseconds, set a, not the speed of light.

Carrier sense does not prevent every collision. It prevents collisions only between stations that can hear each other in time; stations hidden from each other are not helped at all.

Quick revision

  • Pure ALOHA (Abramson 1970): send at once; lost frames resent after a random delay; vulnerable period 2 frame times; S = G e to the power -2G; capacity 1/(2e), about 0.184, at G = 0.5 (the report's "0.186" is a slip); unstable past the peak.
  • Slotted ALOHA (Roberts): slots of one frame; vulnerable period 1; S = G e to the power -G; capacity 1/e, about 0.368, at G = 1; needs a common clock. Still used for random access in GPRS, UMTS, TETRA.
  • CSMA: listen before sending; vulnerable period about a (delay before a transmission is heard, as a fraction of a frame).
  • 1-persistent: busy, listen until idle, then send (waiting stations collide). Non-persistent: busy, back off a random time. p-persistent: idle, send with probability p per slot.
  • Capacities at a = 0.01: 0.529 (1-persistent), 0.815 (non-persistent, at G about 9.4); at a = 0.12, 0.439 and 0.483.
  • Sensor radio: 802.15.4 turnaround up to 12 symbols = 192 microseconds, CCA 8 symbols = 128 microseconds; a 50-byte frame (1.6 ms) gives a = 0.12 to 0.2.
  • Energy per delivered frame at best load: ALOHA about 2.7 transmissions; non-persistent CSMA about 1.2 plus short listens. Non-persistent lets the radio sleep while it waits: B-MAC and 802.15.4 follow it.
munotes.in306

Contention: ALOHA and CSMA

Test yourself

1. Explain pure ALOHA and derive its maximum throughput. A station transmits a frame whenever it has one; if no acknowledgement comes it retransmits after a random delay. A frame starting at t collides with any frame started in the frame time before or after t, so its vulnerable period is two frame times. If attempts form a Poisson stream of G per frame time, the probability of no other start in two frame times is e to the power -2G, and the throughput is S = G e to the power -2G. Its derivative is zero at G = 0.5, where S = 1/(2e), about 0.184.

2. How does slotted ALOHA double the throughput? Time is divided into slots one frame long, and stations may start only at slot boundaries. Frames that collide then overlap completely, and a frame is vulnerable only to frames started in the same slot, one frame time instead of two. So S = G e to the power -G, whose maximum is 1/e, about 0.368, at G = 1. It needs all stations to share a clock.

3. Compare 1-persistent, non-persistent and p-persistent CSMA. All listen first and transmit if the channel is idle. If it is busy, a 1-persistent station keeps listening and transmits as soon as it becomes idle, so stations that were waiting collide; a non-persistent station waits a random time and senses again, which avoids that collision but leaves the channel idle at times; a p-persistent station waits for idle and then transmits with probability p in each slot, spreading the waiting stations out. 1-persistent does best at light load, non-persistent and p-persistent at heavy load.

4. What is a in CSMA, and why does a sensor radio have a large one? It is the time before a transmission can be sensed by other stations, as a fraction of a frame's transmission time; collisions can happen only between stations that start within about a of each other. For a sensor radio it is set not by propagation but by the radio's switch from receiving to transmitting (up to 192 microseconds on 802.15.4) and the 128 microsecond channel assessment. For a 50-byte frame lasting 1.6 ms, a is 0.12 to 0.2, and CSMA's capacity falls towards slotted ALOHA's.

munotes.in307

Contention: ALOHA and CSMA

5. Abramson's consoles each send a 34 ms packet every 30 seconds. How many can a pure ALOHA channel support? Each uses 0.034 / 30, about 0.001133, of the channel's time. The channel can carry at most 1/(2e), so the number is 1 / (2e × 0.001133); 2e × 0.001133 is about 0.00616, and 1 / 0.00616 is about 162 active consoles.

6. Which CSMA rule suits a battery-powered node, and why? Non-persistent: a node that finds the channel busy waits a random time before sensing again and can switch its radio off meanwhile, whereas a 1-persistent node must keep its receiver on until the channel goes idle. B-MAC and 802.15.4's CSMA-CA both back off and sense again in this way.

Contents This chapter on its own page

munotes.in308

Chapter Forty-Seven

Hidden and Exposed Terminals, and RTS and CTS

Syllabus topic Module 1, "Medium Access Control (MAC) in WSN: Fundamentals of MAC protocols for sensor networks"

In one line

Carrier sensing listens at the sender, but collisions happen at the receiver; a hidden terminal is a sender that cannot be heard yet can still spoil the reception, an exposed terminal is one that is heard yet could not, and the RTS and CTS handshake answers both by making the receiver speak first, so that everyone near the receiver learns to keep quiet.

In the wording a student can write in an examination: the hidden terminal problem arises when two stations, A and C, are both in range of a receiver B but out of range of each other: each senses the channel idle and transmits, and their frames collide at B. The exposed terminal problem arises when a station C hears a transmission from B to A and defers, although its own transmission to D, out of A's range, would not have disturbed it: the channel is wasted. In the RTS/CTS handshake (MACA, used in IEEE 802.11 and S-MAC), the sender first sends a short Request to Send (RTS) carrying the length of the transmission; the receiver answers with a Clear to Send (CTS) carrying the same duration; the sender then sends DATA, and the receiver an ACK. A station that overhears a CTS defers until the exchange ends, which stops hidden terminals; a station that hears an RTS but no CTS may transmit, which relieves exposed terminals. The overheard duration is kept in the network allocation vector (NAV): a station treats the medium as busy while its NAV is non-zero (virtual carrier sense) or while it hears a signal (physical carrier sense). Collisions can still happen between RTS frames, but they are short and cheap. For frames little longer than an RTS, the handshake costs more than it saves.

Why listening at the sender is not enough

A CSMA station listens before it sends, at its own position. But a frame succeeds or fails at the receiver's position, and the two can hear different things. Karn, writing for amateur packet radio in 1990, put the consequence in two sentences: "When hidden terminals exist, lack of carrier doesn't always mean it's OK to transmit. Conversely, when exposed terminals exist, presence of carrier doesn't always mean that it's bad to transmit."

Radio ranges make both situations ordinary. Bonaventure's examples are the everyday ones: "two devices separated by a wall may not be able to receive each other's signal while they could both be receiving the signal produced by a third host", and a hill does the same outdoors. In a sensor network, where links in the transitional region come and go ([TOSSIM: Simulating Motes, Radio Gain and Packet Loss]) and nodes lie on the ground among obstacles, some pairs of neighbours of every node will fail to hear each other.

munotes.in309

Hidden and Exposed Terminals, and RTS and CTS

Two panels. Hidden terminal: A, B and C in a line, the ranges of A and C overlapping only around B, with arrows from A and C both into B. Exposed terminal: A, B, C and D in a line, B sending to A, C in B's range with a dashed arrow towards D

Figure 47.1 The two problems, with the ranges that cause them

The hidden terminal

"In the classic hidden terminal situation, station Y can hear both stations X and Z, but X and Z cannot hear each other. X and Z are therefore unable to avoid colliding with each other at Y." In the figure the stations are A, B and C. A is sending to B; C, which cannot hear A, senses an idle channel and starts sending to B too. B receives two signals at once and decodes neither. Neither sender knows until no acknowledgement comes.

Between two hidden stations, carrier sensing does nothing: they cannot hear each other, so to each other they behave like pure ALOHA stations. [Contention: ALOHA and CSMA] showed what that costs: a frame is vulnerable to anything the other starts within one frame time either side of it, so the longer the frame, the likelier the collision.

The exposed terminal

The opposite mistake wastes the channel instead of frames. Karn: "In the exposed terminal case (figure 2), a well-sited station X can hear far away station Y. Even though X is too far from Y to interfere with its traffic to other nearby stations, X will defer to it unnecessarily, thus wasting an opportunity to reuse the channel locally."

In the figure, B is sending to A. C is in B's range, hears the carrier and holds back its frame for D. But A, the receiver that matters, is out of C's range, so C's frame could not have disturbed A's reception, and D is out of B's range, so B's frame would not have disturbed D's. Two transmissions could have gone at once; CSMA allowed one.

RTS and CTS: make the receiver speak

The idea came from the handshake of Apple's LocalTalk network, where, as Karn describes it, a station "first sends a short Request To Send (RTS) packet to the destination. The receiver responds with a Clear to Send (CTS) packet." LocalTalk used it to let the receiver prepare; Karn saw that it also tells the receiver's neighbours that a reception is about to happen. His protocol, MACA (Multiple Access with Collision Avoidance), makes that the rule:

"When a station overhears an RTS addressed to another station, it inhibits its own transmitter long enough for the addressed station to respond with a CTS. When a station overhears a CTS addressed to another station, it inhibits its own transmitter long enough for the other station to send its data."

To know how long "long enough" is, the RTS carries the length of the coming data and the CTS repeats it: "We have X, the initiator of the dialogue, include in its RTS packet the amount of data it plans to send, and we have Y, the responder, echo that information in its CTS packet."

munotes.in310

Hidden and Exposed Terminals, and RTS and CTS

The hidden terminal, answered. C cannot hear A's RTS, but it hears B's CTS, learns how long A's data will take, and stays silent for that long. B's reception is protected by the one station that matters, B itself, speaking before the data arrives.

The exposed terminal, relieved. C hears B's RTS to A but not A's CTS, because A is out of C's range. Karn's rule: "if a station hears no response to an overheard RTS, then it may assume that the intended recipient of the RTS is either down or out of range", and so C may transmit while B's data is going out. MACA dropped carrier sensing altogether ("let's get rid of the CS in CSMA/CA"), relying on the RTS and CTS alone.

Collisions that remain. Two stations can still send RTS frames at the same moment. MACA backs off "similar to that used in regular CSMA", choosing a random wait whose slot "is the duration of an RTS packet", and doubling the average wait after each failure. What makes this worthwhile is size: "MACA still has the advantage over CSMA as long as the RTS packets are significantly smaller than the data packets." A collision of two short RTS frames wastes a few bytes of airtime; a collision of two data frames wastes all of them.

The acknowledgement, and what it does to the exposed terminal. Later protocols added an ACK from the receiver after the data, so the sequence became RTS, CTS, DATA, ACK: IEEE 802.11 uses it (Bonaventure), and S-MAC states that "Unicast packets follow the sequence of RTS/CTS/DATA/ACK between the sender and the receiver." The ACK makes loss visible at once, but it undoes much of the exposed-terminal relief: B must receive A's ACK at the end of its data, and C, in B's range, would be transmitting over it. An exposed station can safely send only if its own exchange ends before that ACK, which in practice it cannot know.

The NAV: carrier sense without hearing a carrier

A station that heard an RTS or a CTS must remember how long to keep quiet, even if the channel it hears goes silent (C, near B, hears nothing while A sends its data). S-MAC, following 802.11, describes the mechanism exactly: "There is a duration field in each transmitted packet that indicates how long the remaining transmission will be. So if a node receives a packet destined to another node, it knows how long it has to keep silent. The node records this value in an variable called the network allocation vector (NAV)", and sets a timer for it.

munotes.in311

Hidden and Exposed Terminals, and RTS and CTS

The node counts the NAV down, and "When a node has data to send, it first looks at the NAV. If its value is not zero, the node determines that the medium is busy. This is called virtual carrier sense." Listening to the radio itself remains: "Physical carrier sense is performed at the physical layer by listening to the channel for possible transmissions." And the two combine: "The medium is determined as free if both virtual and physical carrier sense indicate that it is free."

A timeline of one exchange: A sends RTS, B replies CTS, A sends DATA, B sends ACK. C, near B, is shaded as deferring from the end of the CTS to the end of the ACK; D, near A, from the end of the RTS to the end of the ACK

Figure 47.2 One exchange, and how long each neighbour's NAV keeps it silent (schematic)

In the figure, D is near A and hears only A: A's RTS sets D's NAV to cover the CTS, the data and the ACK. C is near B and hears only B: B's CTS sets C's NAV to cover the data and the ACK. Bonaventure sums up why both frames carry the duration: "all hosts that could collide with either the sender or the reception of the data frame are informed of the reservation."

What the handshake costs, worked out

The handshake is not free: two extra frames per data frame, which a sensor network pays for in airtime and in energy ([MAC Protocols for Sensor Networks: The Job and Where the Energy Goes] counted it as control overhead). Whether it pays depends on how often the hidden terminal strikes and on how long the data frame is.

The program models A sending to B while a hidden C starts transmissions at random, 20 or 100 a second, at 250 kbit/s, with control frames of 11 bytes on air. Without the handshake, A's frame is lost if C starts anything within one frame time of it, the pure ALOHA window; a simulation lays C's starts on a timeline and drops A's frames on it to check that. With the handshake, only the RTS is exposed, because B's CTS silences C for the rest. For each case the program gives the chance an attempt survives and the airtime spent per delivered frame, counting every retry.

# A and C cannot hear each other; both send to B, between them. Carrier sense
# cannot protect them from each other, so between them the channel is pure
# ALOHA: A's frame is lost if C starts anything that overlaps it at B.
# C starts its transmissions at random, a Poisson stream of `rate` a second.
import bisect
import random
from math import exp

BYTE = 8 / 250_000             # seconds per byte on air at 250 kbit/s
CTRL = 11                      # an RTS, a CTS or an ACK: bytes on air

def plain(data, rate):
    """A sends DATA and B answers with an ACK. Returns the chance a frame
    survives, and the airtime spent per delivered frame, in ms."""
    t = data * BYTE
    ok = exp(-2 * rate * t)    # nothing from C may start within t either side
    return ok, (t / ok + CTRL * BYTE) * 1000

def rts_cts(data, rate):
    """RTS, CTS, DATA, ACK: only the short RTS is exposed to C, because B's
    CTS tells C to keep quiet for the rest of the exchange."""
    t = CTRL * BYTE
    ok = exp(-2 * rate * t)
    return ok, (t / ok + CTRL * BYTE + data * BYTE + CTRL * BYTE) * 1000

def simulated(data, rate, rnd, seconds=2000, frames=20000):
    """Check plain(): C's starts laid on a timeline, A's frames dropped on it."""
    t, starts, x = data * BYTE, [], 0.0
    while x < seconds:
        x += rnd.expovariate(rate)
        starts.append(x)
    ok = 0
    for _ in range(frames):
        s = rnd.uniform(1, seconds - 1)
        i = bisect.bisect_left(starts, s - t)      # C's first start after s - t
        ok += starts[i] >= s + t                   # none before s + t: survived
    return ok / frames

rnd = random.Random(47)
print("C/s  bytes |  plain: ok  (sim)    ms | RTS/CTS: ok     ms")
for rate in (20, 100):
    for data in (20, 50, 127):
        p_ok, p_ms = plain(data, rate)
        r_ok, r_ms = rts_cts(data, rate)
        print("%3d %6d | %10.3f %6.3f %6.2f | %11.3f %6.2f"
              % (rate, data, p_ok, simulated(data, rate, rnd), p_ms, r_ok, r_ms))
munotes.in312

Hidden and Exposed Terminals, and RTS and CTS

C/s  bytes |  plain: ok  (sim)    ms | RTS/CTS: ok     ms
 20     20 |      0.975  0.975   1.01 |       0.986   1.70
 20     50 |      0.938  0.934   2.06 |       0.986   2.66
 20    127 |      0.850  0.846   5.13 |       0.986   5.12
100     20 |      0.880  0.878   1.08 |       0.932   1.72
100     50 |      0.726  0.726   2.56 |       0.932   2.68
100    127 |      0.444  0.449   9.51 |       0.932   5.15

Reading it. The simulated survival agrees with the formula to within a few thousandths in every row, so the hidden pair really does behave as pure ALOHA.

Short frames: the handshake loses. With C sending 20 times a second, a 20-byte frame survives 0.975 of the time on its own; the handshake raises that to 0.986 but adds two control frames, and the airtime per delivered frame rises from 1.01 ms to 1.70. At 50 bytes it still loses, 2.66 ms against 2.06.

Long frames under pressure: the handshake wins. At 127 bytes, the largest 802.15.4 frame, the two break even when C sends 20 times a second (5.12 against 5.13 ms). When C sends 100 times a second, a plain 127-byte frame survives only 0.444 of the time and each delivered frame costs 9.51 ms; the handshake, whose short RTS survives 0.932 of the time, needs 5.15. Karn's own advice follows: "If the data packets are of comparable size to the RTS packets, the overhead of the RTS/CTS dialogue may be excessive", and a station may then send "its data without the dialogue."

munotes.in313

Hidden and Exposed Terminals, and RTS and CTS

The model is simple on purpose: C's traffic ignores A's, lost CTS and ACK frames are not counted, and retries are independent. It is enough to show why the answer depends on frame length and load, and why practice varies. In 802.11, Bonaventure notes, reservation "is an optimization that is useful when collisions are frequent", and "Some devices only turn on RTS/CTS after transmission errors."

RTS and CTS in sensor networks

Sensor frames are short: 802.15.4 carries at most 127 bytes, and a typical reading is a few tens. The sensor MACs divide accordingly.

S-MAC keeps the handshake, for unicast frames only: "We adopt the RTS/CTS mechanism to address the hidden terminal problem". It also puts the handshake to a second use: a neighbour that overhears an RTS or CTS goes to sleep for the exchange, which stops it overhearing the data ([S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing]).

B-MAC leaves it out. "Since B-MAC does not have the RTS-CTS mechanism or synchronization requirements of S-MAC", its authors write, "the implementation is both simpler and smaller"; a service that wants a handshake can build one above it.

IEEE 802.15.4 defines none. Its 2006 text has no RTS, no CTS and no mention of hidden terminals; its CSMA-CA relies on backoff and a clear channel assessment ([CSMA-CA, Data Transfer and Frames in 802.15.4]).

Broadcasts get no protection anywhere. A handshake needs one receiver to answer; a broadcast has many, whose CTS frames would collide. S-MAC: "Broadcast packets are sent without using RTS/CTS." Flooding and route discovery, which broadcast ([Flooding, Gossiping and the Broadcast Storm]), meet hidden terminals unprotected.

Distinctions

Hidden terminalExposed terminal
SituationSender cannot hear another sender that reaches its receiverSender hears another sender that cannot reach its receiver
CSMA's mistakeTransmits when it should notDefers when it need not
ResultCollision at the receiverWasted channel
RTS/CTS answerThe receiver's CTS silences the hidden stationA station that hears an RTS but no CTS may send
Physical carrier senseVirtual carrier sense
HowListening to the radio for a signalCounting down the NAV from overheard durations
CatchesTransmissions it can hear nowExchanges it heard announced, even if it cannot hear the data
MissesHidden sendersExchanges whose RTS and CTS it did not hear
munotes.in314

Hidden and Exposed Terminals, and RTS and CTS

Without the handshakeWith RTS and CTS
Exposed to a hidden senderThe whole data frameOnly the short RTS
OverheadAn ACKRTS, CTS and ACK
Better forShort frames, light loadLong frames, frequent collisions

What it does not mean

Hidden does not mean far away. A hidden station is one out of range of the sender but in range of the receiver; a wall or the ground can do it at a few metres.

RTS and CTS do not end collisions. RTS frames still collide, and a station that cannot decode a CTS it can still interfere with is not silenced; Karn notes that MACA "does not guarantee that they will never occur."

The handshake does not always help. For frames close to an RTS in length it costs more airtime and energy than the collisions it prevents.

The NAV is not a timer on the radio. It is a count the MAC keeps; the radio may even sleep while it runs, which is exactly what S-MAC does.

Quick revision

  • Collisions happen at the receiver; carrier sense listens at the sender.
  • Hidden terminal: A and C both reach B, not each other; both sense idle and collide at B. Between them the channel is pure ALOHA.
  • Exposed terminal: C hears B sending to A and defers, though its frame to D could not disturb A: wasted channel.
  • RTS/CTS (MACA, Karn 1990; 802.11; S-MAC): RTS carries the duration, CTS echoes it; overhear an RTS, wait for the CTS; overhear a CTS, stay silent for the data. Hears an RTS but no CTS: may send.
  • RTS/CTS/DATA/ACK in 802.11 and S-MAC; the ACK undoes most of the exposed-terminal relief.
  • NAV: duration from overheard frames, counted down; non-zero means busy (virtual carrier sense); free only if virtual and physical carrier sense agree.
  • Cost: at 20 hidden starts a second the handshake loses for 20 and 50 bytes and breaks even at 127; at 100 a second it wins at 127 bytes (5.15 ms against 9.51).
  • Sensor MACs: S-MAC uses it for unicast; B-MAC and 802.15.4 do not; broadcasts never use it.

Test yourself

1. Explain the hidden terminal problem with a diagram. Stations A and C are both within range of B but not of each other. A starts sending to B. C senses the channel, hears nothing because A is out of its range, and also starts sending to B. The two frames overlap at B, which receives neither. Carrier sensing failed because it was done at the senders, which cannot hear each other, while the collision happened at the receiver.

2. Explain the exposed terminal problem. B is sending to A. C, in range of B but not of A, wants to send to D, which is out of B's range. C hears B's carrier and defers, although its transmission could not disturb A's reception and B's could not disturb D's. The channel could have carried both transmissions at once, and one opportunity is wasted.

munotes.in315

Hidden and Exposed Terminals, and RTS and CTS

3. How does the RTS/CTS handshake solve the hidden terminal problem? The sender sends a short RTS giving the duration of the coming exchange, and the receiver answers with a CTS repeating it. Every station in range of the receiver, including a hidden one that could not hear the sender, hears the CTS and stays silent until the exchange ends, so the data frame cannot be hit by it. Only the short RTS frames can still collide.

4. What is the NAV, and what is virtual carrier sense? The network allocation vector is a counter a station sets from the duration field of an overheard frame not addressed to it, and counts down to zero. While it is non-zero the station treats the medium as busy, even if it hears no signal: this is virtual carrier sense. Physical carrier sense listens to the radio; the medium is free only when both say so.

5. When is the RTS/CTS handshake not worth using? Use numbers. When data frames are short compared with the RTS and CTS, the two extra frames cost more than the collisions they prevent. With a hidden station starting 20 transmissions a second at 250 kbit/s, a 20-byte frame costs 1.01 ms of airtime per delivered frame without the handshake and 1.70 ms with it; only a long frame under heavy hidden traffic gains, such as 127 bytes at 100 starts a second (9.51 ms without, 5.15 ms with).

6. Why do broadcasts in a sensor network get no RTS/CTS protection? A handshake needs a single receiver to answer the RTS with a CTS. A broadcast is for every neighbour; if all answered, their CTS frames would collide. So protocols such as S-MAC send broadcasts without RTS and CTS, and broadcast traffic remains exposed to hidden terminals.

Contents This chapter on its own page

munotes.in316

Chapter Forty-Eight

CSMA/CA Worked Step by Step

Syllabus topic Module 1, "Medium Access Control (MAC) in WSN: Fundamentals of MAC protocols for sensor networks" (and the paired practical, "Implement CSMA/CA protocol simulation and evaluate collision avoidance mechanisms")

In one line

A radio cannot hear a collision while it transmits, so CSMA/CA avoids collisions instead: a station waits for the channel to be idle, then waits a random number of slots more, drawn from its contention window, and doubles the window each time a frame fails, so that the more stations compete, the more widely they spread out.

In the wording a student can write in an examination: in CSMA/CA (carrier sense multiple access with collision avoidance), a station with a frame first senses the channel. In IEEE 802.11 it waits until the channel has been idle for a fixed gap, the DIFS, then chooses a random backoff of 0 to CW - 1 slots, where CW is its contention window, and counts it down, one idle slot at a time; while the channel is busy, the countdown is frozen. At zero it transmits. The receiver answers with an ACK after a shorter gap, the SIFS, which no other station can use because every other station waits at least a DIFS. If no ACK arrives, the station assumes a collision, doubles its contention window (binary exponential backoff), up to a maximum, and tries again; after a success, the window returns to its minimum. Two stations collide only if their backoffs end in the same slot, which a larger window makes less likely. IEEE 802.15.4 uses a simpler form suited to batteries: wait a random 0 to 2 to the power BE minus 1 backoff periods, then make one clear channel assessment; if busy, increase BE and wait again, and give up after a few tries.

Why avoid collisions instead of detecting them

On a wired Ethernet a station listens while it sends and stops the moment it hears a collision (CSMA/CD). A radio cannot: its own signal drowns everything else at its antenna, and the collision happens at the receiver anyway ([MAC Protocols for Sensor Networks: The Job and Where the Energy Goes]). Bonaventure draws the consequence: "it is impossible in CSMA/CA to detect collisions as they happen. With CSMA/CA, a collision may affect the entire frame while with CSMA/CD it can only affect the beginning of the frame."

A wireless collision therefore wastes a whole frame, and is noticed only when no acknowledgement comes. So the effort goes into avoiding it, and the lesson comes from [Contention: ALOHA and CSMA]: 1-persistent CSMA collides whenever two stations wait for the same busy channel, because they both start the moment it goes idle. CSMA/CA makes every waiting station add a random wait, so that they almost never start together.

The steps, for one frame

Bonaventure sets out 802.11's version. A station with a frame to send does this:

munotes.in317

CSMA/CA Worked Step by Step

  1. Wait for an idle gap. It waits until the channel has been idle for a DIFS (distributed coordination function inter frame space), or, if the last frame it received was corrupted, for the longer EIFS.
  2. Choose a backoff. It chooses a random whole number of slots from its contention window.
  3. Count down, and freeze. It counts the backoff down by one for each slot the channel stays idle. If another station starts sending, the count stops where it is; in Bonaventure's words, "the back off timer must be frozen until the channel becomes free again."
  4. Transmit when the count reaches zero.
  5. Wait for the ACK. The receiver checks the frame and, after a short gap, the SIFS, sends an acknowledgement. Because SIFS is shorter than DIFS, no station waiting for a DIFS can start before the ACK.
  6. On failure, double the window. No ACK means the frame was probably lost in a collision. The station doubles its contention window, up to a maximum, chooses a new backoff and returns to step 1. After a success the window returns to its minimum; after too many failures the frame is dropped.
Three stations on a timeline. A sends DATA and its receiver an ACK; B and C then wait a DIFS and count down backoffs of 3 and 5 in shaded slots. B reaches zero and sends; C is frozen at 2 until after B's ACK and another DIFS, then counts 2 and 1 and sends

Figure 48.1 Step 3 on a timeline: C's countdown is frozen, not restarted, while B sends (schematic)

In the figure, B and C became ready while A was sending. Both wait for A's ACK and a DIFS, then count: B from 3, C from 5. After three idle slots B reaches zero and sends; C, at 2, freezes. When B's exchange and another DIFS are over, C resumes from 2, not from a new random number, and sends two slots later. Freezing is what makes CSMA/CA fair over time: a station that has already waited keeps its place.

Why a random backoff avoids collisions

Two stations collide only if their countdowns reach zero in the same slot. If each draws its backoff uniformly from W slots, the chance that two stations draw the same number is 1/W: 1/8 = 0.125 with a window of 8, and 1/64, about 0.016, with 64.

The trouble is the number of stations. With many contenders and a small window, ties become likely, and every tie wastes a whole frame time for every station involved. No fixed window suits every load: a large one wastes time when only two stations compete, a small one collides when twenty do. Binary exponential backoff lets the window find the load. Each collision is evidence of crowding, and doubling the window after it spreads the colliding stations over twice as many slots. Bonaventure: "The range grows exponentially with the retransmissions as in CSMA/CD." Karn's MACA did the same, "doubling the average interval on each successive attempt".

munotes.in318

CSMA/CA Worked Step by Step

The 802.15.4 way: sleep, then look once

IEEE 802.15.4, the standard under most sensor radios, keeps the random backoff but changes what the radio does while waiting. Its unslotted CSMA-CA, used when there are no beacons, keeps two numbers for each frame: NB, how many times it has had to back off, starting at 0, and BE, the backoff exponent, starting at macMinBE. Then:

  1. "The MAC sublayer shall delay for a random number of complete backoff periods", anywhere from 0 to 2 to the power BE, minus 1, a backoff period being 20 symbols, 320 microseconds at 2.4 GHz.
  2. It then asks the radio for one clear channel assessment (CCA), which listens for 8 symbols, 128 microseconds.
  3. Idle: it transmits at once.
  4. Busy: "the MAC sublayer shall increment both NB and BE by one, ensuring that BE shall be no more than macMaxBE." If NB is still at most macMaxCSMABackoffs, it returns to step 1; otherwise "the CSMA-CA algorithm shall terminate with a channel access failure status."

The defaults are macMinBE 3, macMaxBE 5 and macMaxCSMABackoffs 4, so the first wait is 0 to 7 backoff periods, the window grows to 0 to 31, and a frame that finds the channel busy five times is given up. A frame sent but not acknowledged is retried up to macMaxFrameRetries, 3, times, each with a fresh CSMA-CA.

The difference from 802.11 is in step 1. 802.11 counts idle slots, so it must listen to every slot to know whether it was idle. 802.15.4 simply waits a random time and looks once at the end; the standard enables the receiver "during the CCA analysis portion of this algorithm", and nothing in the delay needs it. A sensor radio can sleep through the backoff and wake to listen for 128 microseconds. The price is that the wait no longer adapts to what happened during it, and a node that finds the channel busy a few times in a row gives up. The slotted version, with its contention window of two assessments, is in [CSMA-CA, Data Transfer and Frames in 802.15.4].

An event, simulated

In a sensor network the hard case is not steady traffic but an event: something happens, and every node that sensed it has a report at the same moment. The program puts n nodes in range of each other, gives each one report at time zero, and runs until every report is delivered or dropped. Time runs in backoff periods of 320 microseconds; a 50-byte frame lasts 5 of them (1.6 ms), and with the turnaround and the ACK the channel is busy for 7.

It compares three ways of backing off. The first two are 802.11-style, counting idle slots and freezing, with a window of 8 that either stays fixed or doubles after each collision up to 64. The third is 802.15.4's sleep-and-look with the default parameters, and the fourth the same with the widest the standard allows (BE up to 8, five backoffs). For each it reports, averaged over 2,000 events: frames lost in collisions per report, slots until the burst is over, slots each node spent listening, and reports dropped.

munotes.in319

CSMA/CA Worked Step by Step

# An event: n nodes, all in range of each other, each have one report to send
# at the same moment. Time runs in backoff slots of 320 microseconds; a 50-byte
# frame and its acknowledgement keep the channel busy for 7 slots. Two nodes
# that start in the same slot collide, and both must try again.
import random

FRAME = 7                                   # slots: 5 for the frame, 2 for turnaround and ACK

def listen_and_count(n, cw_min, cw_max, rnd):
    """802.11 style: count a random backoff down in idle slots only, frozen
    while the channel is busy, so the radio listens the whole time it waits.
    After a collision the window doubles, up to cw_max."""
    cw = [cw_min] * n
    left = [rnd.randrange(cw_min) for _ in range(n)]
    waiting, t, lost, heard = set(range(n)), 0, 0, 0
    while waiting:
        ready = [i for i in waiting if left[i] == 0]
        if ready:
            heard += FRAME * (len(waiting) - len(ready))   # the rest listen to it
            if len(ready) == 1:
                waiting.discard(ready[0])
            else:
                lost += len(ready)
                for i in ready:
                    cw[i] = min(2 * cw[i], cw_max)
                    left[i] = rnd.randrange(cw[i])
            t += FRAME
        else:
            for i in waiting:
                left[i] -= 1
            heard += len(waiting)                          # an idle slot, listened to
            t += 1
    return lost, t, heard / n, 0

def sleep_and_sense(n, rnd, min_be=3, max_be=5, max_backoffs=4, max_retries=3):
    """802.15.4 style: sleep a random 0 to 2**BE - 1 slots, then assess the
    channel once (8 symbols, 0.4 of a slot). Busy: BE grows, up to max_be, and
    after max_backoffs + 1 busy assessments the frame is dropped."""
    wake = {i: rnd.randrange(2 ** min_be) for i in range(n)}
    nb, be, tries = [0] * n, [min_be] * n, [0] * n
    t, busy_until, lost, heard, dropped = 0, 0, 0, 0.0, 0
    while wake:
        now = [i for i in wake if wake[i] == t]
        heard += 0.4 * len(now)
        if now and t < busy_until:                 # the channel is busy: back off
            for i in now:
                nb[i], be[i] = nb[i] + 1, min(be[i] + 1, max_be)
                if nb[i] > max_backoffs:
                    del wake[i]
                    dropped += 1
                else:
                    wake[i] = t + 1 + rnd.randrange(2 ** be[i])
        elif now:                                  # idle: everyone who looked sends
            busy_until = t + FRAME
            if len(now) == 1:
                del wake[now[0]]
            else:
                lost += len(now)
                for i in now:
                    tries[i] += 1
                    heard += 2                     # waited for an ACK that never came
                    if tries[i] > max_retries:
                        del wake[i]
                        dropped += 1
                    else:
                        nb[i], be[i] = 0, min_be
                        wake[i] = busy_until + rnd.randrange(2 ** min_be)
        t += 1
    return lost, t, heard / n, dropped

rnd = random.Random(48)
RUNS = 2000
print("nodes  method               lost/frame  slots  listened  dropped")
for n in (5, 20):
    for name, run in (("window 8, fixed", lambda: listen_and_count(n, 8, 8, rnd)),
                      ("window 8, doubling", lambda: listen_and_count(n, 8, 64, rnd)),
                      ("802.15.4, BE 3 to 5", lambda: sleep_and_sense(n, rnd)),
                      ("802.15.4, BE 3 to 8", lambda: sleep_and_sense(n, rnd, max_be=8,
                                                                    max_backoffs=5))):
        tot = [0, 0, 0, 0]
        for _ in range(RUNS):
            for k, v in enumerate(run()):
                tot[k] += v
        lost, slots, heard, dropped = (v / RUNS for v in tot)
        print("%5d  %-20s %10.2f %6.1f %9.1f %8.2f"
              % (n, name, lost / n, slots, heard, dropped))
munotes.in320

CSMA/CA Worked Step by Step

nodes  method               lost/frame  slots  listened  dropped
    5  window 8, fixed            0.67   56.5      23.1     0.00
    5  window 8, doubling         0.54   61.5      25.4     0.00
    5  802.15.4, BE 3 to 5        0.32   56.1       1.8     0.07
    5  802.15.4, BE 3 to 8        0.31   74.6       1.8     0.01
   20  window 8, fixed            6.32  439.0     213.0     0.00
   20  window 8, doubling         2.08  329.4     167.3     0.00
   20  802.15.4, BE 3 to 5        0.97  128.9       4.4    10.87
   20  802.15.4, BE 3 to 8        0.87  354.4       4.1     1.84

Reading it: five nodes. With a fixed window of 8, each report loses 0.67 frames in collisions on the way; doubling the window cuts that to 0.54 but makes the burst slightly longer (61.5 slots against 56.5), because the wider windows it grows into are more than five nodes need. Either way each node listens for about 23 to 25 slots. 802.15.4's sleep-and-look loses 0.32 frames per report, finishes in 56.1 slots, and each node listens for only 1.8 slots: about thirteen times less.

Reading it: twenty nodes. Now a fixed window of 8 is far too small for 20 contenders: 6.32 frames lost per report and 439 slots, 140 ms, to clear the burst. Doubling is the backoff doing its job: 2.08 lost per report, and 329 slots. The 802.11-style nodes each listen for well over 150 slots.

The default 802.15.4 settings give up. With BE limited to 5 and five assessments allowed, only 0.97 frames are lost per report and each node listens for 4.4 slots, but 10.87 of the 20 reports are dropped with a channel access failure: a node that finds the channel busy five times stops trying. With the widest settings the standard allows, the drops fall to 1.84, at the price of a burst lasting 354 slots.

munotes.in321

CSMA/CA Worked Step by Step

In energy. A slot is 320 microseconds. 213 slots of listening at the CC2420's 18.8 mA is 213 × 0.32 = 68.16 ms, about 1.28 mC of charge per node for one report; 4.4 slots is 1.408 ms, about 26.5 microcoulombs. Counting idle slots costs a sensor node roughly fifty times as much as sleeping through them.

The model is simple: every node hears every other, there is no noise, and frames are the same length. It is enough to show the three forces at work: the backoff trades collisions for delay, how the radio waits decides the energy, and a MAC that gives up easily loses data in exactly the moments, events, when the network has most to say.

Distinctions

CSMA/CD (wired Ethernet)CSMA/CA (wireless)
CollisionsDetected while sending; transmission stoppedCannot be detected; avoided by random backoff
Evidence of a collisionThe station hears itNo acknowledgement arrives
Cost of a collisionThe start of a frameThe whole frame
802.11-style backoff802.15.4 unslotted CSMA-CA
WaitingCounts idle slots down, frozen while busyWaits a random time, then one CCA
Radio while waitingListening throughoutCan sleep; on only for the 128 microsecond CCA
Busy channelFreeze and resumeNB and BE up by one; wait again
Gives upAfter a retry limit on unacknowledged framesAlso after macMaxCSMABackoffs busy assessments
WindowDoubles after each collision2 to the power BE, BE from 3 to 5 by default
Fixed windowBinary exponential backoff
Few contendersGood if the window is smallGood: starts small
Many contendersCollisions pile upThe window grows to match
After a successUnchangedBack to the minimum

What it does not mean

Collision avoidance is not collision prevention. Two stations whose backoffs end in the same slot still collide; so do hidden stations, which never hear each other at all ([Hidden and Exposed Terminals, and RTS and CTS]).

A bigger window is not always better. It cuts collisions but lengthens every wait; with few contenders it only adds delay.

Freezing is not restarting. A frozen station keeps its remaining count; drawing a new number each time would push unlucky stations to the back again and again.

A channel access failure is not a collision. It means the node found the channel busy too many times and stopped trying; the frame never went out.

Quick revision

  • Radio cannot detect collisions, which cost whole frames: CSMA/CA avoids them.
  • 802.11 steps: idle for DIFS (EIFS after a corrupted frame); random backoff from the contention window; count down in idle slots, frozen while busy; transmit at zero; ACK after SIFS (SIFS < DIFS < EIFS).
  • No ACK: double the window (binary exponential backoff) up to a maximum; success: back to the minimum; retry limit.
  • Two stations with a window of W collide with probability 1/W: 0.125 for 8.
  • 802.15.4 unslotted: NB = 0, BE = macMinBE (3); wait random 0 to 2 to the power BE minus 1 periods of 320 microseconds; one CCA (8 symbols); idle, send; busy, NB and BE up (BE at most macMaxBE, 5); NB above macMaxCSMABackoffs (4): channel access failure. Radio can sleep through the wait.
  • Event of 20 nodes (program): fixed window of 8, 6.32 frames lost per report; doubling, 2.08; 802.15.4 defaults, 0.97 lost and 4.4 slots listened, but 10.87 of 20 dropped; widest settings, 1.84 dropped, 354 slots.
munotes.in322

CSMA/CA Worked Step by Step

Test yourself

1. Why does a wireless MAC use collision avoidance rather than collision detection? A radio cannot listen while it transmits, because its own signal drowns any other at its antenna, and the collision happens at the receiver, which the sender may not even hear. So a collision cannot be detected as it happens, destroys the whole frame, and is noticed only when no acknowledgement arrives. The protocol therefore tries to make collisions unlikely in the first place, with random backoff.

2. List the steps a station follows in CSMA/CA to send one frame. It waits until the channel has been idle for a DIFS; it chooses a random backoff from its contention window; it counts the backoff down one slot for each idle slot, freezing the count while the channel is busy; at zero it transmits; the receiver replies with an ACK after a SIFS. If no ACK comes, the station doubles its contention window, up to a maximum, chooses a new backoff and tries again; after a success the window returns to its minimum.

3. Why is the ACK sent after a SIFS rather than a DIFS? SIFS is shorter than DIFS, and every other station must see the channel idle for a full DIFS before it counts down or sends. The receiver can therefore always send its ACK before anyone else can start, so the acknowledgement cannot be collided with by stations that heard the data frame.

4. What is binary exponential backoff, and why is it used? After each failed attempt the contention window is doubled, up to a maximum, and the new backoff is drawn from the larger window. A collision suggests many stations are competing, and a larger window spreads them over more slots, so the chance that two choose the same slot falls; with a window of 8 two stations tie with probability 1/8, with 64, 1/64. After a success the window is reset, so light load keeps short waits.

munotes.in323

CSMA/CA Worked Step by Step

5. How does IEEE 802.15.4's unslotted CSMA-CA differ from 802.11's, and why does it suit sensor nodes? 802.15.4 does not count idle slots: it waits a random 0 to 2 to the power BE minus 1 backoff periods, makes a single 128 microsecond clear channel assessment, sends if the channel is idle, and otherwise increases BE and waits again, giving up with a channel access failure after macMaxCSMABackoffs busy assessments. Because nothing in the wait needs the receiver, the radio can sleep through it, which saves most of the listening energy that 802.11's countdown costs.

6. In the program's 20-node event, what did the 802.15.4 defaults gain and lose? They kept each node listening for only about 4.4 slots against well over 150 for the 802.11-style station, and lost only 0.97 frames per report in collisions, but 10.87 of the 20 reports were dropped after finding the channel busy five times. Allowing BE up to 8 and five backoffs cut the drops to 1.84 but made the burst last 354 slots.

Contents This chapter on its own page

munotes.in324

Chapter Forty-Nine

TDMA and Schedule-based MAC

Syllabus topic Module 1, "Medium Access Control (MAC) in WSN: Fundamentals of MAC protocols for sensor networks" (and the paired practical, "Simulate TDMA slot allocation and compare energy efficiency with CSMA")

In one line

In TDMA each node owns a time slot in a repeating frame and transmits only in it, so there are no collisions and a node's radio can sleep outside its own slots; the price is organisation: slots must be allocated, clocks kept in step, and every reading waits for its slot.

In the wording a student can write in an examination: TDMA (time division multiple access) divides time into repeating frames and each frame into slots; each node is allocated a slot and transmits only in it. Because no two interfering nodes share a slot, there are no collisions, no contention overhead and no idle listening: a node wakes only for its own slot, for slots in which it must receive, and for the synchronisation beacon. Slots may be allocated centrally (a cluster head or the sink builds the schedule, as in LEACH) or distributively (nodes negotiate, as in SMACS). In a multi-hop network, a slot can be reused by nodes far enough apart: two nodes within two hops of each other must have different slots, or a common neighbour would hear both at once. Clocks drift, so slots need guard times that grow with the time since the last synchronisation. TDMA's drawbacks are latency (a reading waits for its slot), wasted slots when a node has nothing to send, poor adaptability when nodes join, leave or change their traffic, and the cost of synchronisation.

Frames and slots

Dargie and Poellabauer put the case for schedules in one sentence: in contention-free MAC protocols "access to the medium is strictly regulated, eliminating collisions and allowing sensor nodes to shut down their radios when no communications are expected." Ye, Heidemann and Estrin say the same from the other side: "TDMA protocols have a natural advantage of energy conservation compared to contention protocols, because the duty cycle of the radio is reduced and there is no contention-introduced overhead and collisions."

Of [MAC Protocols for Sensor Networks: The Job and Where the Energy Goes]'s four wastes, a perfect schedule removes all four: no collisions, because no two interfering nodes share a slot; no overhearing, because a node listens only in the slots addressed to it; no control frames beyond the schedule itself; and no idle listening, because the node knows exactly when to wake.

LEACH is the clearest example in the syllabus. Once the clusters of a round have formed, "The cluster head node sets up a TDMA schedule and transmits this schedule to the nodes in the cluster. This ensures that there are no collisions among data messages and also allows the radio components of each non-cluster head node to be turned off at all times except during their transmit time". Then "The steady-state operation is broken into frames, where nodes send their data to the cluster head at most once per frame during their allocated transmission slot." ([LEACH: Clusters That Take Turns] covers the rest of the protocol.)

munotes.in325

TDMA and Schedule-based MAC

One TDMA frame: a beacon B, ten slots, a long stretch of sleep and the next beacon. The head is shaded as awake for the beacon and all ten slots; member 3 only for the beacon, slightly early, and for slot 3. Below, slot 3 opened up into guard, DATA, ACK, guard

Figure 49.1 One frame of a ten-member cluster, and one slot with its guards

Allocating the slots

In a cluster, one node decides. LEACH's head gives each member one slot, and every slot has the same length, so the time to send a frame grows with the number of nodes in the cluster. The head, which must receive every slot, is awake far longer than any member: "The cluster head must be awake to receive all the data from the nodes in the cluster." LEACH rotates the role for exactly that reason.

Between clusters, neighbouring clusters' slots overlap in time. LEACH separates them by code instead of time: each cluster uses its own direct-sequence spreading code, and the paper notes the cost: "the drawback of using DSSS is the need for tight timing synchronization".

In a multi-hop network, a slot can be used again by nodes far enough apart, which is TDMA's spatial reuse. The rule is set by the receivers. Two nodes that are neighbours cannot share a slot (each may be the other's receiver), and two nodes with a common neighbour cannot either, because that neighbour would hear both at once. So nodes within two hops of each other need different slots; nodes three or more hops apart may share.

Worked example: a chain. Seven nodes in a line, A to G, each hearing only its neighbours, send hop by hop towards A. Giving every node its own slot needs a frame of 7. With the two-hop rule, slots can repeat every third node:

NodeABCDEFG
Slot1231231

Check one pair: B and E both use slot 2. B's neighbours are A and C, E's are D and F; no node hears both, so their frames cannot collide. A, D and G share slot 1 for the same reason. The frame shrinks from 7 slots to 3, and a node's reading waits for a frame less than half as long. Any two nodes one or two hops apart (A and B, A and C) have different slots, so 3 is also the least possible. The paired practical does the same on a grid, where a node and its four neighbours are all within two hops of each other and at least 5 slots are needed.

Distributed allocation needs no head. Akyildiz and colleagues describe SMACS, in which nodes "discover their neighbors and establish transmission/reception schedules for communication without the need for any local or global master nodes", each link being "a pair of time slots operating at a randomly chosen but fixed frequency". Its drawback, in the same survey, is that members of different subnets "might never get connected".

munotes.in326

TDMA and Schedule-based MAC

Keeping the schedule: synchronisation and guard times

A slot is only useful if sender and receiver agree when it starts, and their clocks drift. [Time Synchronisation and Localisation] gave the figure: a mote's clock may be up to 40 ppm out, 40 microseconds per second, and 802.15.4 demands the same ±40 ppm of the radio's crystal, "This accuracy must also take ageing and temperature drift into consideration." Two clocks each 40 ppm out, in opposite directions, drift apart by up to 80 ppm.

So a TDMA node must wake early, and a slot must carry a guard on each side, as wide as the drift since the last synchronisation.

Worked example. If the head's beacon resynchronises everyone once a minute, the drift before the next beacon can reach 80 ppm × 60 s = 0.0048 s, 4.8 ms. A 50-byte frame lasts only 1.6 ms, so each slot needs 4.8 ms of guard on each side of it: the guards are six times the frame. Beaconing every 10 s cuts the guard to 0.8 ms; every second, to 80 microseconds.

The guards cost the listener, who must be awake from the earliest moment the frame might start to the latest. LEACH simply assumes the problem away ("We assume that the nodes are all time synchronized"), suggesting that the base station could send "synchronization pulses". A real schedule pays for it in beacons and in guards, and the program finds the balance.

TDMA against CSMA, for the same traffic

The program charges a LEACH-style cluster of a head and 10 members for one hour. Each member sends one 50-byte reading a minute in its own slot and hears an 11-byte ACK; the head sends a 20-byte beacon every sync seconds, which resets the members' clocks. Members wake early by the guard to hear each beacon; the head listens to every slot for its full width, guards included. Everyone sleeps at 0.02 mA otherwise. For comparison, the same member traffic is charged with CSMA and the receiver always on. Lifetimes are for 2,500 mAh.

# One cluster, as in LEACH: a head and 10 members, each member sending one
# 50-byte reading a minute. TDMA: the frame is one minute and each member owns
# one slot in it. The head's beacon resets the members' clocks every `sync`
# seconds; in between, two clocks each up to 40 ppm out drift apart by up to
# 80 ppm, so a slot needs a guard on each side, and whoever listens for a
# beacon or a slot must wake that much early.
RX, TX, SLEEP = 18.8, 17.4, 0.02           # mA, the CC2420
BYTE = 8 / 250_000                         # seconds per byte on air
DATA, ACK, BEACON = 50, 11, 20             # bytes on air
DRIFT = 2 * 40e-6                          # two clocks, each up to 40 ppm out
MEMBERS, FRAME, HOUR = 10, 60.0, 3600.0

def charge(rx, tx):                        # mA-s over the hour, asleep otherwise
    return rx * RX + tx * TX + (HOUR - rx - tx) * SLEEP

def tdma(sync):
    guard = DRIFT * sync                   # the worst drift since the last beacon
    beacons, frames = HOUR / sync, HOUR / FRAME
    slot = DATA * BYTE + ACK * BYTE + 2 * guard
    member = charge(rx=beacons * (guard + BEACON * BYTE) + frames * ACK * BYTE,
                    tx=frames * DATA * BYTE)
    head = charge(rx=frames * MEMBERS * (2 * guard + DATA * BYTE),
                  tx=beacons * BEACON * BYTE + frames * MEMBERS * ACK * BYTE)
    return slot, member, head

def life(mas):                             # days on 2,500 mAh at this mA-s an hour
    return 2500 / (mas / HOUR) / 24

print("beacon every   slot    member mA-s   days    head mA-s   days")
for sync in (1, 10, 60, 600):
    slot, member, head = tdma(sync)
    print("%9d s %6.1f ms %10.1f %7.0f %11.1f %6.0f"
          % (sync, slot * 1000, member, life(member), head, life(head)))
# the same member traffic with CSMA and the receiver always on
frames = HOUR / FRAME
csma = charge(rx=HOUR - frames * DATA * BYTE, tx=frames * DATA * BYTE)
print("CSMA, always on        %10.1f %7.1f" % (csma, life(csma)))
munotes.in327

TDMA and Schedule-based MAC

beacon every   slot    member mA-s   days    head mA-s   days
        1 s    2.1 ms      122.7    3055       135.5   2767
       10 s    3.6 ms       83.8    4475       115.7   3240
       60 s   11.6 ms       80.2    4676       202.5   1851
      600 s   98.0 ms       79.5    4714      1175.5    319
CSMA, always on           67679.9     5.5

Reading it: the members. A TDMA member spends between 79.5 and 122.7 mA-s an hour, against 67,679.9 for the same traffic with CSMA listening all the time: some 550 to 850 times less. Of its 80.2 mA-s at a one-minute beacon, 72 are the sleep current alone (3,600 s × 0.02 mA). The computed lifetimes, 3,055 to 4,714 days, are 8 to 13 years, beyond the alkaline cell's 10-year shelf life that [How Long a Node Lasts: The Energy Budget Worked Out] warned about: once a schedule has removed idle listening, the battery's chemistry sets the lifetime, not the MAC.

The members hardly care how often the beacon comes. A member's guard grows with the interval, but it hears proportionally fewer beacons, and the two cancel: 80 ppm of every hour, 0.288 s, is spent in guards whatever the interval. Only the beacons' own airtime changes, which is why beaconing every second costs the members most.

munotes.in328

TDMA and Schedule-based MAC

The head does care. It listens to every slot across both guards. With a beacon every 10 s it spends 115.7 mA-s an hour; every minute, 202.5; every 10 minutes, 1,175.5, when each slot has grown to 98.0 ms, almost all guard, and the head lasts 319 days. Every second, the beacons themselves cost it 135.5. The best interval here is near 10 s, and it depends on the clocks: better crystals push it out, worse ones pull it in.

What TDMA costs besides energy

Latency. A reading waits for its node's slot, on average half a frame: 30 s in a one-minute frame. In a multi-hop network it waits again at every hop unless the slots are ordered along the route. S-MAC's authors, whose protocol is not TDMA, face the same trade in its sleep schedule ([S-MAC: Latency, Adaptive Listening and the Energy Saved]).

Wasted slots. A slot belongs to its owner whether or not it has anything to send, and a node cannot borrow another's. Ye, Heidemann and Estrin note of Sohrabi and Pottie's super frame that "A drawback of the scheme is its low bandwidth utilization. For example, if a node only has packets to be sent to one neighbor, it cannot reuse the time slots scheduled to other neighbors."

Change. "When the number of nodes within a cluster changes, it is not easy for a TDMA protocol to dynamically change its frame length and time slot assignment. So its scalability is normally not as good as that of a contention-based protocol." Every node that dies or joins, and every change in traffic, means a new schedule. LEACH rebuilds its clusters and schedules every round.

The synchronisation itself. Beacons, guards, and a node that must never miss its beacon.

These are why most sensor MACs are hybrids: S-MAC shares a coarse schedule of listen and sleep and contends inside it ([S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step]); 802.15.4's superframe offers a contention access period and, for devices that need them, guaranteed time slots ([The 802.15.4 Superframe and Guaranteed Time Slots]).

Distinctions

TDMACSMA
AccessOwn slot in a repeating frameCompete when there is data
CollisionsNone within the schedulePossible
Idle listeningNone: wake only for own, receiving and beacon slotsUnless duty-cycled, all the time
NeedsSlot allocation and synchronisationNothing shared
Load changesSchedule must be rebuiltAdapts at once
LatencyWait for the slotLow at light load
munotes.in329

TDMA and Schedule-based MAC

Centralised allocationDistributed allocation
Who decidesA cluster head or the sink (LEACH)The nodes, by negotiating (SMACS)
StrengthSimple, collision-free within the clusterNo master; survives node loss
WeaknessThe decider must know the topology and stay awakeSlower to converge; subnets may not connect

What it does not mean

TDMA does not mean every node needs its own slot. Nodes three or more hops apart can share one; the chain of seven needed three.

A schedule is not free once built. Clocks drift from the moment they are set, and the guards and beacons that absorb the drift cost energy every frame.

Rare beacons are not always cheaper. The members save a little; the head, listening across ever wider guards, pays far more.

No idle listening does not mean no waste. An owned slot with nothing to send is channel wasted, and a reading waiting for its slot is time wasted.

Quick revision

  • TDMA: frames of slots; each node transmits only in its own slot: no collisions, no idle listening; radio off otherwise.
  • LEACH: the head "sets up a TDMA schedule"; members send "at most once per frame"; the frame grows with the cluster; the head stays awake for all slots; clusters separated by DSSS codes.
  • Two-hop rule: nodes within two hops need different slots; farther ones may reuse a slot. A chain of 7 needs 3 slots (1, 2, 3, 1, 2, 3, 1); a grid needs at least 5.
  • SMACS: distributed, pairs of slots on a random fixed frequency, no master.
  • Guard = relative drift × time since the last sync: 80 ppm × 60 s = 4.8 ms each side, against a 1.6 ms frame.
  • Program: member about 80 mA-s an hour (mostly sleep current) against 67,679.9 for always-on CSMA; head cheapest with a beacon near every 10 s (115.7), worst at 10 minutes (1,175.5, slots of 98 ms).
  • Costs: latency, wasted slots, poor adaptability, synchronisation. Hence hybrids: S-MAC, 802.15.4's GTS.

Test yourself

1. What is TDMA, and why does it save energy in a sensor network? Time is divided into repeating frames and each frame into slots, and each node is allocated a slot in which alone it transmits. Because interfering nodes never share a slot, there are no collisions and no contention, and each node knows exactly when it must transmit or receive, so it can keep its radio off at all other times and avoid idle listening and overhearing.

2. Explain the two-hop rule for reusing slots, with an example. Two nodes that are neighbours cannot share a slot, and two nodes with a common neighbour cannot either, because that neighbour would receive both transmissions at once. So nodes within two hops of each other need different slots, while nodes three or more hops apart can share one. In a chain A to G, slots 1, 2, 3, 1, 2, 3, 1 are enough: B and E share slot 2, and no node hears both.

munotes.in330

TDMA and Schedule-based MAC

3. Why does a TDMA schedule need guard times? Compute one. Clocks drift, so a sender's slot may start a little earlier or later than the receiver expects. A guard on each side of the slot, as wide as the worst drift since the last synchronisation, absorbs this. With clocks up to 40 ppm out each (80 ppm between two) and a beacon once a minute, the guard is 80 ppm × 60 s = 4.8 ms, three times the 1.6 ms of a 50-byte frame.

4. How does the beacon interval affect the energy of the members and of the cluster head? Members wake early for each beacon by a guard proportional to the interval, but hear proportionally fewer beacons, so their guard cost per hour is fixed and only the beacons' airtime falls with longer intervals. The head listens across the guards of every slot, which widen with the interval, so its cost rises steeply for long intervals, while very short intervals cost it many beacon transmissions. In the program the head spent least with a beacon every 10 s.

5. What are the disadvantages of TDMA in a wireless sensor network? Readings wait for their slots, adding latency at every hop; a slot is wasted when its owner has nothing to send; schedules must be rebuilt whenever nodes join, leave or change their traffic, so it scales and adapts poorly; and all nodes must be kept synchronised, which costs beacons and guard times.

6. How does LEACH use TDMA and CDMA together? Inside each cluster the cluster head builds a TDMA schedule and sends it to the members, who transmit in their own slots and turn their radios off otherwise, so there are no collisions within the cluster. Between clusters, whose slots overlap in time, each cluster uses a different direct-sequence spreading code, which keeps neighbouring clusters from interfering but requires tight timing synchronisation.

Contents This chapter on its own page

munotes.in331

Chapter Fifty

Duty Cycling: Preamble Sampling, B-MAC and X-MAC

Syllabus topic Module 1, "Medium Access Control (MAC) in WSN: Fundamentals of MAC protocols for sensor networks" (and the paired practical, "Implement duty-cycling MAC protocol and evaluate energy conservation in sensor networks")

In one line

A duty-cycled radio sleeps most of the time and wakes briefly to check the channel; with low power listening the sender makes itself heard by sending a preamble longer than the check interval (B-MAC), and X-MAC cuts that preamble into short packets naming the receiver, so that other nodes go back to sleep and the receiver can stop it early.

In the wording a student can write in an examination: duty cycling keeps the radio asleep except for short, periodic wake-ups; the duty cycle is the fraction of time it is on. A sleeping node cannot receive, so sender and receiver must meet (the rendezvous problem). Synchronous protocols (S-MAC, T-MAC) agree a common schedule of listen and sleep. Asynchronous protocols use low power listening (LPL), also called preamble sampling: every node wakes once per check interval and samples the channel; a sender transmits a preamble at least as long as the check interval before the data, so the receiver is sure to wake during it, stay awake and receive the data. B-MAC is the classic LPL protocol: clear channel assessment, backoffs, optional acknowledgements and LPL, with the check interval and preamble exposed for tuning. LPL's weaknesses are overhearing (every neighbour that wakes during the preamble stays awake until the data), a preamble longer than needed (the receiver wakes half-way through on average), latency of at least the preamble at every hop, and no support for packet radios such as the CC2420. X-MAC sends a strobed preamble: short preamble packets containing the target's address, separated by gaps. A node that is not the target goes back to sleep after one; the target answers with an early acknowledgement in a gap, and the sender sends the data at once.

Why sleep, and the problem sleeping creates

[MAC Protocols for Sensor Networks: The Job and Where the Energy Goes] found idle listening to be almost the whole energy bill of a radio left on, and [How Long a Node Lasts: The Energy Budget Worked Out] gave the arithmetic of the cure. A node awake a fraction d of the time draws on average d × I(awake) + (1 - d) × I(asleep); for the Telos parts at 1 per cent that is 0.01 × 19.3 + 0.99 × 0.022 = 0.21478 mA, and 2,500 mAh lasts about 1.3 years instead of 5.4 days.

Sleeping creates the rendezvous problem: a frame sent while its receiver sleeps is lost. X-MAC's authors divide the answers in two. Synchronised protocols, such as S-MAC and T-MAC, "negotiate a schedule that specifies when nodes are awake and asleep within a frame", so neighbours wake together; chapters 51 to 53 take S-MAC apart. The other family keeps no schedule at all: asynchronous protocols such as B-MAC and WiseMAC "rely on low power listening (LPL), also called preamble sampling, to link together a sender with data to a receiver who is duty cycling."

munotes.in332

Duty Cycling: Preamble Sampling, B-MAC and X-MAC

Low power listening: B-MAC

B-MAC, from Berkeley in 2004, was built to be small and configurable. Its authors describe the core in a few sentences: "Each time the node wakes up, it turns on the radio and checks for activity. If activity is detected, the node powers up and stays awake for the time required to receive the incoming packet. After reception, the node returns to sleep. If no packet is received (a false positive), a timeout forces the node back to sleep."

The sender's side follows from the receiver's. A receiver checks the channel once per check interval, so a sender must keep the channel busy for at least that long before its data: "If the channel is checked every 100 ms, the preamble must be at least 100 ms long for a node to wake up, detect activity on the channel, receive the preamble, and then receive the message." B-MAC offers eight modes, check intervals of 10, 20, 50, 100, 200, 400, 800 and 1600 ms, and lets a protocol set its own.

The cost of listening moves from the receivers to the sender, which is the point: in a quiet network almost nobody sends, and the nodes' only cost is the brief check. "Idle listening occurs when the node wakes up to sample the channel and there is no activity", and the check is short. Its accuracy matters: a check that mistakes noise for a preamble keeps the node awake for nothing, so "Accurate channel assessment (CCA) is critical to achieving low power operation with this method." B-MAC estimates the noise floor and treats a reading as activity only if it stands out from it.

B-MAC is also small. Its paper compares code sizes on the Mica2: B-MAC with low power listening and acknowledgements takes 4,386 bytes of ROM and 172 of RAM; S-MAC takes 6,274 and 516.

The check interval: B-MAC's own model, run

A short check interval makes every node wake often; a long one makes every packet carry a long preamble, which the sender transmits and every neighbour hears. The best interval depends on how many neighbours there are and how often they send. B-MAC's paper models this exactly, and the program runs its equations with its own Mica2 numbers: 19.2 kbit/s (416 microseconds a byte), 20 mA to transmit and 15 to receive, 17.3 microjoules per channel sample, 0.030 mA asleep, a 36-byte packet, a reading every five minutes costing 1.1 s of sensing at 20 mA, and a 3 V, 2,500 mAh battery. Following the paper, every one of a node's n neighbours is counted as receiving every packet, preamble and all.

munotes.in333

Duty Cycling: Preamble Sampling, B-MAC and X-MAC

# B-MAC's lifetime model (Polastre, Hill and Culler 2004, section 4.1), with
# the paper's own Mica2/CC1000 numbers from its Tables 2 and 3. Energies are
# in mW (mJ per second). The preamble must last at least one check interval.
from math import ceil

V, T_BYTE = 3.0, 416e-6                 # volts; seconds per byte, sent or received
C_TX, C_RX, C_SLEEP = 20, 15, 0.030     # mA
E_SAMPLE = 17.3e-3                      # mJ for one channel sample (their Figure 3)
T_SAMPLE = 350e-6 + 1.5e-3 + 250e-6 + 350e-6   # s the radio is up for one sample
T_DATA, C_DATA = 1.1, 20                # sensing: 1.1 s at 20 mA per reading
PACKET, RATE = 36, 1 / 300              # bytes; one reading every five minutes

def power(check, n):
    preamble = ceil(check / T_BYTE)                     # bytes, the constraint
    t_tx = RATE * (preamble + PACKET) * T_BYTE           # equation 3
    t_rx = n * RATE * (preamble + PACKET) * T_BYTE       # equation 4, every neighbour's
    t_listen = T_SAMPLE / check                          # equation 5
    t_d = T_DATA * RATE                                  # equation 2
    t_sleep = 1 - t_rx - t_tx - t_d - t_listen           # equation 6
    return (t_tx * C_TX * V + t_rx * C_RX * V + E_SAMPLE / check
            + t_d * C_DATA * V + t_sleep * C_SLEEP * V)

CHECKS = (0.010, 0.020, 0.050, 0.100, 0.200, 0.400, 0.800, 1.600)   # B-MAC's 8 modes
SIZES = (5, 10, 20, 50)
print("check    " + "".join("   n = %-3d" % n for n in SIZES) + "  (mW)")
for check in CHECKS:
    print("%5d ms " % (check * 1000) + "".join("%10.3f" % power(check, n) for n in SIZES))
for n in SIZES:
    best = min(CHECKS, key=lambda c: power(c, n))
    years = 2500 * V / power(best, n) / 24 / 365                    # 2,500 mAh at 3 V
    print("n = %2d: best check interval %4d ms, %.3f mW, %.2f years"
          % (n, best * 1000, power(best, n), years))
check       n = 5     n = 10    n = 20    n = 50   (mW)
   10 ms      2.042     2.061     2.099     2.213
   20 ms      1.197     1.224     1.277     1.435
   50 ms      0.713     0.762     0.860     1.153
  100 ms      0.590     0.676     0.848     1.366
  200 ms      0.599     0.760     1.082     2.048
  400 ms      0.746     1.057     1.678     3.543
  800 ms      1.104     1.714     2.935     6.597
 1600 ms      1.852     3.061     5.479    12.734
n =  5: best check interval  100 ms, 0.590 mW, 1.45 years
n = 10: best check interval  100 ms, 0.676 mW, 1.27 years
n = 20: best check interval  100 ms, 0.848 mW, 1.01 years
n = 50: best check interval   50 ms, 1.153 mW, 0.74 years
munotes.in334

Duty Cycling: Preamble Sampling, B-MAC and X-MAC

Reading it. Every column falls and then rises: too short an interval and the node spends its energy sampling (2.042 mW at 10 ms with 5 neighbours), too long and it spends it on preambles, its own and its neighbours' (1.852 mW at 1600 ms). With 10 neighbours the best is 100 ms, 0.676 mW, and the battery lasts 1.27 years; B-MAC's default check interval on the Mica2 was 100 ms.

The more neighbours, the shorter the best interval. Each neighbour's packets must be heard, preamble and all, so a crowded neighbourhood favours short preambles: with 50 neighbours the best interval falls to 50 ms and the lifetime to 0.74 years. The paper's own reading of its figure is that "a check interval of 50 ms is optimal for a neighborhood size of 20, but if the neighborhood size is only 5, a check interval of 100 ms is optimal". The model agrees for 5; for 20 it puts 100 ms and 50 ms within 1.5 per cent of each other (0.848 against 0.860 mW), close enough that the preamble the implementation actually sends (271 bytes at 100 ms by default, against the 241-byte minimum the model uses) decides it.

Too short is worse than too long, when neighbours are few. With 5 neighbours, halving the best interval to 50 ms costs 21 per cent more power (0.713 against 0.590 mW), doubling it to 200 ms only 1.5 per cent (0.599). The authors saw the same: "the penalty for more idle listening than required by the traffic pattern, left of the maximum lifetime point in Figure 6, is much more severe than the penalty for sending packets that are longer than necessary."

What is wrong with the long preamble

X-MAC's authors list LPL's faults, and each is visible in the upper half of the figure below.

Overhearing. "Non-target receivers who wake and sample the medium while a preamble is being sent must wait until the end of the extended preamble before finding out that they are not the target and should go back to sleep." Every neighbour pays for every packet, which is the n in B-MAC's equation 4.

A preamble longer than it needs to be. "The sender sends the entire preamble even though, on average, the receiver will wake up half way through the preamble." There is no way for the sender to know the receiver is already awake.

Latency. The receiver waits for the whole preamble, so every hop takes at least one check interval, and over many hops the delays add up.

Packet radios cannot do it. The CC2420, on the TelosB and MICAz, sends packets, not raw bits: "with these radios the application cannot send a preamble of arbitrary length. This precludes the use of LPL protocols that depend on an extended preamble." B-MAC was written for the Mica2's bit-streaming CC1000.

munotes.in335

Duty Cycling: Preamble Sampling, B-MAC and X-MAC

Two halves. Low power listening: the sender's radio is on for a preamble a whole check interval long and then the data; the target wakes mid-way and stays on to the end; a neighbour wakes earlier and also stays on to the end. X-MAC: the sender sends short strobes with listening gaps between them; the target wakes, hears a strobe, sends an ACK in the gap and receives the data; a neighbour wakes, hears one strobe and sleeps

Figure 50.1 One packet with a long preamble and with a strobed one (schematic)

X-MAC: the strobed preamble

X-MAC, from the University of Colorado in 2006, keeps the asynchronous idea and fixes the preamble in two steps.

Name the target in the preamble. The long preamble is divided "into a series of short preamble packets, each containing the ID of the target node". A node that wakes and hears one "looks at the target node ID that is included in the packet. If the node is not the intended recipient, the node returns to sleep immediately". Overhearing shrinks from a whole preamble to one short packet.

Leave gaps, and let the target answer. Between the short packets the sender pauses to listen. "These gaps enable the receiver to send an early acknowledgment packet back to the sender by transmitting the acknowledgment during the short pause between preamble packets. When a sender receives an acknowledgment from the intended receiver, it stops sending preambles and sends the data packet." Because the sender alternates a short packet with a short wait, the authors call it a strobed preamble. The preamble now ends when the receiver wakes, half-way on average, instead of always running to the end.

Two refinements complete it. A second sender waiting for the same receiver that overhears the early acknowledgement "will back-off a random amount and then send its data without a preamble", and the receiver stays awake briefly after each packet to catch it. And a node may adapt its own sleep period to the traffic it sees. Short packets suit every radio, the CC2420 included.

One packet, both ways

The second program charges one packet from a sender to a sleeping neighbour, with 8 other neighbours in range, using X-MAC's own measurements on the TelosB: 86.2 mW transmitting, 96.6 mW receiving, 1.98 ms for a short preamble packet, 1.84 ms for the early acknowledgement, 3.8 ms for the data. With LPL the preamble lasts the whole check interval, and the target and every other neighbour wake at a random moment in it and stay awake to the end. With X-MAC the target wakes at a random moment (20,000 of them are drawn), waits for the next strobe, answers and receives; each other neighbour hears one strobe and sleeps.

# One packet sent to a duty-cycled neighbour, with plain low power listening
# (a preamble as long as the check interval) and with X-MAC's strobed preamble
# (short packets naming the target, with gaps for an early ACK). The TelosB
# constants are X-MAC's own (its Table 3); every other neighbour also wakes
# once during the preamble. Energy in microjoules (mW x ms), latency in ms.
import random

P_TX, P_RX = 86.2, 96.6          # mW
STROBE, ACK, DATA = 1.98, 1.84, 3.8   # ms: short preamble packet, ACK, data packet
PERIOD = STROBE + ACK            # one strobe and the gap that fits an early ACK
OTHERS = 8                       # neighbours that are not the target

def lpl(check):
    """The sender sends the whole preamble; a waking node, target or not, waits
    for the rest of it (half, on average) and then the data."""
    wait = check / 2
    return {"sender": (check + DATA) * P_TX,
            "target": (wait + DATA) * P_RX,
            "others": OTHERS * (wait + DATA) * P_RX,
            "latency": check + DATA}

def xmac(check, woke):
    """`woke` is when, in the check interval, the target wakes. The sender
    strobes until the target has heard a whole strobe; others hear one, sleep."""
    strobes = int(woke // PERIOD) + 2          # the one under way, then a whole one
    until_next = PERIOD - woke % PERIOD         # listening before the next strobe
    return {"sender": strobes * (STROBE * P_TX + ACK * P_RX) + DATA * P_TX,
            "target": (until_next + STROBE + DATA) * P_RX + ACK * P_TX,
            "others": OTHERS * (PERIOD / 2 + STROBE) * P_RX,
            "latency": strobes * PERIOD + DATA}

rnd = random.Random(50)
print("check   method   sender  target  8 others  total uJ  latency ms")
for check in (50, 100, 500):
    runs = [xmac(check, rnd.uniform(0, check)) for _ in range(20000)]
    x = {k: sum(r[k] for r in runs) / len(runs) for k in runs[0]}
    for name, r in (("LPL", lpl(check)), ("X-MAC", x)):
        total = r["sender"] + r["target"] + r["others"]
        print("%4d ms  %-6s %8.0f %7.0f %9.0f %9.0f %11.1f"
              % (check, name, r["sender"], r["target"], r["others"], total, r["latency"]))
munotes.in336

Duty Cycling: Preamble Sampling, B-MAC and X-MAC

check   method   sender  target  8 others  total uJ  latency ms
  50 ms  LPL        4638    2782     22257     29676        53.8
  50 ms  X-MAC      3136     902      3006      7044        34.6
 100 ms  LPL        8948    5197     41577     55721       103.8
 100 ms  X-MAC      5434     902      3006      9342        59.8
 500 ms  LPL       43428   24517    196137    264081       503.8
 500 ms  X-MAC     23752     901      3006     27659       260.6

Reading it. At a 100 ms check interval, one packet costs the neighbourhood 55,721 microjoules with LPL and 9,342 with X-MAC, about six times less. The largest single saving is overhearing: the 8 other neighbours spend 41,577 microjoules staying awake through the long preamble, and 3,006 hearing one strobe each.

The target and the sender both gain. The target no longer waits for the end of the preamble (902 microjoules against 5,197), and the sender, stopped by the early acknowledgement half-way on average, spends 5,434 instead of 8,948.

munotes.in337

Duty Cycling: Preamble Sampling, B-MAC and X-MAC

Latency halves. A hop costs 103.8 ms with LPL and 59.8 with X-MAC at a 100 ms interval: the preamble ends when the receiver wakes, not when the interval does.

The longer the interval, the bigger the gain. At 500 ms LPL's cost grows almost in proportion (264,081 microjoules) while X-MAC's target and neighbours pay the same small amounts as before, so the ratio rises to about 9.5. Long check intervals, which make the idle node cheap, are exactly where the long preamble hurts most.

X-MAC is not free. Its receivers must stay awake at each wake-up long enough to catch a strobe, across a gap (its Table 3 gives a listen time of at least 1.84 ms), where an LPL receiver needs only a brief sample; and the model here, by counting every neighbour as waking during the strobes, overstates X-MAC's overhearing, since strobing usually stops before all of them wake.

Distinctions

Synchronous duty cyclingAsynchronous duty cycling
RendezvousNeighbours agree when to be awakeThe sender's preamble outlasts the receiver's sleep
NeedsSchedules and synchronisationNothing shared
Who paysEveryone, for schedule upkeep and common listen timeThe sender, with a long preamble, and its overhearers
ExamplesS-MAC, T-MACB-MAC, WiseMAC, X-MAC
Low power listening (B-MAC)Strobed preamble (X-MAC)
PreambleOne long transmission, a whole check intervalShort packets naming the target, with gaps
Non-target that wakesStays awake to the end of the preamble and dataSleeps after one short packet
TargetWaits for the end of the preambleSends an early ACK; data follows at once
Latency per hopAt least one check intervalAbout half a check interval on average
RadiosBit-streaming (CC1000)Any, including packet radios (CC2420)
Short check intervalLong check interval
Idle nodeWakes often: costlyWakes rarely: cheap
Each packetShort preambleLong preamble, heard by every neighbour
SuitsMany neighbours, frequent trafficFew neighbours, rare traffic

What it does not mean

A low duty cycle is not free energy. The checks, the preambles and the overhearing are all paid for; B-MAC's own model shows the best check interval moving with the neighbourhood.

The preamble is not wasted by design. It is the price of not keeping schedules; X-MAC shortens it, but a sender still strobes until its receiver wakes.

LPL does not remove idle listening. It shortens each episode to a brief channel check, which is why an accurate CCA matters.

X-MAC does not need synchronised clocks. Each node keeps its own sleep schedule; the strobes find the receiver whenever it wakes.

Quick revision

  • Duty cycle d: average current d × I(awake) + (1 - d) × I(asleep); 1 per cent on Telos parts, about 1.3 years.
  • Rendezvous: synchronous (S-MAC, T-MAC) or asynchronous (B-MAC, WiseMAC, X-MAC).
  • LPL / preamble sampling (B-MAC): wake every check interval, sample; preamble at least one check interval; check intervals 10 to 1600 ms; accurate CCA against false positives.
  • B-MAC's model with its numbers: 10 neighbours, best 100 ms, 0.676 mW, 1.27 years; 50 neighbours, 50 ms, 0.74 years. Too short costs more than too long.
  • LPL's faults: overhearing, excess preamble (receiver wakes half-way on average), latency at least the check interval per hop, no packet radios.
  • X-MAC: short preamble packets with the target ID; non-targets sleep; gaps for an early ACK; data at once; waiting senders skip the preamble.
  • One packet at 100 ms (program): LPL 55,721 microjoules, X-MAC 9,342; overhearing 41,577 against 3,006; latency 103.8 against 59.8 ms.
munotes.in338

Duty Cycling: Preamble Sampling, B-MAC and X-MAC

Test yourself

1. What is low power listening, and why must the preamble be at least as long as the check interval? Each node sleeps and wakes once per check interval to sample the channel briefly; if it detects activity it stays awake to receive, otherwise it sleeps again. Because nodes are not synchronised, a sender cannot know when its receiver will next wake, so it sends a preamble before the data that lasts at least one check interval; the receiver is then certain to wake during the preamble, detect it and stay awake for the data.

2. How does the check interval trade energy, and what did B-MAC's model give as the best? A short check interval makes idle nodes wake and sample often; a long one makes every packet's preamble long, costing the sender and every neighbour that hears it. With B-MAC's Mica2 numbers, a reading every five minutes and 10 neighbours, the model's best interval is 100 ms at 0.676 mW, a lifetime of about 1.27 years on 2,500 mAh at 3 V; with 50 neighbours it falls to 50 ms.

3. What are the weaknesses of the long preamble? Every neighbour that wakes during it must stay awake until the end to learn whether the packet is for it (overhearing); the sender always sends the whole preamble although the receiver wakes half-way on average; each hop is delayed by at least the preamble; and packet radios such as the CC2420 cannot send a preamble of arbitrary length.

4. Explain X-MAC's strobed preamble. The sender replaces the long preamble with a series of short preamble packets, each carrying the target's address, with short gaps between them in which it listens. A node that wakes and hears a strobe for another node goes back to sleep at once. The target, on hearing a strobe addressed to it, sends an early acknowledgement in the next gap, and the sender immediately sends the data, so the preamble ends as soon as the receiver is awake.

munotes.in339

Duty Cycling: Preamble Sampling, B-MAC and X-MAC

5. With X-MAC's TelosB numbers and a 100 ms check interval, compare the cost of one packet to a neighbourhood of a sender, a target and 8 other nodes. With LPL the total is about 55,721 microjoules, 41,577 of it the 8 other nodes overhearing the preamble; with X-MAC it is about 9,342, the others spending 3,006. The hop latency falls from 103.8 ms to about 59.8 ms.

6. Distinguish synchronous and asynchronous duty-cycling MAC protocols, with examples. Synchronous protocols such as S-MAC and T-MAC make neighbours agree on a common schedule of listen and sleep, so senders transmit when receivers are known to be awake, at the cost of schedule exchange and clock synchronisation. Asynchronous protocols such as B-MAC and X-MAC keep no common schedule; each node wakes on its own, and a sender uses a long or strobed preamble to catch the receiver when it wakes.

Contents This chapter on its own page

munotes.in340

Chapter Fifty-One

S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step

Syllabus topic Module 1, "Medium Access Control (MAC) in WSN: Sensor-MAC (S-MAC) / Sensor-MAC case study"

In one line

S-MAC makes every node sleep for most of each frame and wake for a short listen interval, and makes neighbours share the same schedule so that they are awake together, forming virtual clusters; it trades latency and per-node fairness for energy.

In the wording a student can write in an examination: S-MAC (Sensor-MAC), proposed by Ye, Heidemann and Estrin in 2002, is a contention-based MAC protocol designed for sensor networks, whose main goal is to cut energy consumption from all four sources of waste (idle listening, collisions, overhearing and control overhead) while supporting scalability and collision avoidance. It has three main components: periodic listen and sleep, collision and overhearing avoidance, and message passing, to which the 2004 version adds adaptive listening. Time is divided into frames, each a listen period followed by a sleep period, during which the radio is off; the duty cycle is the listen interval divided by the frame length. Neighbours synchronise their schedules by broadcasting SYNC packets containing the sender's address and the time until its next sleep. A node that hears no schedule when it starts becomes a synchronizer and chooses its own; a node that hears one becomes a follower and adopts it. Nodes on a common schedule form a virtual cluster; a border node between two virtual clusters follows both schedules. Timestamps are relative and the listen interval is long compared with clock drift, so synchronisation can be loose, but SYNC packets are repeated periodically, and nodes periodically listen for a whole synchronisation period to discover neighbours on other schedules. The price is latency: a packet waits for its receiver's next listen interval, at every hop.

What S-MAC is for

S-MAC was designed at USC's Information Sciences Institute for networks that are idle for long periods and then suddenly busy. Its authors' abstract states the priorities that set it apart from 802.11: "energy conservation and self-configuration are primary goals, while per-node fairness and latency are less important."

The design rests on assumptions about the application, stated in the 2002 paper. The nodes are many and deployed casually, so they "must therefore self-configure". The network serves one application, "thus rather than node-level fairness (like in the Internet), we focus on maximizing system-wide application performance." Data is processed in the network, "as whole messages at a time in store-and-forward fashion". And "applications will have long idle periods and can tolerate some latency."

From these come the three components: "periodic listen and sleep, collision and overhearing avoidance, and message passing". This chapter takes the first, which attacks idle listening; [S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing] takes the other two; [S-MAC: Latency, Adaptive Listening and the Energy Saved] takes the cost in delay and the adaptive listening that reduces it.

munotes.in341

S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step

Periodic listen and sleep

"Each node sleeps for some time, and then wakes up and listens to see if any other node wants to talk to it. During sleeping, the node turns off its radio, and sets a timer to awake itself later."

The journal paper names the parts. "We call a complete cycle of listen and sleep a frame." The listen interval is fixed by the radio and the MAC ("e.g., the radio bandwidth and the contention window size"), and "The duty cycle is defined as the ratio of the listen interval to the frame length." To change the duty cycle, S-MAC changes the sleep, not the listen.

Worked example. On the Mica motes the listen interval was 115 ms. A 10 per cent duty cycle therefore means a frame of 0.115 / 0.10 = 1.15 s, as the paper states; a 1 per cent duty cycle, a frame of 0.115 / 0.01 = 11.5 s, most of it asleep.

Why is the listen interval so long, when a TDMA slot can be a few milliseconds? Because S-MAC contends for the channel inside it, with a SYNC part and a data part each divided into contention slots (15 and 31 on the Mica), and because a long listen tolerates clock drift. The paper contrasts it with TDMA: "Compared to TDMA schemes with very short time slots, S-MAC requires much looser time synchronization."

Choosing a schedule: synchronizers and followers

A node could sleep on its own timetable, but then a sender would need to know every neighbour's. S-MAC prefers agreement: "All nodes are free to choose their own listen/sleep schedules. However, to reduce control overhead, we prefer neighboring nodes to synchronize together. That is, they listen at the same time and go to sleep at the same time." Each node keeps a schedule table of its neighbours' schedules and follows these steps (journal version):

  1. Listen first. A new node listens for at least one synchronisation period. If it hears no schedule, it chooses its own, starts following it and announces it in a SYNC packet. The 2002 paper calls such a node a synchronizer, "since it chooses its schedule independently and other nodes will synchronize with it."
  2. Follow if you hear one. "If the node receives a schedule from a neighbor before choosing or announcing its own schedule, it follows that schedule by setting its schedule to be the same." Such a node is a follower; it waits a random delay before rebroadcasting the schedule, "so that multiple followers triggered from the same synchronizer do not systematically collide".
  3. A second schedule arrives. If the node has no other neighbours, it drops its own schedule and follows the new one. If it already shares its schedule with at least one neighbour, "it adopts both schedules by waking up at the listen intervals of the two schedules."
munotes.in342

S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step

In a network where every node hears every other, the first node to finish listening sets the schedule for all. In a multi-hop network, nodes far apart may choose independently, and the nodes between them end up on two schedules.

Top: four nodes C, A, B and D in a line, C and A ringed as schedule 1, B and D as schedule 2. Middle: their timelines, C listening at schedule 1's times, D at schedule 2's, and the border nodes A and B at both. Bottom: one listen interval divided into a part for SYNC with 15 contention slots and a part for RTS with 31, followed by sleep

Figure 51.1 Virtual clusters and border nodes, after the paper's Fig. 2, and one listen interval

Virtual clusters and border nodes

Nodes that share a schedule form a virtual cluster. It is not a cluster in LEACH's sense: there is no head and nobody relays for anybody. The journal: "Unlike clustering protocols, S-MAC does not require coordination through cluster heads. Instead, nodes form virtual clusters around common schedules, but communicate directly with peers."

In the figure, C and A follow schedule 1, B and D schedule 2, and A and B are neighbours. Any node can still reach any neighbour, whatever its schedule: "if node A wants to talk to node B, it waits until B is listening." The border nodes A and B have a choice. Adopting both schedules means waking for both listen intervals, so that a broadcast needs sending only once; "The disadvantage is that these border nodes have less time to sleep and consume more energy than others." Adopting only one keeps their duty cycle, but a broadcast must be sent twice, once in each schedule's listen time.

Keeping neighbours in step

Clocks drift, and a schedule agreed today is wrong tomorrow unless it is refreshed. S-MAC uses two defences. First, "all exchanged timestamps are relative rather than absolute": a SYNC says how long until the sender sleeps, not at what hour. Second, the listen interval is long compared with the drift. The journal's measurement: "the clock drift between two nodes does not exceed 0.2 ms per second."

Worked example. At 0.2 ms per second, two nodes drift apart by 0.2 × 10 = 2 ms in one 10-second synchronisation period, against a listen interval of 115 ms. Without any SYNC at all, the drift would use up the whole listen interval only after 115 / 0.2 = 575 seconds, nearly ten minutes. Compare a TDMA slot of a few milliseconds ([TDMA and Schedule-based MAC]).

SYNC packets. Each node still broadcasts its schedule periodically. "The SYNC packet is very short, and includes the address of the sender and the time of its next sleep." The time is counted "relative to the moment that the sender starts transmitting the SYNC packet", and the receiver subtracts the packet's transmission time before adjusting its own timer. The period between a node's SYNC packets is the synchronisation period, 10 s in the Mica implementation. SYNC packets are sent in the first part of the listen interval, data exchanges in the second.

munotes.in343

S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step

Neighbour discovery. A node may still miss a neighbour: its SYNC collided, or was delayed, or the two schedules never overlap. So "each node periodically listens for the whole synchronization period". The paper adds a warning: "Since the energy cost is high during the neighbor discovery, it should not be performed too often. In our current implementation, the synchronization period is 10 s, and a node performs neighbor discovery every 2 min if it has at least one neighbor." The program shows how high.

An idle S-MAC node's day, charged

The program takes a node with no data to send and charges its radio for a day, with the journal's own Mica implementation: a 115 ms listen interval; a SYNC every 10 s, a 10-byte control packet sent at 20 kbit/s with Manchester coding, which doubles every bit; neighbour discovery every 2 minutes, listening for the whole 10-second period; and the TR3000 radio's 14.4 mW receiving, 36 mW transmitting and 15 microwatts asleep. It compares an ordinary node, a border node on two schedules, a node that discovers only every 20 minutes, and one that never does, at duty cycles of 1 and 10 per cent.

# Where an S-MAC node's radio time goes when there is no traffic at all, from
# the parameters of the S-MAC implementation on Mica motes (Ye, Heidemann and
# Estrin 2004): a 115 ms listen interval, a SYNC every 10 s, neighbour discovery
# (listening for a whole synchronisation period) every 2 minutes, 10-byte
# control packets at 20 kbit/s with Manchester coding (two bits per data bit).
LISTEN, SYNC_EVERY, DISCOVER_EVERY = 0.115, 10.0, 120.0       # seconds
SYNC_AIR = 10 * 8 * 2 / 20_000                                 # s on air per SYNC
P_RX, P_TX, P_SLEEP = 14.4, 36.0, 0.015                        # mW: TR3000 radio

def day(duty, schedules=1, discover_every=DISCOVER_EVERY):
    """Fraction of time listening and transmitting, and mJ per day."""
    listen = min(1.0, duty * schedules)
    # discovery: one whole synchronisation period awake, of which the node's
    # scheduled listening would have covered the fraction `listen` anyway
    discovery = SYNC_EVERY / discover_every * (1 - listen) if discover_every else 0
    tx = SYNC_AIR / SYNC_EVERY * schedules          # its own SYNCs, one per schedule
    on = listen + discovery + tx
    energy = 86400 * ((listen + discovery) * P_RX + tx * P_TX + (1 - on) * P_SLEEP)
    return listen, discovery, tx, energy

print("duty  node            listen  discovery  SYNC tx   radio on   J per day")
for duty in (0.01, 0.10):
    for name, kw in (("ordinary", {}), ("border", {"schedules": 2}),
                     ("discover 20 min", {"discover_every": 1200}),
                     ("no discovery", {"discover_every": 0})):
        listen, disc, tx, energy = day(duty, **kw)
        print("%3.0f%%  %-15s %6.2f%% %9.2f%% %8.3f%% %9.2f%% %11.1f"
              % (duty * 100, name, listen * 100, disc * 100, tx * 100,
                 (listen + disc + tx) * 100, energy / 1000))
print("always on:%55.1f" % (86400 * P_RX / 1000))
print("frame length at 1%%: %.1f s; at 10%%: %.2f s" % (LISTEN / 0.01, LISTEN / 0.10))
munotes.in344

S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step

duty  node            listen  discovery  SYNC tx   radio on   J per day
  1%  ordinary          1.00%      8.25%    0.080%      9.33%       118.7
  1%  border            2.00%      8.17%    0.160%     10.33%       132.6
  1%  discover 20 min   1.00%      0.83%    0.080%      1.91%        26.5
  1%  no discovery      1.00%      0.00%    0.080%      1.08%        16.2
 10%  ordinary         10.00%      7.50%    0.080%     17.58%       221.3
 10%  border           20.00%      6.67%    0.160%     26.83%       337.7
 10%  discover 20 min  10.00%      0.75%    0.080%     10.83%       137.4
 10%  no discovery     10.00%      0.00%    0.080%     10.08%       128.1
always on:                                                 1244.2
frame length at 1%: 11.5 s; at 10%: 1.15 s

Reading it. A radio left on costs 1,244.2 J a day. S-MAC at a 1 per cent duty cycle with no discovery costs 16.2 J, about 77 times less: periodic sleep does what it promises. The SYNC packets themselves are negligible, 0.080 per cent of the time.

Neighbour discovery can cost more than the duty cycle. With discovery every 2 minutes, the ordinary node at 1 per cent spends 8.25 per cent of its time listening for new neighbours against 1.00 per cent in its own schedule, and its daily energy rises from 16.2 J to 118.7. The 10 seconds of discovery every 120 are one twelfth of the time, whatever the duty cycle. Discovering every 20 minutes instead brings the bill down to 26.5 J. The paper was right that discovery "should not be performed too often"; with its own settings, a 1 per cent duty cycle becomes, in effect, more than 9.

Border nodes pay twice. Following two schedules doubles the scheduled listening: 2 per cent instead of 1, and at 10 per cent a border node is on 26.83 per cent of the time against 17.58 for an ordinary one.

In battery terms, two AA cells at 3 V and 2,500 mAh hold about 27,000 J. Always on, that lasts about 21.7 days; at 1 per cent with 2-minute discovery, about 227 days; at 1 per cent without discovery, about 1,667 days, some four and a half years, before the other parts of the node and the battery's own shelf life are counted.

Distinctions

SynchronizerFollower
WhenHeard no schedule during its initial listenHeard a neighbour's schedule first
ScheduleChooses its ownAdopts the neighbour's
ThenBroadcasts it in a SYNCRebroadcasts it after a random delay
munotes.in345

S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step

Virtual cluster (S-MAC)Cluster (LEACH)
What binds itA shared listen/sleep scheduleA cluster head
HeadNoneYes, collecting and relaying members' data
CommunicationDirectly between neighboursMembers through the head
At the boundaryBorder nodes follow two schedulesClusters separated by codes
Adopt both schedulesAdopt one schedule
Border node's listeningDoubledUnchanged
A broadcastSent onceSent twice, once per schedule

What it does not mean

Synchronised does not mean globally synchronised. Only neighbours need to agree, and a large network may run several schedules side by side.

A virtual cluster is not a cluster. It has no head and routes nothing; it is only nodes that happen to sleep at the same time.

Sleeping does not make nodes unreachable. A node that knows a neighbour's schedule sends when that neighbour listens, even if their schedules differ.

The duty cycle is not the whole radio-on time. SYNC packets, neighbour discovery and second schedules all add to it; discovery, with the paper's own settings, adds the most.

Quick revision

  • S-MAC (Ye, Heidemann, Estrin, 2002; journal 2004): contention-based, energy first; gives up latency and per-node fairness. Components: periodic listen and sleep, collision and overhearing avoidance, message passing, and (2004) adaptive listening.
  • Frame = listen + sleep; duty cycle = listen interval / frame length; the listen interval is fixed (115 ms on Mica), the sleep varies: 10 per cent, a 1.15 s frame; 1 per cent, 11.5 s.
  • Synchronizer: heard no schedule, chooses one, broadcasts SYNC. Follower: adopts the first schedule heard, rebroadcasts after a random delay. A node on one schedule that hears another adopts both.
  • Virtual cluster: nodes on a common schedule, no head. Border nodes follow two schedules (more energy) or one (broadcast twice).
  • Synchronisation: relative timestamps, long listen interval; drift at most 0.2 ms per second (2 ms per 10 s period); SYNC holds the sender's address and time to its next sleep; SYNC period 10 s; neighbour discovery (a whole period awake) every 2 minutes.
  • Program: always on 1,244.2 J a day; 1 per cent without discovery 16.2 J; with 2-minute discovery 118.7 J (radio on 9.33 per cent).

Test yourself

1. What are the design goals of S-MAC, and what does it trade away? Its primary goals are energy efficiency, by cutting idle listening, collisions, overhearing and control overhead, and self-configuration, with good scalability and collision avoidance. It accepts lower per-node fairness and higher latency, which matter less in a sensor network serving one application that is idle most of the time and can tolerate some delay.

munotes.in346

S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step

2. Explain periodic listen and sleep. Define frame and duty cycle, with a numerical example. Each node sleeps with its radio off for most of the time and wakes periodically to listen for a short, fixed interval in which neighbours can contact it. A complete cycle of listen and sleep is a frame, and the duty cycle is the listen interval divided by the frame length. With a listen interval of 115 ms, a 10 per cent duty cycle gives a frame of 1.15 s, and a 1 per cent duty cycle a frame of 11.5 s.

3. How does a node choose its schedule in S-MAC? It first listens for at least one synchronisation period. If it hears no schedule, it chooses its own, follows it and broadcasts it in a SYNC packet, becoming a synchronizer. If it hears a neighbour's schedule first, it adopts it, becoming a follower, and rebroadcasts it after a random delay. If a node already sharing a schedule with a neighbour later hears a different schedule, it adopts both, waking in both listen intervals; if it has no other neighbours, it simply switches to the new schedule.

4. What is a virtual cluster, and what is a border node? A virtual cluster is a group of neighbouring nodes that follow the same listen/sleep schedule; it has no cluster head, and nodes communicate directly. A border node lies between two virtual clusters and follows both schedules, so it can reach both sides and broadcast once, at the cost of listening twice as long; alternatively it may follow one schedule and send broadcasts twice.

5. How does S-MAC keep neighbours synchronised, and why can the synchronisation be loose? Nodes periodically broadcast SYNC packets giving their address and the time until their next sleep, measured relative to the moment the packet is sent, and receivers adjust their timers. The synchronisation can be loose because the listen interval (115 ms) is long compared with clock drift (at most 0.2 ms per second, so 2 ms over a 10 s synchronisation period), unlike TDMA's short slots.

6. Why is neighbour discovery costly, and how costly is it with the implementation's own settings? Discovery means staying awake for a whole synchronisation period to catch neighbours on other schedules. With a 10 s period every 2 minutes, a node is awake for one twelfth of the time on discovery alone, 8.25 per cent at a 1 per cent duty cycle, raising its radio's daily energy from 16.2 J to 118.7 J. Discovering every 20 minutes brings it to 26.5 J.

Contents This chapter on its own page

munotes.in347

Chapter Fifty-Two

S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing

Syllabus topic Module 1, "Medium Access Control (MAC) in WSN: Sensor-MAC (S-MAC) / Sensor-MAC case study"

In one line

Inside each listen interval S-MAC contends for the channel as 802.11 does, with carrier sense and RTS and CTS; every neighbour that hears the RTS or CTS goes to sleep until the exchange is over, and a long message is sent as a burst of fragments under one RTS and CTS.

In the wording a student can write in an examination: collision avoidance in S-MAC follows IEEE 802.11: a node uses physical carrier sense (listening to the channel) and virtual carrier sense (the NAV, set from the duration field of overheard packets), with its carrier-sense time chosen at random within a contention window; unicast packets use the RTS/CTS/DATA/ACK exchange, and broadcast packets are sent without RTS and CTS. Contention happens only in the receiver's listen interval, which is split into a part for SYNC packets and a part for RTS packets. Overhearing avoidance, inspired by PAMAS but using only the data channel, puts to sleep every node that hears an RTS or CTS addressed to another node; since collisions happen at the receiver, all immediate neighbours of both the sender and the receiver should sleep until the exchange ends, which avoids overhearing the long data packets and ACKs. Message passing sends a long message as a burst of fragments reserved by one RTS and one CTS; every fragment is acknowledged, every fragment and ACK carries the remaining duration, and a lost fragment is resent at once by extending the reservation. This cuts control overhead and message latency at the expense of per-node fairness.

Collision avoidance inside the listen interval

S-MAC takes its contention rules from 802.11 ([Hidden and Exposed Terminals, and RTS and CTS], [CSMA/CA Worked Step by Step]): "S-MAC follows similar procedures, including virtual and physical carrier sense, and the RTS/CTS exchange for the hidden terminal problem". The virtual half is the NAV, set from the duration field every packet carries. The physical half is listening, and "Carrier sense time is randomized within a contention window to avoid collisions and starvations. The medium is determined as free if both virtual and physical carrier sense indicate that it is free."

What sleeping adds is when. A sender contends only when its receiver is listening, so everyone with something for that receiver contends at the start of the same listen interval. The interval is divided in two, the first part for SYNC packets and the second for RTS packets, each with its own contention slots (15 and 31 on the Mica motes). A sender picks a random slot in which to finish its carrier sense; if the channel is still idle at the end of that slot, it transmits.

Two rules keep this cheap. A node that loses "goes to sleep and wakes up when the receiver is free and listening again", rather than staying awake to wait. And a pair that has won keeps going through its sleep period: "After the successful exchange of RTS and CTS, the two nodes will use their normal sleep time for data packet transmission." As in every RTS/CTS protocol, "Broadcast packets are sent without using RTS/CTS. Unicast packets follow the sequence of RTS/CTS/DATA/ACK between the sender and the receiver."

munotes.in348

S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing

Overhearing avoidance: who should sleep?

In 802.11 every node listens to everything. The journal is blunt about the cost: "In 802.11 each node keeps listening to all transmissions from its neighbors in order to perform effective virtual carrier sense. As a result, each node overhears many packets that are not directed to itself." S-MAC borrows the answer of PAMAS, an earlier protocol, but without PAMAS's second radio channel: "S-MAC tries to avoid overhearing by letting interfering nodes go to sleep after they hear an RTS or CTS packet. Since DATA packets are normally much longer than control packets, the approach prevents neighboring nodes from overhearing long DATA packets and following ACKs."

Which nodes should sleep? The paper argues it on a line of six nodes, E, C, A, B, D and F, each hearing only its immediate neighbours, with A sending to B.

Top: the line E, C, A, B, D, F with an arrow from A to B; C and D shaded as sleeping, E and F marked free. Below: A sends an RTS and three data fragments, B a CTS and an ACK after each; C sleeps from the end of the RTS and D from the end of the CTS, for the whole message

Figure 52.1 Who sleeps while A sends to B (after the paper's Fig. 5), and message passing

  • D must sleep. It is B's neighbour; anything it sends would collide with A's data at B. "Remember that collision happens at the receiver."
  • E and F need not. They are out of range of both A and B, and cause no interference.
  • C should sleep too, although it is two hops from B and its own transmission would not disturb B. C could send to E, but E's reply would collide at C with A's transmission, which C hears; and C's transmission could corrupt B's ACK arriving at A. "So C's transmission is simply a waste of energy."

The conclusion: "all immediate neighbors of both the sender and receiver should sleep after they hear the RTS or CTS until the current transmission is over". C learns of the exchange from A's RTS, D from B's CTS, and each sets its NAV from the duration field: "a node should sleep to avoid overhearing if its NAV is not zero. It can wake up when its NAV becomes zero." C's case is [Hidden and Exposed Terminals, and RTS and CTS]'s exposed terminal, and S-MAC settles it the same way the acknowledgement does in 802.11: the exposed node stays quiet, and here it also sleeps.

munotes.in349

S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing

Overhearing is not always waste. The paper notes that some algorithms rely on it "to gather neighborhood information for network monitoring, reliable routing or distributed queries", and S-MAC can be configured to allow it; but it suggests that algorithms that do not need it suit energy-limited networks better, and uses explicit acknowledgements rather than inferring delivery from overheard forwarding.

Message passing

In-network processing works on messages: "A message is the collection of meaningful, interrelated units of data. The receiver usually needs to obtain all the data units before it can perform in-network data processing or aggregation." A long message poses a dilemma. Sent as one long packet, a few corrupted bits force the whole of it to be sent again. Sent as independent small packets, "we have to pay the penalty of large control overhead and longer delay", because each packet contends and runs its own RTS and CTS.

S-MAC's answer is to "fragment the long message into many small fragments, and transmit them in a burst. Only one RTS and one CTS are used. They reserve the medium for transmitting all the fragments." Three details make it work:

  1. Every fragment is acknowledged. "If it fails to receive the ACK, it will extend the reserved transmission time for one more fragment, and re-transmit the current fragment immediately."
  2. Every packet carries the time left. The duration field in each fragment and ACK is the time still needed for all the remaining fragments and ACKs, so "if a node wakes up or a new node joins in the middle of a transmission, it can properly go to sleep no matter if it is the neighbor of the sender or the receiver."
  3. The ACKs guard against hidden terminals. A node that is a neighbour of the receiver only cannot hear the fragments; without frequent ACKs it "may mistakenly infer from its carrier sense that the medium is clear" and start sending, corrupting the transfer at the receiver.

Against 802.11's fragmentation. 802.11 can also fragment, but its RTS and CTS reserve the channel only for the first fragment and its ACK, and each fragment then reserves the next. A neighbour therefore learns only that one more fragment follows, "So it has to keep listening until all the fragments are sent." And when a fragment fails, 802.11, designed for fairness, "must give up the transmission and re-contend for the medium so that other nodes have a chance to transmit", while S-MAC extends and resends at once, with "less contention and a small latency." S-MAC caps the number of extensions in case the receiver has died.

The fairness it gives up. Message passing is unfair by design: a node with a long message holds the channel, and neighbours with a short packet wait. The 2002 paper accepts it: "a node who has more data to send gets more time to access the medium." In a network serving one application, "application-level performance is the goal as opposed to per-node fairness."

munotes.in350

S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing

One message, three ways

The program sends one message of 10 fragments from A to B on the paper's line, where C is A's other neighbour and D is B's. Timings follow the journal's Table I for the Mica motes: 20 kbit/s with Manchester coding, so 0.8 ms per byte; 10-byte control packets (8 ms); fragments of 40 bytes of data, the size used in the paper's two-hop test, plus an 8-byte MAC header (38.4 ms). Powers are the TR3000's: 14.4 mW receiving or listening, 36 mW transmitting, 0.015 mW asleep. It compares a handshake for every fragment, 802.11's fragmentation (one handshake, neighbours listening throughout), and S-MAC's message passing (one handshake, C asleep after the RTS and D after the CTS). Contention time and the gaps between frames are left out of all three.

# One message of 10 fragments from A to B, on a line E - C - A - B - D - F where
# each node hears only its neighbours: C hears A, D hears B. Mica numbers from
# the S-MAC journal paper: 20 kbit/s with Manchester coding (0.8 ms a byte),
# 10-byte control packets, and fragments of 40 bytes plus an 8-byte header.
# Energy in mJ (mW x s); radio powers of the TR3000.
BYTE = 8 * 2 / 20_000                    # s per byte on air
CTRL, FRAG = 10 * BYTE, 48 * BYTE        # an RTS, CTS or ACK; one data fragment
P_RX, P_TX, P_SLEEP = 14.4, 36.0, 0.015  # mW
N = 10                                   # fragments

def exchange(handshakes, neighbours_sleep):
    """handshakes: how many RTS/CTS pairs the message needs. Returns the time on
    air and each node's energy. Without overhearing avoidance, C and D keep
    their receivers on for the whole exchange."""
    a_tx = handshakes * CTRL + N * FRAG               # RTSs and fragments
    b_tx = handshakes * CTRL + N * CTRL               # CTSs and ACKs
    total = a_tx + b_tx
    energy = {"A": a_tx * P_TX + b_tx * P_RX,
              "B": b_tx * P_TX + a_tx * P_RX}
    if neighbours_sleep:          # C sleeps after the RTS, D after the CTS
        energy["C"] = CTRL * P_RX + (total - CTRL) * P_SLEEP
        energy["D"] = 2 * CTRL * P_RX + (total - 2 * CTRL) * P_SLEEP
    else:
        energy["C"] = energy["D"] = total * P_RX
    return total, energy

print("scheme                        time ms     A      B      C      D    total mJ")
for name, handshakes, sleep in (("RTS/CTS for every fragment", N, False),
                                ("802.11 fragmentation", 1, False),
                                ("S-MAC message passing", 1, True)):
    total, e = exchange(handshakes, sleep)
    print("%-28s %8.1f %6.2f %6.2f %6.2f %6.2f %9.2f"
          % (name, total * 1000, e["A"], e["B"], e["C"], e["D"], sum(e.values())))
munotes.in351

S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing

scheme                        time ms     A      B      C      D    total mJ
RTS/CTS for every fragment      624.0  19.01  12.44   8.99   8.99     49.42
802.11 fragmentation            480.0  15.38   8.81   6.91   6.91     38.02
S-MAC message passing           480.0  15.38   8.81   0.12   0.24     24.55

Reading it. A handshake for every fragment keeps the channel busy for 624 ms and costs 49.42 mJ in all. One handshake for the burst, as 802.11's fragmentation and S-MAC both use, saves 9 RTS/CTS pairs: 480 ms and 38.02 mJ. So far S-MAC and 802.11 are equal.

Overhearing avoidance makes the difference. In 802.11 the neighbours C and D listen through all 480 ms, 6.91 mJ each. In S-MAC, C sleeps after the 8 ms RTS and D after the RTS and CTS, and they spend 0.12 and 0.24 mJ. The total falls to 24.55 mJ, about half of the first scheme, and every other neighbour of A or B would save as much again.

The sender and receiver pay the same in all three but the first. Message passing does not change what A and B must send; it removes the repeated handshakes, and overhearing avoidance removes the listening of everyone else. In a dense network, where a node has many neighbours, the second saving dominates.

The paper measured the same pattern. On its two-hop test with two sources, under heavy traffic, "802.11 MAC uses more than twice the energy used by S-MAC", and since idle listening was rare at that load, "S-MAC achieves energy savings mainly by avoiding overhearing and efficiently transmitting long messages." Under light traffic, periodic sleep took over as the main saving ([S-MAC: Latency, Adaptive Listening and the Energy Saved]).

Distinctions

802.11 fragmentationS-MAC message passing
ReservationRTS/CTS cover the first fragment; each fragment reserves the nextOne RTS/CTS reserve the whole message
Duration fieldOne more fragmentAll the remaining fragments and ACKs
NeighboursListen until the last fragmentSleep for the whole message
A fragment lostGive up and contend again (fairness)Extend and resend at once
GoalPer-node fairnessMessage-level latency and energy
PAMASS-MAC overhearing avoidance
SignallingA separate signalling channel (a second radio)In-channel: the RTS and CTS themselves
Who sleepsNodes that would overhearEvery neighbour of sender and receiver, by its NAV
UnicastBroadcast
HandshakeRTS/CTS/DATA/ACKNone: carrier sense only
Overhearing avoidanceNeighbours sleep on RTS or CTSNot applicable: everyone is a receiver

What it does not mean

Overhearing avoidance is not collision avoidance. The NAV stops a neighbour from transmitting; sleeping also stops it from listening. S-MAC does both with the same duration field.

munotes.in352

S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing

Message passing is not a longer packet. Each fragment is short and separately acknowledged; only the reservation is long.

Sleeping neighbours do not lose track. Every fragment and ACK carries the time left, so a node that wakes in the middle can go back to sleep for exactly the right time.

Giving up fairness is not giving up delivery. Short messages still get through; they wait for the long one, which suits a network where the long message is what the application needs.

Quick revision

  • Collision avoidance as in 802.11: physical and virtual carrier sense (NAV), carrier-sense time randomised in a contention window; contention in the receiver's listen interval, split into SYNC and RTS parts (15 and 31 slots on Mica).
  • Loser sleeps until the receiver's next listen; winners use their sleep time to finish. Unicast: RTS/CTS/DATA/ACK. Broadcast: no RTS/CTS.
  • Overhearing avoidance (after PAMAS, in-channel): nodes that hear an RTS or CTS for another node sleep while their NAV is non-zero. On E, C, A, B, D, F with A sending to B: D and C sleep, E and F are free: all immediate neighbours of sender and receiver.
  • Message passing: fragments in a burst, one RTS and one CTS; ACK per fragment; duration = all remaining fragments and ACKs; lost fragment: extend and resend; a cap on extensions. Trades per-node fairness for message latency and energy.
  • Program (10 fragments, Mica numbers): handshake per fragment 49.42 mJ; 802.11 fragmentation 38.02; S-MAC message passing 24.55, neighbours 0.12 and 0.24 mJ against 6.91 each.

Test yourself

1. How does S-MAC avoid collisions? It follows IEEE 802.11: before sending, a node performs virtual carrier sense, checking that its NAV (set from the duration fields of overheard packets) is zero, and physical carrier sense, listening to the channel, with its carrier-sense time randomised within a contention window. Unicast packets use the RTS/CTS/DATA/ACK exchange, which protects against hidden terminals; broadcasts use carrier sense only. All contention takes place in the receiver's listen interval, divided into parts for SYNC and for RTS packets.

2. Which nodes should sleep when node A transmits to node B, and why? All immediate neighbours of both A and B. B's neighbours must not transmit, because they would collide with the data at B. A's neighbours could transmit without disturbing B, but they could not receive any reply, since A's transmission would collide with it at them, and their transmissions could corrupt B's ACK at A; so transmitting would waste energy. Nodes farther away are free.

3. Explain overhearing avoidance in S-MAC. When a node hears an RTS or CTS addressed to another node, it reads the duration field, sets its NAV and goes to sleep until the NAV reaches zero, so it does not overhear the long data packet and the ACK that follow. The idea comes from PAMAS, but S-MAC uses only the ordinary RTS and CTS on the data channel rather than a separate signalling channel.

munotes.in353

S-MAC: Collision Avoidance, Overhearing Avoidance and Message Passing

4. What is message passing, and how does it differ from 802.11's fragmentation? A long message is split into fragments sent in a burst, reserved by a single RTS and CTS whose duration covers all of them; each fragment is acknowledged, a lost one is resent at once by extending the reservation, and every fragment and ACK carries the remaining duration. In 802.11 the RTS and CTS reserve only the first fragment and each fragment reserves the next, so neighbours must keep listening, and a failed fragment makes the sender give up and contend again for the sake of fairness.

5. What does message passing trade off, and why is that acceptable in a sensor network? It gives up per-node fairness: a node with a long message holds the channel while neighbours with short packets wait. It gains lower message-level latency and less control overhead and contention. In a sensor network the nodes serve one application and in-network processing needs whole messages, so application-level performance matters more than fairness between nodes.

6. Using the program's figures, how much energy do message passing and overhearing avoidance save? Sending a 10-fragment message with a handshake for every fragment cost 49.42 mJ over the four nodes; one handshake for the burst (802.11 fragmentation) cost 38.02 mJ, with each neighbour spending 6.91 mJ listening; S-MAC's message passing with overhearing avoidance cost 24.55 mJ, the neighbours spending only 0.12 and 0.24 mJ.

Contents This chapter on its own page

munotes.in354

Chapter Fifty-Three

S-MAC: Latency, Adaptive Listening and the Energy Saved

Syllabus topic Module 1, "Medium Access Control (MAC) in WSN: Sensor-MAC (S-MAC) / Sensor-MAC case study" (and the paired practical, "Implement duty-cycling MAC protocol and evaluate energy conservation in sensor networks")

In one line

Sleeping makes a packet wait at every hop for the next node to wake, about one frame per hop; adaptive listening, in which the neighbours that overheard an exchange wake briefly when it ends, lets a packet cross two hops per frame and halves the sleep delay, while periodic sleep still cuts the idle energy in proportion to the duty cycle.

In the wording a student can write in an examination: at each hop a packet meets carrier sense delay, backoff delay, transmission delay, propagation delay, processing delay and queueing delay; periodic sleep adds a sleep delay, the wait for the receiver's next listen interval. With a frame of length T(f), S-MAC's average latency over N hops is about N T(f), precisely N T(f) - T(f)/2 + t(cs) + t(tx), against N (t(cs) + t(tx)) without sleep. In adaptive listening, a node that overhears an RTS or CTS wakes up for a short time when that exchange ends; if it is the next hop, the packet is passed on at once instead of waiting a frame. Nodes two hops away usually cannot overhear, so the sleep delay is avoided at every other hop, and the latency falls to about N T(f)/2. Energy falls roughly in proportion to the duty cycle; latency grows roughly in inverse proportion. S-MAC with adaptive listening suits networks whose traffic is intermittent: long quiet periods with bursts. T-MAC goes further, ending a node's listen period early if nothing is heard, at some cost in throughput and latency.

Where a packet's time goes

The journal paper lists the delays a packet meets at each hop of any contention-based multi-hop network:

  • Carrier sense delay, while the sender senses the channel; its size "is determined by the contention window size."
  • Backoff delay, when carrier sense fails, "either because the node detects another transmission or because collision occurs."
  • Transmission delay, "determined by channel bandwidth, packet length and the coding scheme adopted."
  • Propagation delay, which "can normally be ignored" over sensor distances.
  • Processing delay, while the receiver handles the packet before forwarding it.
  • Queueing delay, which "becomes a dominant factor" under heavy traffic.

S-MAC adds one more. "When a sender gets a packet to transmit, it must wait until the receiver wakes up. We call it sleep delay since it is caused by the sleep of the receiver."

Latency without adaptive listening

The paper analyses the lightest load, one packet travelling alone, so there is no queueing or backoff, and propagation and processing are ignored. Write t(cs) for the mean carrier-sense delay, t(tx) for the transmission time of one hop's exchange, and T(f) for the frame length.

munotes.in355

S-MAC: Latency, Adaptive Listening and the Energy Saved

Without sleep, a node forwards a packet the moment it has it, so each hop costs t(cs) + t(tx) on average, and over N hops (equation 2):

E(D(N)) = N (t(cs) + t(tx)).

With periodic sleep, all nodes on one schedule, contention only begins when a listen interval begins. A packet that crossed a hop in one frame must wait for the next frame before the next node can send it on, so each hop after the first takes a whole frame. The first hop is different: the source's packet appears at a random moment in the frame and waits, on average, half a frame. Adding these up gives equation 8:

E(D(N)) = N T(f) - T(f)/2 + t(cs) + t(tx).

"The slope of the line is the frame length" instead of t(cs) + t(tx), and at low duty cycles the frame is far longer. The 2002 paper put it plainly: "the latency requirement of the application places a fundamental limit on the sleep time."

Adaptive listening

The 2004 paper adds a mechanism for the moment an event produces traffic. "The basic idea is to let the node who overhears its neighbor's transmissions (ideally only RTS or CTS) wake up for a short period of time at the end of the transmission. In this way, if the node is the next-hop node, its neighbor is able to immediately pass the data to it instead of waiting for its scheduled listen time. If the node does not receive anything during the adaptive listening, it will go back to sleep until its next scheduled listen time."

The duration fields make it possible: the neighbours of the sender learn from the RTS, and the neighbours of the receiver from the CTS, exactly when the exchange will end.

Two rows of hops over three frames, each frame a shaded listen interval and a long sleep. Without adaptive listen, hop 1, hop 2 and hop 3 each start at the next frame's listen interval. With adaptive listen, hops 1 and 2 happen back to back in the first frame, hops 3 and 4 in the second, hops 5 and 6 in the third

Figure 53.1 One hop per frame, and two with adaptive listening (after the paper's Fig. 4)

Why only every other hop. Take a chain i, j, k, l, and a packet going from i to j at a scheduled listen time. The next hop, k, is j's neighbour and hears j's CTS, so it wakes when i's transmission ends, and j can send to k at once. But l is two hops from j and cannot hear j's CTS; when j sends to k during what is normally sleep time, l is asleep and knows nothing of it. So k must wait for l's next scheduled listen. "Therefore, the sleep delay occurs at every other hop in S-MAC with adaptive listen", and the latency becomes equation 13:

E(D(N)) = N T(f)/2 + 2 t(cs) + 2 t(tx) - T(f)/2.

The slope of the line is now T(f)/2, half that of equation 8. The paper adds that real radios do better than this model, because range does not end sharply: a node two hops away may still catch some of the short CTS packets. If two-hop neighbours can receive from each other 20 to 30 per cent of the time, the paper suggests, more hops go adaptively and the latency falls further.

munotes.in356

S-MAC: Latency, Adaptive Listening and the Energy Saved

Two limits apply. A SYNC is sent only at scheduled listen times, so an adaptive exchange is not attempted if the next scheduled listen would begin too soon. And an RTS sent during adaptive listening may get no CTS if the intended receiver did not overhear the previous exchange, in which case "it just goes back to sleep and will try again at the next normal listen time."

Energy against latency, computed

The program puts numbers into the three equations, with the journal's Mica values: a 115 ms listen interval, and one hop's exchange (RTS, CTS, a 100-byte message with its 8-byte header, ACK) of 138 bytes at 0.8 ms a byte, 110.4 ms. The mean carrier-sense delay is not given in the paper; 10 ms is assumed. For each duty cycle it gives the frame length, the energy an idle node's radio uses in a day (listening and sleeping only, with the TR3000's 14.4 mW and 15 microwatts), and the latency over 10 hops without and with adaptive listening. Then it checks equations 8 and 13 by walking 20,000 packets along a timeline: each is born at a random moment, waits at each hop for the next listen interval unless the previous hop's CTS woke the next node, and draws its carrier-sense delay at random.

# Latency over N hops with S-MAC's periodic sleep, from the journal paper's
# equations 2, 8 and 13, checked by walking a packet along a timeline, and the
# energy each duty cycle costs an idle node. Mica numbers: 115 ms listen
# interval; one hop's exchange (RTS, CTS, a 100-byte message with its 8-byte
# header, ACK) at 0.8 ms a byte is 110.4 ms. The mean carrier-sense delay is
# not given in the paper: 10 ms is assumed.
import random

LISTEN, T_TX, T_CS = 0.115, (10 + 10 + 108 + 10) * 0.0008, 0.010
P_RX, P_SLEEP = 14.4, 0.015                    # mW, the TR3000
HOPS = 10

def formula(duty, adaptive):
    if duty >= 1:
        return HOPS * (T_CS + T_TX)                              # equation 2
    tf = LISTEN / duty
    if not adaptive:
        return HOPS * tf - tf / 2 + T_CS + T_TX                  # equation 8
    return HOPS * tf / 2 + 2 * T_CS + 2 * T_TX - tf / 2          # equation 13

def walk(duty, adaptive, rnd):
    """One packet: born at a random moment, it waits at each hop for the next
    listen interval, unless adaptive listening let the next hop overhear the
    CTS just before, in which case it goes at once (every other hop)."""
    tf = LISTEN / duty
    t = rnd.uniform(0, tf)                      # when the source has the packet
    born = t
    for hop in range(HOPS):
        overheard = adaptive and hop % 2 == 1
        if not overheard:
            t = (int(t / tf) + 1) * tf          # sleep until the next listen
        t += rnd.uniform(0, 2 * T_CS) + T_TX    # carrier sense, then the exchange
    return t - born

print("duty   frame s   J a day   10 hops: no adaptive   adaptive  (seconds)")
for duty in (1.0, 0.2, 0.1, 0.05, 0.01):         # equation 13 needs a frame over 2 exchanges
    energy = 86400 * (duty * P_RX + (1 - duty) * P_SLEEP) / 1000
    frame = LISTEN / duty
    if duty >= 1:
        print("%4.0f%% %8s %9.1f %22.2f %10s" % (duty * 100, "none", energy,
                                                 formula(duty, False), "same"))
    else:
        print("%4.0f%% %8.2f %9.1f %22.2f %10.2f" % (duty * 100, frame, energy,
                                                     formula(duty, False), formula(duty, True)))
rnd = random.Random(53)
for adaptive in (False, True):
    runs = [walk(0.1, adaptive, rnd) for _ in range(20000)]
    print("check at 10%%, %s: simulated %.2f s, equation %.2f s"
          % ("adaptive" if adaptive else "no adaptive", sum(runs) / len(runs),
             formula(0.1, adaptive)))
munotes.in357

S-MAC: Latency, Adaptive Listening and the Energy Saved

duty   frame s   J a day   10 hops: no adaptive   adaptive  (seconds)
 100%     none    1244.2                   1.20       same
  20%     0.57     249.9                   5.58       2.83
  10%     1.15     125.6                  11.05       5.42
   5%     2.30      63.4                  21.97      10.59
   1%    11.50      13.7                 109.37      51.99
check at 10%, no adaptive: simulated 11.04 s, equation 11.05 s
check at 10%, adaptive: simulated 5.42 s, equation 5.42 s

Reading it. The walked packets agree with the equations to a hundredth of a second, at a 10 per cent duty cycle both without adaptive listening (11.04 against 11.05 s) and with it (5.42 against 5.42 s).

Energy and latency move in opposite directions, almost in proportion. Always on, a node's radio uses 1,244.2 J a day and a packet crosses 10 hops in 1.20 s. At 10 per cent it uses 125.6 J, a tenth, and the 10 hops take 11.05 s, about nine times as long. At 1 per cent, 13.7 J a day, and 109.37 s. Each hop costs about one frame, and the frame is the listen interval divided by the duty cycle.

Adaptive listening halves the delay at no cost in idle energy. At every duty cycle the adaptive column is about half the other: 5.42 against 11.05 s at 10 per cent, 51.99 against 109.37 at 1 per cent. The idle energy is unchanged, because adaptive listening only wakes nodes that have just overheard an exchange; with no traffic it never fires.

munotes.in358

S-MAC: Latency, Adaptive Listening and the Energy Saved

The rows show why S-MAC's duty cycle is an application's choice. A network that must report within a few seconds over 10 hops cannot run at 1 per cent without adaptive listening, whatever the energy saved.

What the paper measured

On a ten-hop chain of Mica motes at a 10 per cent duty cycle (frames of 1.15 s), the paper measured energy, latency and throughput.

Energy. "S-MAC with periodic sleep achieves substantial energy savings over the MAC without periodic sleep in the multihop network, especially when traffic load is light." Under heavy load, adaptive listening also saved energy, because "the adaptive listen largely reduces the overall time needed to pass the fixed amount of data through the network." The 2002 paper's first test on a source node found an 802.11-like MAC using two to six times the energy of S-MAC, for messages sent every 1 to 10 s.

Latency. Without adaptive listening, latency grew by a frame per hop: "The reason is that each message has to wait for one sleep cycle on each hop." With it, latency was "very close to that of the MAC without any periodic sleep", about twice the fully active MAC's average, which is much better than equation 13's halving of the non-adaptive latency; the paper's reason is that "adaptive listening often allows S-MAC to immediately send a message to the next hop". The spread of latencies was much wider when sleeping, "and it increases with the number of hops", because some messages miss a node's listen interval.

Throughput. "As expected, periodic sleeping reduces throughput", and adaptive listening "significantly improves the end-to-end throughput"; at light load all modes delivered alike, since "Nothing happens during the long time between two messages."

The paper's summary is the chapter's: "periodic sleeping provides excellent energy performance at light traffic load, but adaptive listening is able to adjust to traffic and provide energy performance as good as no-sleep at heavy load. It makes S-MAC with adaptive listening ideal for sensor networks where traffic is intermittent."

After S-MAC: T-MAC

S-MAC's listen interval is fixed: a node stays awake for all of it even when nobody sends. T-MAC, from Delft in 2003, removes that waste. In the X-MAC paper's account, T-MAC "improves on the design of S-MAC by shortening the awake period if the channel is idle": a node listens only briefly after the synchronisation phase and, "if no data is received during this window, the node returns to sleep mode. If data is received, the node remains awake until no further data is received or the awake period ends." Its authors reported that "for variable workloads, T-MAC uses one fifth of the energy used by S-MAC", but "these gains come at the cost of reduced throughput and increased latency", and a later comparison found that "T-MAC was not able to handle as heavy a load as LPL and S-MAC due to the early sleeping problem": a node that goes back to sleep early can miss traffic meant for it.

munotes.in359

S-MAC: Latency, Adaptive Listening and the Energy Saved

The two families of chapters 50 to 53 thus meet the same trade in different ways: B-MAC and X-MAC keep no schedules and pay in preambles; S-MAC and T-MAC keep schedules and pay in synchronisation and latency.

Distinctions

S-MAC without adaptive listeningS-MAC with adaptive listening
Hops per frameOneTwo (the next hop overheard the CTS)
Latency over N hopsN T(f) - T(f)/2 + t(cs) + t(tx)N T(f)/2 + 2 t(cs) + 2 t(tx) - T(f)/2
Extra energy when idleNoneNone: it only fires after an exchange
SuitsVery light, delay-tolerant trafficIntermittent traffic with bursts
Sleep delayQueueing delay
CauseThe receiver is asleepEarlier packets are waiting at the node
When it mattersEvery hop, at any loadHeavy load
CureAdaptive listening, a higher duty cycleMore capacity, less traffic
S-MACT-MAC
Listen periodFixed lengthEnds early if the channel is idle
Energy under variable loadHigherAbout one fifth of S-MAC's (its authors)
WeaknessIdle listening within the listen periodEarly sleeping: less throughput, more latency

What it does not mean

Adaptive listening is not a higher duty cycle. It wakes only the neighbours of an exchange, only briefly, and only after traffic; an idle network's energy is unchanged.

Latency is not fixed by the duty cycle alone. The frame is the listen interval divided by the duty cycle, so a shorter listen interval at the same duty cycle means shorter frames and less delay.

The analysis is not the measurement. Equation 13 assumes nodes two hops apart never overhear each other; on real motes they sometimes do, and measured latency with adaptive listening was lower than the equation's.

Halving the latency does not make it small. At 1 per cent, even with adaptive listening, 10 hops take about 52 s in the program.

Quick revision

  • Per-hop delays: carrier sense, backoff, transmission, propagation, processing, queueing; S-MAC adds sleep delay, waiting for the receiver to wake.
  • No sleep: E(D(N)) = N (t(cs) + t(tx)). S-MAC: E(D(N)) = N T(f) - T(f)/2 + t(cs) + t(tx), slope T(f).
  • Adaptive listening: nodes that overhear an RTS or CTS wake briefly when the exchange ends; the next hop can forward at once; works at every other hop; E(D(N)) = N T(f)/2 + 2 t(cs) + 2 t(tx) - T(f)/2, slope T(f)/2.
  • Program (Mica numbers, 10 hops): always on 1.20 s, 1,244.2 J a day; 10 per cent 11.05 s or 5.42 s adaptive, 125.6 J; 1 per cent 109.37 s or 51.99 s, 13.7 J.
  • Measured: large energy savings at light load; adaptive listening saves energy at heavy load and brings latency near the no-sleep MAC's (about twice); sleeping widens the spread of latency and cuts throughput. Ideal for intermittent traffic.
  • T-MAC: ends the listen period early when idle; about one fifth of S-MAC's energy under variable load; early sleeping costs throughput and latency.
munotes.in360

S-MAC: Latency, Adaptive Listening and the Energy Saved

Test yourself

1. What delays does a packet meet at each hop, and which one does S-MAC add? Carrier sense delay, backoff delay, transmission delay, propagation delay, processing delay and queueing delay occur in any contention-based multi-hop network. S-MAC adds sleep delay: a sender with a packet must wait until its receiver wakes for its next listen interval.

2. Derive the average latency of S-MAC over N hops without adaptive listening. With all nodes on one schedule, contention begins only at the start of a listen interval. A packet that crosses one hop in a frame must wait for the next frame for the next hop, so each hop after the first takes one frame, T(f). The first hop starts after the packet has waited, on average, half a frame from its random creation time, then takes t(cs) + t(tx). Adding: E(D(N)) = T(f)/2 + (N - 1) T(f) + t(cs) + t(tx) = N T(f) - T(f)/2 + t(cs) + t(tx).

3. Explain adaptive listening and why it halves the sleep delay. A node that overhears a neighbour's RTS or CTS learns from its duration field when the exchange will end, and wakes briefly at that moment; if it is the next hop, the packet is forwarded at once instead of waiting for the next listen interval. The node after that, two hops from the last sender, did not overhear the CTS and is asleep, so the packet must wait for its scheduled listen. Thus the sleep delay is avoided at every other hop, and the latency's slope falls from T(f) to T(f)/2.

4. With a 115 ms listen interval, one hop's exchange of 110.4 ms and a carrier-sense delay of 10 ms, what is the latency over 10 hops at a 10 per cent duty cycle, with and without adaptive listening? T(f) = 0.115 / 0.10 = 1.15 s. Without: 10 × 1.15 - 0.575 + 0.010 + 0.1104 = 11.0454, about 11.05 s. With: 10 × 0.575 + 2 × 0.010 + 2 × 0.1104 - 0.575 = 5.4158, about 5.42 s. Without sleep it would be 10 × 0.1204 = 1.204 s.

munotes.in361

S-MAC: Latency, Adaptive Listening and the Energy Saved

5. How do energy and latency change with the duty cycle in S-MAC? The energy an idle node spends is roughly proportional to the duty cycle, since the radio listens only during the listen interval. The frame length is the listen interval divided by the duty cycle, and the latency is roughly one frame per hop (half a frame with adaptive listening), so it grows in inverse proportion. Cutting the duty cycle tenfold cuts the idle energy about tenfold and multiplies the latency about tenfold.

6. What does T-MAC change in S-MAC, and at what cost? T-MAC does not keep the listen period fixed: a node listens for only a short time after the synchronisation phase and goes back to sleep if no activity is heard, staying awake only while traffic continues. Its authors reported about one fifth of S-MAC's energy for variable workloads, but the early return to sleep reduces throughput, increases latency and can make a node miss traffic under heavy load.

Contents This chapter on its own page

munotes.in362

Chapter Fifty-Four

Routing Challenges and Design Issues in WSNs

Syllabus topic Module 1, "Routing in WSN: Routing challenges and design issues in WSNs"

In one line

Routing in a sensor network must carry data from many sources to a base station over many hops, on nodes with little energy, no global addresses and changing links, so its design is shaped by how the nodes are deployed, how they report, how much they can aggregate, and how long the network must last.

In the wording a student can write in an examination: routing in WSNs differs from routing in other networks because there are too many nodes for global addressing, traffic flows from many sources to one base station, nodes are constrained in energy, processing and storage, most nodes are stationary but links change, networks are application specific, position matters, and the data is redundant. The main design issues, following Al-Karaki and Kamal, are: node deployment (deterministic or random); energy consumption without losing accuracy, since each node is both a sender and a router; the data reporting model (time-driven, event-driven, query-driven or hybrid); node and link heterogeneity; fault tolerance; scalability to hundreds or thousands of nodes; network dynamics (moving nodes, base stations or phenomena); the wireless transmission medium; connectivity; coverage; data aggregation; and quality of service, such as bounded latency, traded against lifetime.

Why routing is different here

A router on the Internet forwards packets between addresses for millions of independent users. A sensor network does something narrower and harder. Al-Karaki and Kamal set out the differences.

  • No global addressing. With so many nodes, "it is not possible to build a global addressing scheme for the deployment of a large number of sensor nodes as the overhead of ID maintenance is high. Thus, traditional IP-based protocols may not be applied to WSNs." What matters is the data: "sometimes getting the data is more important than knowing the IDs of which nodes sent the data."
  • Many to one. "almost all applications of sensor networks require the flow of sensed data from multiple sources to a particular BS."
  • Tight constraints on energy, processing and storage, which "require careful resource management."
  • Mostly stationary nodes. Unlike a MANET's, the nodes rarely move, but they fail, run down and sleep, so the topology still changes.
  • Application specific. "the challenging problem of low-latency precision tactical surveillance is different from that required for a periodic weather-monitoring task."
  • Position matters, "since data collection is normally based on the location", though GPS on every node is not feasible ([Time Synchronisation and Localisation]).
  • Redundant data. Readings come from a common phenomenon, so routing should exploit the redundancy; sensor networks are data-centric, requesting data by attribute, such as temperature above 60 degrees F, rather than by node ([Design Principles: Data Centricity, Location, Activity and Heterogeneity]).
munotes.in363

Routing Challenges and Design Issues in WSNs

The design issues

Al-Karaki and Kamal list the factors that "must be overcome before efficient communication can be achieved in WSNs". Each is a question the designer of a routing protocol must answer.

1. Node deployment. "The deployment can be either deterministic or randomized." Placed by hand, nodes can use pre-determined paths; scattered at random, they must organise themselves, and uneven density may call for clustering. Because radios are short-range, "it is most likely that a route will consist of multiple wireless hops."

2. Energy consumption without losing accuracy. "In a multihop WSN, each node plays a dual role as data sender and data router." A node that dies of flat batteries does not only lose its readings; it breaks routes, and "might require rerouting of packets and reorganization of the network." ([Energy-aware Routing] took this up for ad hoc networks.)

3. The data reporting model. Reporting "can be categorized as either time-driven (continuous), event-driven, query-driven, and hybrid." Time-driven reporting suits periodic monitoring; in event- and query-driven reporting, nodes "react immediately to sudden and drastic changes in the value of a sensed attribute due to the occurrence of a certain event or a query is generated by the BS", which suits time-critical applications. "The routing protocol is highly influenced by the data reporting model with regard to energy consumption and route stability." The program below shows how much.

4. Node and link heterogeneity. Nodes may differ in sensors, in reporting rates and in capability; "hierarchical protocols designate a clusterhead node different from the normal sensors", and cluster heads may be more powerful nodes, which then carry the burden of transmission to the base station.

5. Fault tolerance. "The failure of sensor nodes should not affect the overall task of the sensor network." If nodes fail, routing must form new links and routes, perhaps "rerouting packets through regions of the network where more energy is available", so redundancy is needed at several levels.

6. Scalability. The network may hold "hundreds or thousands, or more" nodes, and routing must work at that size and respond to events, while most nodes sleep until something happens.

7. Network dynamics. Most designs assume stationary nodes, but base stations or sensors may move, and the phenomenon itself may be static (a forest watched for fire) or moving (a tracked target), which generates periodic traffic.

8. The transmission medium. Links suffer the wireless channel's fading and errors, rates are low ("on the order of 1-100 kb/s"), and the MAC below routing matters ([MAC Protocols for Sensor Networks: The Job and Where the Energy Goes]).

9. Connectivity. High density means nodes are expected to be "highly connected", but failures shrink the network, and connectivity "depends on the, possibly random, distribution of nodes."

munotes.in364

Routing Challenges and Design Issues in WSNs

10. Coverage. Each sensor sees only a limited area, so "area coverage is also an important design parameter" ([Deployment and Coverage: Random Against Grid]).

11. Data aggregation. "similar packets from multiple nodes can be aggregated so that the number of transmissions is reduced", by functions such as "duplicate suppression, minima, maxima and average", or by signal processing (data fusion).

12. Quality of service. Some data is useless if late, so "bounded latency for data delivery is another condition for time-constrained applications", but "in many applications, conservation of energy, which is directly related to network lifetime, is considered relatively more important than the quality of data sent." The routing protocol must strike the balance the application needs.

Akyildiz and colleagues give the same themes as principles for the network layer: power efficiency is always important; sensor networks are mostly data-centric; data aggregation is useful only when it does not hinder the collaborative effort of the nodes; and an ideal sensor network has attribute-based addressing and location awareness.

One field, four ways of reporting

The program scatters 100 nodes at random over a 100 m by 100 m field with the base station at a corner and a radio range of 20 m, builds each node's shortest-hop route to the base station, and counts the transmissions each node makes in a day, its own and those it relays. The traffic is illustrative: time-driven reporting every 5 minutes, first relayed as it is and then aggregated so that each node sends one packet per period for its whole subtree; event-driven reporting of 4 events a day at random places, each reported by the nodes within 15 m; and query-driven reporting of 24 queries a day, each flooded to every node and answered by the nodes in a random 30 m square. Events and queries are averaged over a year.

# One field, several routing design issues at once. 100 nodes scattered at
# random over 100 m by 100 m, the base station at a corner, radio range 20 m.
# Each node's route is its shortest-hop path to the base station, and every
# packet is transmitted once by its source and once by every relay on its way.
import random
from collections import deque

rnd = random.Random(54)
N, SIDE, RANGE = 100, 100.0, 20.0
pts = [(0.0, 0.0)] + [(rnd.uniform(0, SIDE), rnd.uniform(0, SIDE)) for _ in range(N)]

def near(a, b, r=RANGE):
    return (a[0] - b[0]) ** 2 + (a[1] - b[1]) ** 2 <= r * r

# the shortest-hop tree: breadth-first search outwards from the base station (0)
parent, hops, queue = {0: None}, {0: 0}, deque([0])
while queue:
    u = queue.popleft()
    for v in range(1, N + 1):
        if v not in hops and near(pts[u], pts[v]):
            hops[v], parent[v] = hops[u] + 1, u
            queue.append(v)
reached = [v for v in hops if v]

def carried(reports):
    """reports: packets each node originates per day. Returns the packets each
    node transmits per day: its own and every one it relays."""
    tx = dict.fromkeys(reached, 0)
    for src, k in reports.items():
        v = src
        while v:
            tx[v] += k
            v = parent[v]
    return tx

def year_of(events, per_day, radius):
    """Reports per day, averaged over a year of events at random places: every
    node within `radius` of an event (or inside a query's square) reports once."""
    count = dict.fromkeys(reached, 0)
    for _ in range(per_day * 365):
        spot = (rnd.uniform(0, SIDE), rnd.uniform(0, SIDE))
        for v in reached:
            if events(spot, pts[v], radius):
                count[v] += 1
    return {v: c / 365 for v, c in count.items()}

in_circle = near
in_square = lambda s, p, half: abs(s[0] - p[0]) <= half and abs(s[1] - p[1]) <= half

models = {
    "time-driven, every 5 min": carried(dict.fromkeys(reached, 288)),
    "time-driven, aggregated": dict.fromkeys(reached, 288),   # one packet per period each
    "event-driven, 4 events/day": carried(year_of(in_circle, 4, 15.0)),
}
replies = carried(year_of(in_square, 24, 15.0))      # 24 queries a day, 30 m squares
models["query-driven, 24/day"] = {v: replies[v] + 24 for v in reached}  # + the query flood

far = max(hops.values())
print("reached %d of %d nodes; hops to the base station: 1 to %d, mean %.2f"
      % (len(reached), N, far, sum(hops[v] for v in reached) / len(reached)))
print("the base station's own neighbours: %d" % sum(1 for v in reached if hops[v] == 1))
print("%-28s %10s %9s %9s %7s" % ("reporting model", "total/day", "mean", "busiest", "ratio"))
for name, tx in models.items():
    total, busiest = sum(tx.values()), max(tx.values())
    print("%-28s %10.0f %9.1f %9.1f %7.1f"
          % (name, total, total / len(tx), busiest, busiest / (total / len(tx))))
munotes.in365

Routing Challenges and Design Issues in WSNs

reached 100 of 100 nodes; hops to the base station: 1 to 9, mean 4.94
the base station's own neighbours: 2
reporting model               total/day      mean   busiest   ratio
time-driven, every 5 min         142272    1422.7   24480.0    17.2
time-driven, aggregated           28800     288.0     288.0     1.0
event-driven, 4 events/day          121       1.2      21.5    17.7
query-driven, 24/day               3286      32.9     184.0     5.6
The 100 nodes of the field joined by their routes to the base station at the bottom-left corner. Every route ends through one of the two nodes beside the base station, and most through one of them, shaded darkest with two others that relay for 30 or more nodes; nodes relaying for 5 to 29 are shaded lighter; the rest relay little

Figure 54.1 The program's field: routes to the base station, shaded by how much each node relays under time-driven reporting

Reading it: deployment and multiple hops. Random deployment gave a connected network, every node reaching the base station, but in up to 9 hops, 4.94 on average: routes are long even in a small field.

Reading it: the dual role. With time-driven reporting every node originates 288 packets a day, yet the network makes 142,272 transmissions, about 4.9 per packet, the mean hop count. The busiest node transmits 24,480 a day, 17.2 times the average: it is one of only 2 nodes beside the base station, and relays for most of the field. It will run out of energy first, and when it does, the routes of everything behind it must be rebuilt through the one remaining neighbour. Fault tolerance and energy are the same problem seen twice.

munotes.in366

Routing Challenges and Design Issues in WSNs

Reading it: the reporting model. The same field makes 142,272 transmissions a day time-driven, 3,286 query-driven and 121 event-driven. The model, set by the application, decides the traffic by a factor of more than a thousand, and so decides which routing protocol makes sense: proactive routes are worth their upkeep for continuous traffic, not for a few events a day.

Reading it: aggregation. Aggregating each period's readings on the way cuts the total to 28,800 transmissions and makes every node's load equal, a ratio of 1.0. It works only when the application can use a combined value (a maximum, an average, a count), which is why Akyildiz and colleagues attach the condition that aggregation must not hinder the collaborative task.

Distinctions

Time-drivenEvent-drivenQuery-driven
Who starts itA clock, periodicallyA change in the sensed valueThe base station's query
TrafficSteady, heavyRare, burstyOn demand
SuitsPeriodic monitoringTime-critical detectionInteractive retrieval
In the program142,272 transmissions a day1213,286
Routing in IP networks and MANETsRouting in WSNs
AddressingGlobal, per nodeNo global IDs; data- or attribute-based
TrafficAny node to any nodeMany sources to one base station
ConstraintThroughput, delayEnergy, above all
DataIndependentRedundant, often aggregated
Deterministic deploymentRandom deployment
PlacementBy handScattered
RoutesPre-determined pathsSelf-organised, multi-hop
DensityEvenUneven: clustering may be needed

What it does not mean

Many-to-one is not the only traffic. Queries flow out from the base station, and some applications need node-to-node or multicast flows; the survey notes that many-to-one "does not prevent the flow of data to be in other forms".

A connected network is not a durable one. The field above is connected, but two nodes carry nearly all its traffic, and losing either reshapes the whole network.

Aggregation is not free accuracy. It reduces transmissions only when a combined value is acceptable, and it delays data while a node waits for its children.

The design issues are not independent. Deployment sets connectivity and hop counts, the reporting model sets the load, the load sets which nodes die first, and their death tests fault tolerance.

Quick revision

  • Different from other networks: no global addressing, many sources to one base station, tight energy, processing and storage, mostly stationary nodes, application specific, location matters, redundant data (data-centric, attribute-based).
  • Twelve design issues (Al-Karaki and Kamal): node deployment; energy without losing accuracy (nodes are senders and routers); data reporting model (time-driven, event-driven, query-driven, hybrid); node/link heterogeneity; fault tolerance; scalability; network dynamics; transmission medium (1 to 100 kb/s); connectivity; coverage; data aggregation; quality of service (bounded latency against lifetime).
  • Program (100 random nodes, corner base station): up to 9 hops, mean 4.94; only 2 base-station neighbours; time-driven 142,272 transmissions a day, busiest node 24,480 (17.2 times the mean); aggregated 28,800, all equal; query-driven 3,286; event-driven 121.
munotes.in367

Routing Challenges and Design Issues in WSNs

Test yourself

1. Why can traditional IP-based routing not simply be used in a wireless sensor network? There are too many nodes for a global addressing scheme to be worth its overhead, and the application wants data rather than node identities. Traffic flows mostly from many sources to a single base station; nodes are tightly limited in energy, processing and memory; links change as nodes fail and sleep; and the data from nearby nodes is redundant, so routing should aggregate it and address it by attributes.

2. List and explain any six routing design issues in WSNs. Node deployment: deterministic placement allows fixed paths, random scattering needs self-organising multi-hop routes. Energy consumption without losing accuracy: every node is both a sender and a router, and a node's death forces rerouting. Data reporting model: time-driven, event-driven, query-driven or hybrid reporting determines traffic, energy and route stability. Fault tolerance: failed nodes must not stop the task, so routes must be rebuilt and redundancy kept. Scalability: routing must work with hundreds or thousands of nodes. Data aggregation: combining redundant data from several nodes reduces transmissions. (Others: heterogeneity, network dynamics, transmission medium, connectivity, coverage, quality of service.)

3. Compare the time-driven, event-driven and query-driven reporting models. In the time-driven model nodes sense and report periodically, suiting continuous monitoring but producing steady, heavy traffic. In the event-driven model nodes report only when the sensed value changes sharply, suiting time-critical detection with rare, bursty traffic. In the query-driven model nodes report when the base station asks. In the program, one field made 142,272 transmissions a day time-driven, 3,286 query-driven and 121 event-driven.

4. Why does multi-hop routing to a single base station shorten network lifetime, and how can aggregation help? Every packet must pass through the few nodes near the base station, which therefore relay far more than the rest and run out of energy first, disconnecting the nodes behind them. In the program the busiest node transmitted 24,480 packets a day, 17.2 times the average. If each node combines its subtree's readings into one packet per period, every node transmits the same amount, 288 a day in the program, and the total falls from 142,272 to 28,800.

munotes.in368

Routing Challenges and Design Issues in WSNs

5. What is meant by energy consumption without losing accuracy? Nodes must conserve energy in computation and communication, but not by degrading the data the application needs. Because each node both senses and routes, a node that exhausts its battery removes its own data and breaks routes through it, so energy-efficient routing must spread the load and keep the network's view of the environment accurate for as long as possible.

6. How does quality of service conflict with energy in WSN routing? Some applications need data within a bounded time, which favours short, fast, always-ready routes. But in many applications lifetime matters more than the quality of each report, so as energy runs low the network may reduce the quality of its results, sending less often or aggregating more, to last longer. Energy-aware routing must strike the balance the application requires.

Contents This chapter on its own page

munotes.in369

Chapter Fifty-Five

Routing Strategies in WSNs: A Map

Syllabus topic Module 1, "Routing in WSN: Routing strategies in WSNs"

In one line

WSN routing protocols are classified two ways: by the network's structure (flat, where all nodes are equal; hierarchical, where cluster heads collect and aggregate; location-based, where positions guide the route) and by how the protocol operates (by negotiation, over multiple paths, by queries, with quality of service, or by the processing done on the way).

In the wording a student can write in an examination: by network structure, routing protocols are flat (all nodes have the same role and routing is usually data-centric: flooding, gossiping, SPIN, directed diffusion, rumour routing), hierarchical (nodes are grouped into clusters whose heads aggregate and forward data: LEACH, PEGASIS, TEEN and APTEEN) or location-based (nodes' positions are used to send data only towards the region that needs it: GAF, GEAR, GPSR). By protocol operation they are negotiation-based (metadata is exchanged to suppress redundant data: SPIN), multipath (several paths are kept for reliability or load balance), query-based (the sink sends queries and matching data flows back: directed diffusion, rumour routing), QoS-based (delay, energy or bandwidth targets are met: SAR, SPEED) and based on coherent or non-coherent processing (how much data is processed before it reaches an aggregator). By route discovery they are proactive, reactive or hybrid. A protocol can belong to several classes at once.

Two ways to classify

Al-Karaki and Kamal organise their survey around two questions. The first is about the network: "routing in WSNs can be divided into flat-based routing, hierarchical-based routing, and location-based routing", depending on the network structure. The second is about the protocol: "these protocols can be classified into multipath-based, query-based, negotiation-based, QoS-based, or coherent-based routing techniques depending on the protocol operation."

They add a third, familiar from MANETs ([Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid]): "In proactive protocols, all routes are computed before they are really needed, while in reactive protocols, routes are computed on demand." For a sensor network their advice leans proactive: "When sensor nodes are static, it is preferable to have table driven routing protocols rather than using reactive protocols. A significant amount of energy is used in route discovery and setup of reactive protocols."

Two trees. By network structure: flat (flooding, gossiping, SPIN, directed diffusion, rumour routing), hierarchical (LEACH, PEGASIS, TEEN, APTEEN) and location-based (GAF, GEAR, greedy forwarding and GPSR). By protocol operation: negotiation (SPIN), multipath (braided paths), query (directed diffusion, rumour routing), QoS (SAR, SPEED) and coherent processing (SWE, MWE)

Figure 55.1 The survey's taxonomy, with this book's protocols placed on it

By network structure

Flat routing. "In flat-based routing, all nodes are typically assigned equal roles or functionality." There are too many nodes to give each a global identifier, so flat protocols are usually data-centric: the sink asks for data by what it is, not by who holds it, and nodes aggregate on the way. The simplest are flooding and gossiping ([Flooding, Gossiping and the Broadcast Storm]); SPIN negotiates before sending ([SPIN: Negotiating Before Sending]); directed diffusion sets up gradients from interests, and rumour routing sends agents instead of flooding ([Directed Diffusion and Rumour Routing]). The collection tree of [Routing Tables and What Happens When the Topology Changes] is flat too: every node forwards for others along a tree to the sink.

munotes.in370

Routing Strategies in WSNs: A Map

Hierarchical routing. "In hierarchical-based routing, however, nodes will play different roles in the network." Nodes form clusters, and cluster heads gather their members' data, aggregate it and send it on. LEACH chooses cluster heads at random and rotates the role ([LEACH: Clusters That Take Turns]); PEGASIS builds a chain instead of clusters, and TEEN and APTEEN add thresholds so that nodes report only what matters ([PEGASIS, TEEN and the Other Hierarchical Protocols]).

Location-based routing. "In location-based routing, sensor nodes' positions are exploited to route data in the network." "In this kind of routing, sensor nodes are addressed by means of their locations", and positions come from GPS on a few nodes, from signal strengths, or from localisation ([Time Synchronisation and Localisation]). GAF uses positions to put redundant nodes to sleep ([Energy Efficiency in Ad Hoc Networks: Where the Energy Goes]); greedy forwarding, GPSR and GEAR route towards a position ([Geographic Routing: Greedy Forwarding and GPSR]).

Flat against hierarchical. The survey compares the two in its Table 2. In summary:

Hierarchical routingFlat routing
Channel accessReservation-based scheduling, collisions avoidedContention-based, with collision overhead
Duty cycleReduced by periodic sleepingVaried by controlling nodes' sleep times
AggregationBy the cluster headBy each node on a multi-hop path
Routing"Simple but non-optimal"Can be made optimal, at added complexity
SynchronisationNeeds global and local synchronisationLinks formed on the fly, without it
OverheadForming clusters throughout the networkRoutes formed only where there is data
LatencyLower: cluster heads always availableWaking intermediate nodes and setting up paths
EnergyDissipated uniformly, not controllableDepends on and adapts to the traffic
FairnessFair channel allocationNot guaranteed

The comparison is a trade: hierarchy buys order (schedules, even energy, fairness) at the price of organisation, while flat routing adapts to where the data is but guarantees little.

By protocol operation

Negotiation-based routing. "These protocols use high level data descriptors in order to eliminate redundant data transmissions through negotiation." Flooding sends every node duplicate copies, "implosion and overlap"; negotiation asks before sending. SPIN is the example.

Multipath routing. Several paths are kept between a source and the sink, so that when the primary fails an alternate exists. The survey measures a protocol's fault tolerance "by the likelihood that an alternate path exists between a source and a destination when the primary path fails", and names the price: "network reliability can be increased at the expense of increased overhead of maintaining the alternate paths." Braided multipath, built on directed diffusion, keeps alternates cheap by keeping them close to the primary path.

munotes.in371

Routing Strategies in WSNs: A Map

Query-based routing. "the destination nodes propagate a query for data (sensing task) from a node through the network and a node having this data sends the data which matches the query back to the node, which initiates the query." Directed diffusion's interests are queries; rumour routing sends a query towards an event along paths left by agents.

QoS-based routing. "the network has to balance between energy consumption and data quality", meeting targets for "delay, energy, bandwidth". SAR weighs energy, the QoS of each path and each packet's priority; SPEED gives each packet a target speed across the field so that delay can be predicted.

Coherent and non-coherent processing. In non-coherent processing, nodes process the raw data locally before sending it on to the nodes that process it further, the aggregators; in coherent processing, data goes to the aggregators "after minimum processing", such as time stamping and duplicate suppression. The survey's examples are the single-winner and multiple-winner elections of an aggregator (SWE and MWE).

The protocols of this book, compared

The survey's Fig. 8 compares routing protocols on a dozen properties. For the ones this book teaches, a few of its columns:

ProtocolClassNegotiationAggregationScalabilityMultipathQuery-based
SPINFlatYesYesLimitedYesYes
Directed diffusionFlatYesYesLimitedYesYes
Rumour routingFlatNoYesGoodNoYes
LEACHHierarchicalNoYesGoodNoNo
TEEN and APTEENHierarchicalNoYesGoodNoNo
PEGASISHierarchicalNoNoGoodNoNo
GAFLocationNoNoGoodNoNo
GEARLocationNoNoLimitedNoNo

Two readings help an answer. The flat, data-centric protocols are the query-based ones, and the survey rates their scalability as limited, except rumour routing, which avoids flooding. The hierarchical protocols scale well but, in the survey's columns, depend on a fixed base station. (The survey marks PEGASIS as doing no aggregation, though PEGASIS's own paper fuses data along its chain; [PEGASIS, TEEN and the Other Hierarchical Protocols] shows how.)

Answering an examination question on routing strategies

A complete examination answer has four parts: the two classifications with a line of definition each; one example protocol for each class, with its one idea (SPIN negotiates, LEACH rotates cluster heads, GPSR forwards greedily and walks around voids); the flat against hierarchical comparison; and the remark that classes overlap (directed diffusion is flat, query-based, negotiation-based and multipath at once). The next eight chapters give the detail for each protocol.

The three strategies, run over one field

The map is easier to trust once the strategies have been run. A field of a hundred nodes is laid out, a sink is put in one corner, and every node sends one report; the program counts the transmissions each strategy costs and, more importantly, what the busiest node has to carry.

munotes.in372

Routing Strategies in WSNs: A Map

# The three strategies, run over one field, counted in transmissions.
import math
import random

FIELD, N, RANGE = 200.0, 100, 40.0
SINK = (0.0, 0.0)
rng = random.Random(20260930)
nodes = [(rng.uniform(0, FIELD), rng.uniform(0, FIELD)) for _ in range(N)]

def d(a, b):
    return math.hypot(a[0] - b[0], a[1] - b[1])

pts = [SINK] + nodes                      # 0 is the sink
nbrs = {i: [j for j in range(len(pts)) if j != i and d(pts[i], pts[j]) <= RANGE]
        for i in range(len(pts))}

# 1. Flat multi-hop: every node sends its own report, forwarded hop by hop on
#    the shortest path to the sink.
hops = {0: 0}
queue, head = [0], 0
while head < len(queue):
    cur = queue[head]
    head += 1
    for nb in nbrs[cur]:
        if nb not in hops:
            hops[nb] = hops[cur] + 1
            queue.append(nb)
unreached = [i for i in range(1, len(pts)) if i not in hops]
print("A field of %d nodes in %.0f m square, radio range %.0f m, sink in one corner."
      % (N, FIELD, RANGE))
print("  %d nodes can reach the sink at all; %d are isolated."
      % (len(hops) - 1, len(unreached)))
print("  the furthest is %d hops away, the average %.2f hops."
      % (max(hops.values()), sum(hops[i] for i in hops if i) / float(len(hops) - 1)))

load = dict((i, 0) for i in range(len(pts)))
parent = {}
for i in range(1, len(pts)):
    if i not in hops:
        continue
    cur = i
    while cur != 0:
        nxt = min((nb for nb in nbrs[cur] if hops.get(nb, 10 ** 6) == hops[cur] - 1),
                  key=lambda nb: (d(pts[nb], SINK), nb))
        load[cur] += 1
        cur = nxt
flat_total = sum(load.values())
flat_worst = max(load.values())

# 2. Hierarchical: the field is divided into clusters, a member sends one hop
#    to its head, and the head sends one aggregated report onward.
def cluster_run():
    """Heads spread over the field, not huddled by the sink: one head nearest
    each point of a 4 by 2 grid."""
    targets = [(FIELD * (c + 0.5) / 4, FIELD * (r + 0.5) / 2)
               for r in range(2) for c in range(4)]
    heads, taken = [], set()
    for t in targets:
        h = min((i for i in range(1, len(pts)) if i not in taken),
                key=lambda i: (d(pts[i], t), i))
        heads.append(h)
        taken.add(h)
    head_load = dict((h, 0) for h in heads)
    member_tx, out_of_range = 0, 0
    for i in range(1, len(pts)):
        if i in heads:
            continue
        h = min(heads, key=lambda h: (d(pts[i], pts[h]), h))
        head_load[h] += 1
        member_tx += 1
        if d(pts[i], pts[h]) > RANGE:
            out_of_range += 1
    relay = sum(hops[h] for h in heads if h in hops)
    return member_tx + relay, max(head_load.values()) + 1, out_of_range, len(heads)

# 3. Geographic greedy: forward to the neighbour nearest the sink.
def greedy_run():
    load = dict((i, 0) for i in range(len(pts)))
    delivered, stuck = 0, 0
    for i in range(1, len(pts)):
        cur, seen = i, set()
        while cur != 0:
            if cur in seen:
                stuck += 1
                break
            seen.add(cur)
            better = [nb for nb in nbrs[cur] if d(pts[nb], SINK) < d(pts[cur], SINK)]
            if not better:
                stuck += 1
                break
            load[cur] += 1
            cur = min(better, key=lambda nb: (d(pts[nb], SINK), nb))
        else:
            delivered += 1
    return sum(load.values()), max(load.values()), delivered, stuck

h_total, h_worst, h_far, h_heads = cluster_run()
g_total, g_worst, g_ok, g_stuck = greedy_run()

print()
print("One report from every node, counted in transmissions:")
print("  strategy          total   the busiest node sends   delivered")
print("  flat multi-hop  %7d %20d %11d" % (flat_total, flat_worst, len(hops) - 1))
print("  hierarchical    %7d %20d %11d  (%d heads, one packet each)"
      % (h_total, h_worst, len(hops) - 1, h_heads))
print("  geographic      %7d %20d %11d  (%d got stuck)"
      % (g_total, g_worst, g_ok, g_stuck))
print()
print("Flat routing and geographic routing send nearly the same number of packets,")
print("because both carry every report separately: what geographic routing saves is")
print("the routing table, not the traffic. Hierarchy is the only one that sends fewer,")
print("and only because the head aggregates: %d reports leave the field as %d packets."
      % (len(hops) - 1, h_heads))
print("The model is generous to it: %d of the %d members are further from their head"
      % (h_far, len(hops) - 1 - h_heads))
print("than the radio reaches, so a real cluster needs either more heads or a hop inside")
print("the cluster, and either costs some of the saving back.")
print("The busiest node is the number that decides how long the network lives. Under")
print("flat and geographic routing it sits beside the sink and carries the whole field;")
print("under hierarchy it is a cluster head, and the load is %.1f times lighter."
      % (flat_worst / float(h_worst)))
munotes.in373

Routing Strategies in WSNs: A Map

A field of 100 nodes in 200 m square, radio range 40 m, sink in one corner.
  100 nodes can reach the sink at all; 0 are isolated.
  the furthest is 8 hops away, the average 5.25 hops.

One report from every node, counted in transmissions:
  strategy          total   the busiest node sends   delivered
  flat multi-hop      525                   71         100
  hierarchical        136                   15         100  (8 heads, one packet each)
  geographic          481                   60          89  (11 got stuck)

Flat routing and geographic routing send nearly the same number of packets,
because both carry every report separately: what geographic routing saves is
the routing table, not the traffic. Hierarchy is the only one that sends fewer,
and only because the head aggregates: 100 reports leave the field as 8 packets.
The model is generous to it: 25 of the 92 members are further from their head
than the radio reaches, so a real cluster needs either more heads or a hop inside
the cluster, and either costs some of the saving back.
The busiest node is the number that decides how long the network lives. Under
flat and geographic routing it sits beside the sink and carries the whole field;
under hierarchy it is a cluster head, and the load is 4.7 times lighter.
munotes.in374

Routing Strategies in WSNs: A Map

Distinctions

FlatHierarchicalLocation-based
RolesAll nodes equalCluster heads and membersEqual, but positions known
AddressingBy data (attributes)Through cluster headsBy location
StrengthAdapts to where the data isOrder, even energy, scaleSends only towards the region needed
WeaknessFlooding overhead, limited scaleCluster formation, synchronisationNeeds positions; voids
ProactiveReactive
Routes computedBefore they are neededOn demand
SuitsStatic sensor networks (the survey's advice)Rare, unpredictable traffic
CostKeeping tables currentRoute discovery before sending

What it does not mean

The classes are not exclusive. A protocol has a structure and an operation, and often several operations: directed diffusion appears under flat, query-based, negotiation-based and multipath.

Flat does not mean flooding. Flooding is the simplest flat protocol, but directed diffusion, rumour routing and collection trees are flat too.

Hierarchical does not mean more powerful hardware. Cluster heads may be ordinary nodes taking turns, as in LEACH, or more capable nodes placed on purpose.

Location-based does not mean GPS everywhere. A few anchors, signal strengths or relative coordinates may be enough.

Quick revision

  • By network structure: flat (equal roles; data-centric; flooding, gossiping, SPIN, directed diffusion, rumour routing), hierarchical (cluster heads aggregate; LEACH, PEGASIS, TEEN, APTEEN), location-based (positions; GAF, GEAR, GPSR).
  • By protocol operation: negotiation-based (SPIN: suppress redundant data), multipath (alternate paths for fault tolerance, at maintenance cost), query-based (directed diffusion, rumour routing), QoS-based (SAR, SPEED: delay, energy, bandwidth), coherent / non-coherent processing (minimum processing or local processing before the aggregators).
  • Proactive / reactive / hybrid; for static sensor networks the survey prefers table-driven (proactive).
  • Hierarchical against flat (Table 2): reservation against contention; aggregation by cluster head against by each node; simple non-optimal against optimal but complex; needs synchronisation against links on the fly; uniform energy against traffic-dependent; fair against not guaranteed.

Test yourself

1. Classify WSN routing protocols by network structure, with an example of each. Flat routing, in which all nodes have the same role and routing is usually data-centric, for example SPIN or directed diffusion; hierarchical routing, in which nodes are organised into clusters and cluster heads collect, aggregate and forward their members' data, for example LEACH; and location-based routing, in which nodes' positions are used to route data towards the region of interest, for example GPSR or GEAR.

munotes.in375

Routing Strategies in WSNs: A Map

2. Classify WSN routing protocols by protocol operation. Negotiation-based protocols exchange metadata before sending data to avoid redundant transmissions (SPIN). Multipath protocols maintain several paths to improve reliability or share load (braided multipath). Query-based protocols send queries from the sink and route matching data back (directed diffusion, rumour routing). QoS-based protocols meet delay, energy or bandwidth targets (SAR, SPEED). Coherent and non-coherent processing-based protocols differ in how much data is processed before it reaches the aggregator nodes.

3. Compare hierarchical and flat routing in WSNs. Hierarchical routing uses reservation-based scheduling, avoiding collisions, reduces the duty cycle by periodic sleeping, aggregates at cluster heads, is simple but not optimal, needs synchronisation and cluster formation, gives low latency and uniform energy use, and allocates the channel fairly. Flat routing uses contention, with collision overhead, aggregates at every node on a path, can be optimal at greater complexity, forms links on the fly without synchronisation and routes only where there is data, and its energy use and latency depend on the traffic, with no guarantee of fairness.

4. Why does the survey prefer proactive routing for static sensor networks? When nodes do not move, routes stay valid for a long time, so the cost of computing them in advance is paid rarely, whereas reactive protocols spend a significant amount of energy on route discovery and set-up each time a route is needed.

5. Can a protocol belong to more than one class? Give an example. Yes. Directed diffusion is flat by structure, and by operation it is query-based (the sink sends interests), negotiation-based and multipath (the survey marks both), so it appears in several classes of the same taxonomy.

Contents This chapter on its own page

munotes.in376

Chapter Fifty-Six

Flooding, Gossiping and the Broadcast Storm

Syllabus topic Module 1, "Routing in WSN: Routing strategies in WSNs" (and the paired practical, "Create a multi-node ad-hoc network in TOSSIM and evaluate broadcast communication among nodes")

In one line

Flooding sends every packet to every neighbour and has every receiver do the same, which is simple and robust but wastes energy on duplicates; gossiping forwards to one random neighbour instead, cheaply but slowly; and in a dense network flooding becomes a broadcast storm, cured by rebroadcasting only when it is likely to reach someone new.

In the wording a student can write in an examination: in flooding, a node that receives a packet it has not seen before rebroadcasts it to all its neighbours, until the packet has reached every node or its hop limit; duplicates are recognised by a (source, sequence number) pair and dropped. It needs no routing tables or topology knowledge and finds every path, but it has three faults: implosion (a node receives duplicate copies of the same packet from several neighbours), overlap (nodes sensing overlapping areas send the same data to a common neighbour) and resource blindness (nodes act without regard to their remaining energy). Gossiping forwards each packet to one randomly chosen neighbour, avoiding implosion but spreading data slowly. In a dense network, flooding causes the broadcast storm: redundant rebroadcasts (a rebroadcast adds on average only 41 per cent new coverage, at most 61 per cent, and far less once a node has heard the packet several times), contention among neighbours that rebroadcast together, and collisions, since broadcasts have no RTS/CTS or acknowledgement. Its cures are probabilistic, counter-based, distance-based, location-based and cluster-based rebroadcasting.

Classic flooding

The SPIN paper describes the baseline: "classic flooding, start with a source node sending its data to all of its neighbors. Upon receiving a piece of data, each node then stores and sends a copy of the data to all of its neighbors. This is therefore a straightforward protocol requiring no protocol state at any node, and it disseminates data quickly in a network where bandwidth is not scarce and links are not loss-prone."

Two rules keep it finite. A node must forward each packet once, so it must recognise a copy it has already seen; the broadcast storm paper states the assumption and the usual method: "we assume that a host can detect duplicate broadcast messages. This is essential to prevent endless flooding of a message. One way to do so is to associate with each broadcast message a tuple (source ID, sequence number)". And a packet may carry a hop limit (a time to live) that each node decrements, so that a packet meant for nearby nodes does not travel the whole network.

What flooding does well is exactly what makes it the backbone of other protocols: it needs no knowledge of the topology, it finds every node that can be reached, and it tries every path at once, so it survives lost links. AODV's route requests, directed diffusion's interests and the query dissemination of [Routing Challenges and Design Issues in WSNs] are all floods.

munotes.in377

Flooding, Gossiping and the Broadcast Storm

Flooding's three faults

The SPIN paper names them.

Implosion. "In classic flooding, a node always sends data to its neighbors, regardless of whether or not the neighbor has already received the data from another source." In its Figure 1, A floods to B and C, which both pass the data to D: "The protocol thus wastes resources by sending two copies of the data to D. It is easy to see that implosion is linear in the degree of any node." A node with ten neighbours may hear the same packet ten times.

Overlap. "Sensor nodes often cover overlapping geographic areas, and nodes often gather overlapping pieces of sensor data." Two sensors watching the same spot flood the same observation, and their common neighbour receives it twice. "Overlap is a harder problem to solve than the implosion problem" because "implosion is a function only of network topology, whereas overlap is a function of both topology and the mapping of observed data to sensor nodes."

Resource blindness. "In classic flooding, nodes do not modify their activities based on the amount of energy available to them at a given time." A node nearly out of energy forwards as eagerly as a fresh one.

Left: nodes A, B, C and D; A sends to B and C, and both send to D, which is shaded. Right: two overlapping circles, centred on A and on B at distance r apart; the part of B's circle outside A's is shaded

Figure 56.1 Implosion (after SPIN's Fig. 1), and how little of B's disc a rebroadcast newly reaches (after the storm paper's Fig. 2)

Gossiping

Gossiping "is an alternative to the classic flooding approach that uses randomization to conserve energy. Instead of indiscriminately forwarding data to all its neighbors, a gossiping node only forwards data on to one randomly selected neighbor." It may even send the data straight back to the neighbour it came from, which the paper allows on purpose, since otherwise a node reachable only through that neighbour might never receive it.

Because each node makes only one copy, "Gossiping avoids such implosion". The price is speed: "While gossiping distributes information slowly, it dissipates energy at a slow rate as well." With a single source, the data reaches at most one new node per round. Al-Karaki and Kamal agree: selecting one random node instead of broadcasting "causes delays in propagation of data through the nodes."

The word gossip is used more loosely too: many papers mean probabilistic flooding, where each node rebroadcasts with some probability p. The broadcast storm paper calls that the probabilistic scheme, below.

The broadcast storm

Ni, Tseng, Chen and Sheu studied flooding in dense ad hoc networks and found that, used blindly, it causes "serious redundancy, contention, and collision", which together they call the broadcast storm problem.

munotes.in378

Flooding, Gossiping and the Broadcast Storm

  • Redundant rebroadcasts. "When a mobile host decides to rebroadcast a broadcast message to its neighbors, all its neighbors already have the message."
  • Contention. "After a mobile host broadcasts a message, if many of its neighbors decide to rebroadcast the message, these transmissions (which are all from nearby hosts) may severely contend with each other."
  • Collision. With no RTS/CTS dialogue, no acknowledgement and no collision detection for broadcasts, collisions are both likelier and more damaging ([Hidden and Exposed Terminals, and RTS and CTS] explained why broadcasts get no handshake).

How little a rebroadcast adds. Let A broadcast and B, at distance d, rebroadcast; both have range r. The only area that can benefit from B's rebroadcast is the part of B's disc outside A's. It is largest when B is at the edge of A's range, d = r, where it equals r squared × (π/3 + half the square root of 3), about 0.61 πr squared: a rebroadcast can add at most 61 per cent new coverage, and anything from nothing upwards. Averaged over all positions of B within A's range, it is about 0.41 πr squared: "a rebroadcast can cover only additional 41% area in average." And a node that has already heard the message from two neighbours can add, on average, only about 0.19 of its disc. The paper calls the expected share after k copies EAC(k), the expected additional coverage.

The first program reproduces the analysis. It places k senders at random within a node's range and measures, with random points, how much of the node's disc none of them covered.

# Part 1: how little a rebroadcast adds. A node X (range 1, at the origin) has
# heard the same broadcast from k senders placed at random within its range.
# EAC(k) is the expected share of X's disc that none of them covered: what X's
# own rebroadcast could add. Estimated by random points, as Ni and colleagues did.
import random
from math import pi, sqrt

rnd = random.Random(56)

def in_disc(r=1.0):
    while True:
        x, y = rnd.uniform(-r, r), rnd.uniform(-r, r)
        if x * x + y * y <= r * r:
            return x, y

def eac(k, trials=2000, points=200):
    total = 0.0
    for _ in range(trials):
        senders = [in_disc() for _ in range(k)]
        new = 0
        for _ in range(points):
            px, py = in_disc()
            if all((px - sx) ** 2 + (py - sy) ** 2 > 1 for sx, sy in senders):
                new += 1
        total += new / points
    return total / trials

print("a single rebroadcast adds at most %.3f of the disc" % ((pi / 3 + sqrt(3) / 2) / pi))
for k in range(1, 7):
    print("heard %d times: expected additional coverage %.3f" % (k, eac(k)))
munotes.in379

Flooding, Gossiping and the Broadcast Storm

a single rebroadcast adds at most 0.609 of the disc
heard 1 times: expected additional coverage 0.416
heard 2 times: expected additional coverage 0.192
heard 3 times: expected additional coverage 0.093
heard 4 times: expected additional coverage 0.047
heard 5 times: expected additional coverage 0.025
heard 6 times: expected additional coverage 0.014

Reading it. The maximum, 0.609, and the averages after one and two copies, about 0.41 and 0.19, are the paper's 61, 41 and 19 per cent. After three copies a rebroadcast can add about 9 per cent, after four about 5: the paper's statement that for four or more copies "the expected additional coverage is below 0.05%" is right once the percent sign is read as a slip for a fraction of 0.05 of the disc, which its own Figure 3 plots. A node that has heard a broadcast three or four times has almost nothing to add, and its rebroadcast mostly adds contention and collisions.

The cures

The paper's schemes all ask one question before rebroadcasting: is this node likely to reach anyone new?

  • Probabilistic. Rebroadcast with probability P. In dense networks a small P reaches nearly everyone; in sparse ones P must be larger.
  • Counter-based. Wait a random number of slots before rebroadcasting, counting the copies heard meanwhile; if the count reaches a threshold C, the rebroadcast is inhibited. The paper finds that "a threshold C of 3 or 4 is an appropriate choice", which the EAC curve explains.
  • Distance-based. Rebroadcast only if every sender heard is farther than a threshold D, since a near sender leaves little new area; D is matched to the coverage, for example to EAC(2), about 0.187.
  • Location-based. With positions known, compute the additional coverage directly and rebroadcast only if it exceeds a threshold.
  • Cluster-based. Only cluster heads and gateways rebroadcast; ordinary members stay silent.

The random wait before rebroadcasting matters in every scheme: it spreads neighbours' rebroadcasts in time, reducing contention, and gives a node time to count the copies it hears.

One broadcast, four ways

The second program sends one broadcast from a corner node of a random field of 100 nodes (100 m by 100 m, range 20 m), four ways: flooding; probabilistic rebroadcast with P = 0.6; counter-based with C = 3; and SPIN's one-neighbour gossiping, counted until every reachable node has the data. Each rebroadcast waits a random 0 to 7 slots. Collisions are not modelled, so what is counted is redundancy alone. It reports, averaged over 200 random fields, the sends, the duplicate receptions, the share of reachable nodes reached, and the slots taken.

munotes.in380

Flooding, Gossiping and the Broadcast Storm

# Part 2: one broadcast from a corner node of a field of 100 nodes (100 m by
# 100 m, range 20 m), four ways. Time runs in slots; a transmission is heard
# by every neighbour in the slot it is sent (no collisions are modelled, so
# the cost counted is redundancy alone). Averages over 200 random fields.
import random
from collections import deque

N, SIDE, RANGE, WAIT = 100, 100.0, 20.0, 8   # WAIT: rebroadcast delay, 0..7 slots

def field(rnd):
    pts = [(0.0, 0.0)] + [(rnd.uniform(0, SIDE), rnd.uniform(0, SIDE)) for _ in range(N - 1)]
    nbr = [[j for j in range(N) if j != i and
            (pts[i][0] - pts[j][0]) ** 2 + (pts[i][1] - pts[j][1]) ** 2 <= RANGE ** 2]
           for i in range(N)]
    seen, queue = {0}, deque([0])            # the part of the field node 0 can reach
    while queue:
        u = queue.popleft()
        for v in nbr[u]:
            if v not in seen:
                seen.add(v)
                queue.append(v)
    return nbr, len(seen)

def rebroadcast(nbr, rnd, p=1.0, counter=None):
    """Flooding (p = 1), probabilistic (rebroadcast with probability p) or
    counter-based (give up if the message is heard `counter` times first)."""
    heard, due, sent, received = {0: 1}, {0: 0}, 0, 0
    t = 0
    while due:
        for u in [u for u, s in due.items() if s == t]:
            del due[u]
            if counter and heard[u] >= counter:
                continue                          # enough copies heard: stay quiet
            sent += 1
            for v in nbr[u]:
                received += 1
                if v not in heard:
                    heard[v] = 1
                    if rnd.random() < p:
                        due[v] = t + 1 + rnd.randrange(WAIT)
                else:
                    heard[v] += 1
        t += 1
    return sent, received, len(heard), t

def gossip(nbr, rnd, reachable):
    """SPIN's gossiping: the holder sends to ONE random neighbour, which does
    the same. Count the sends until everyone reachable has the message."""
    have, u, sent = {0}, 0, 0
    while len(have) < reachable and sent < 100000:
        u = rnd.choice(nbr[u]) if nbr[u] else u
        have.add(u)
        sent += 1
    return sent, sent, len(have), sent

rnd = random.Random(56)
rows = {"flooding": [], "probabilistic, p = 0.6": [], "counter-based, C = 3": [],
        "gossiping, one neighbour": []}
for _ in range(200):
    nbr, reach = field(rnd)
    runs = (rebroadcast(nbr, rnd), rebroadcast(nbr, rnd, p=0.6),
            rebroadcast(nbr, rnd, counter=3), gossip(nbr, rnd, reach))
    for name, (sent, received, got, slots) in zip(rows, runs):
        rows[name].append((sent, received - got + 1, 100 * got / reach, slots))
print("%-26s %6s %11s %8s %7s" % ("scheme", "sends", "duplicates", "reach %", "slots"))
for name, r in rows.items():
    m = [sum(col) / len(r) for col in zip(*r)]
    print("%-26s %6.1f %11.1f %8.1f %7.1f" % (name, m[0], m[1], m[2], m[3]))
scheme                      sends  duplicates  reach %   slots
flooding                     96.0       887.9    100.0    30.9
probabilistic, p = 0.6       44.3       378.7     77.2    31.4
counter-based, C = 3         41.9       283.2     98.4    33.4
gossiping, one neighbour   1590.2      1495.3    100.0  1590.2
munotes.in381

Flooding, Gossiping and the Broadcast Storm

Reading it: flooding. Every reachable node sends once, 96 sends on average, and they cause about 888 duplicate receptions, some nine for every node: implosion counted. The broadcast finishes in about 31 slots.

Reading it: the cures. The counter-based scheme with C = 3 reaches 98.4 per cent of the nodes with 41.9 sends, less than half of flooding's, and about a third of the duplicates. The probabilistic scheme with P = 0.6 sends about as little (44.3) but reaches only 77.2 per cent: a random decision ignores whether the node has anything to add, while the counter measures it. At this density the counter is the better rule, as the EAC curve predicts.

Reading it: gossiping. One-neighbour gossiping reaches everyone in the end, but needs about 1,590 sends, one at a time, some sixteen times flooding's total and some fifty times as long: cheap per step, costly and slow overall, because a random walk revisits nodes it has already covered.

Collisions would make flooding look worse still: its rebroadcasts, bunched in time among close neighbours, are the ones most likely to collide.

Broadcast in the practical

The paired practical broadcasts among the motes of a small TOSSIM network. The principles above are what its code must implement: a sequence number with each broadcast so that a mote forwards each message once; optionally a hop count; and a small random delay before rebroadcasting, so that neighbours do not all transmit in the same instant. [TOSSIM: Simulating Motes, Radio Gain and Packet Loss] showed how the simulator's radio model loses packets, which is why a broadcast that works on a clean channel may not reach every mote in the simulated one.

Distinctions

FloodingGossiping
Forward toAll neighboursOne random neighbour
Copies per nodeOne send, many receptionsOne
ImplosionYesAvoided
SpeedFast: every path at onceSlow: about one node per round
ReliabilityHighEventually, if run long enough
ImplosionOverlap
Duplicate fromThe same data by different pathsDifferent nodes sensing the same thing
Depends onTopology onlyTopology and what nodes sense
Cured byDuplicate suppression, negotiationNegotiation on the data itself (SPIN)
ProbabilisticCounter-basedDistance-basedLocation-based
Rebroadcast ifA random draw succeedsFewer than C copies heardEvery sender heard is farComputed new coverage is large
NeedsNothingA counter and a random waitDistances (signal strength)Positions

What it does not mean

Flooding is not useless. It is robust and needs no state, which is why route discovery, interests and queries use it; the cures limit it rather than replace it.

Duplicate suppression does not stop implosion. It stops a node forwarding a packet twice; it does not stop the node receiving the packet from every neighbour.

munotes.in382

Flooding, Gossiping and the Broadcast Storm

Gossiping is not cheap overall. Each step is one send, but covering the network takes many steps.

The broadcast storm is not only about energy. Redundant rebroadcasts also cause contention and collisions, which can make a broadcast reach fewer nodes than a more careful scheme.

Quick revision

  • Flooding: forward every new packet to all neighbours; duplicates dropped by (source ID, sequence number); optional hop limit; no state, robust, fast.
  • Three faults (SPIN): implosion (duplicates by several paths, linear in degree), overlap (overlapping sensors send the same data), resource blindness (ignores energy).
  • Gossiping: forward to one random neighbour; avoids implosion; slow (about one node per round).
  • Broadcast storm (Ni, Tseng, Chen, Sheu 1999): redundant rebroadcasts, contention, collision. A rebroadcast adds at most 61 per cent, on average 41 per cent; after two copies about 19, after four about 5.
  • Cures: probabilistic (P), counter-based (C = 3 or 4), distance-based (D), location-based, cluster-based; always a random delay before rebroadcasting.
  • Program (100 nodes): flooding 96 sends, 888 duplicates; counter-based C = 3: 41.9 sends, 98.4 per cent reach; probabilistic P = 0.6: 77.2 per cent; gossiping about 1,590 sends.

Test yourself

1. Explain flooding and its advantages. A source sends a packet to all its neighbours, and every node that receives a packet for the first time sends it on to all of its neighbours, recognising and dropping copies it has already seen by a source ID and sequence number, and optionally stopping at a hop limit. It needs no routing tables or knowledge of the topology, reaches every reachable node, and tries all paths at once, so it is simple and robust to failures.

2. What are implosion, overlap and resource blindness? Implosion: a node receives several copies of the same packet because several neighbours forward it, wasting a send and a receive for each extra copy. Overlap: nodes whose sensing areas overlap produce the same data and send it to the same neighbour, which receives it twice. Resource blindness: nodes forward without regard to their remaining energy, so nearly exhausted nodes are used as freely as full ones.

3. What is gossiping, and how does it compare with flooding? A gossiping node forwards a packet to one randomly chosen neighbour rather than to all of them. This avoids implosion, since each node makes only one copy, and uses little energy per step, but the data spreads slowly, about one node per round, and covering the whole network can take many transmissions. In the program, gossiping needed about 1,590 sends to reach every node, against flooding's 96.

munotes.in383

Flooding, Gossiping and the Broadcast Storm

4. What is the broadcast storm problem? When flooding is used in a dense network, many nodes rebroadcast the same message close together in space and time, causing redundant rebroadcasts (the neighbours already have the message), contention for the channel among the rebroadcasting neighbours, and collisions, which are likely because broadcasts use no RTS/CTS, no acknowledgement and no collision detection.

5. Show why a rebroadcast is often redundant. If A broadcasts and B, within A's range, rebroadcasts, only the part of B's disc outside A's disc can gain new receivers. This area is largest when B is at the edge of A's range, about 61 per cent of B's disc, and averages about 41 per cent over all positions of B. After a node has heard the message twice, a rebroadcast adds on average only about 19 per cent, and after four times about 5 per cent.

6. Explain the counter-based scheme and why a threshold of 3 or 4 works. After first hearing a broadcast, a node waits a random number of slots and counts how many more copies it hears. If the count reaches the threshold C before its turn, it cancels its rebroadcast. Since the expected additional coverage after three or four copies is only about 9 or 5 per cent, a node that has heard the message that often has almost nothing to add. In the program, C = 3 reached 98.4 per cent of nodes with less than half of flooding's sends.

Contents This chapter on its own page

munotes.in384

Chapter Fifty-Seven

SPIN: Negotiating Before Sending

Syllabus topic Module 1, "Routing in WSN: Routing strategies in WSNs"

In one line

In SPIN a node with new data first advertises a short description of it, sends the data only to neighbours that ask for it, and uses its remaining energy to decide how much to take part, so that nothing is sent twice and nothing is sent that nobody wants.

In the wording a student can write in an examination: SPIN (Sensor Protocols for Information via Negotiation, Heinzelman, Kulik and Balakrishnan, 1999) is a family of flat, negotiation-based protocols that disseminate each node's data to all nodes, treating every node as a potential sink. It rests on two ideas: nodes exchange meta-data, short descriptions of their data, before exchanging the data itself; and nodes are resource-aware, adapting to their remaining energy. SPIN uses three messages: ADV (an advertisement carrying the meta-data of new data), REQ (a request for the data) and DATA (the data itself, with a meta-data header). In SPIN-1, a node with new data sends an ADV to its neighbours; a neighbour that has not already received or requested that data replies with a REQ; the advertiser sends the DATA; and the receiver then advertises it to its own neighbours. SPIN-2 adds a low-energy threshold: near it, a node takes part in a round only if it can complete all its stages. Negotiation solves flooding's implosion and overlap, and resource awareness its resource blindness. Its main limit is that advertisements do not guarantee delivery: data may not reach a node that wants it if the nodes in between are not interested.

Two ideas

SPIN was designed after an analysis of flooding and gossiping ([Flooding, Gossiping and the Broadcast Storm]). Its authors state its basis in two sentences. "First, to operate efficiently and to conserve energy, sensor applications need to communicate with each other about the data that they already have and the data they still need to obtain. Exchanging sensor data may be an expensive network operation, but exchanging data about sensor data need not be. Second, nodes in a network must monitor and adapt to changes in their own energy resources to extend the operating lifetime of the system."

The task SPIN solves is dissemination: getting every node's observations to every other node, "treating all sensors as potential sink nodes", which replicates the whole network's view for fault tolerance, or spreads an alarm everywhere.

Meta-data

"Sensors use meta-data to succinctly and completely describe the data that they collect." Three rules make it useful: the meta-data x for data X must be shorter than X, or SPIN gains nothing; distinguishable data must have distinguishable meta-data; and "two pieces of indistinguishable data should share the same meta-data representation."

SPIN leaves the format to the application. Sensors covering separate areas can use their IDs, so that x means "all the data gathered by sensor x"; a camera might use a position and an orientation. This is what lets SPIN cure overlap: two sensors that saw the same thing name it the same way, so a node that already has it can say so.

munotes.in385

SPIN: Negotiating Before Sending

The three messages

  • ADV, "new data advertisement. When a SPIN node has data to share, it can advertise this fact by transmitting an ADV message containing meta-data."
  • REQ, "request for data. A SPIN node sends an REQ message when it wishes to receive some actual data."
  • DATA, "data message. DATA messages contain actual sensor data with a meta-data header."

"Because ADV and REQ messages contain only meta-data, they are smaller, and cheaper to send and receive, than their corresponding DATA messages." In the paper's tests, meta-data was 16 bytes and data 500.

SPIN-1: the three-stage handshake

"The SPIN-1 protocol is a simple handshake protocol for disseminating data through a lossless network. It works in three stages (ADV-REQ-DATA)." A node with new data sends an ADV naming it; "Upon receiving an ADV, the neighboring node checks to see whether it has already received or requested the advertised data. If not, it responds by sending an REQ message for the missing data back to the sender"; the advertiser answers with a DATA message. The new holder then advertises in turn, to everyone but the node it came from.

Six small panels of nodes A, B, C, D and E, with B at the centre. 1: A sends an ADV to B. 2: B sends a REQ to A. 3: A sends DATA to B. 4: B sends ADVs to C, D and E. 5: C and D send REQs to B, while E, which already has the data, stays silent. 6: B sends DATA to C and D

Figure 57.1 SPIN-1's handshake in six steps, after the paper's Fig. 3

Two remarks from the paper complete it. A node that has data of its own may aggregate it with the data it received and advertise the combination. And a node need not answer every message: E in the figure, which already has the data, sends no REQ, which is how negotiation prevents implosion. Losses can be handled by re-advertising periodically and re-requesting data that does not arrive; in a mobile network, a node whose neighbour list changes can re-advertise all its data.

SPIN-1's strength is its simplicity: "each node only needs to know about its single-hop network neighbors", so it runs in an unconfigured network, and topology changes "only have to travel one hop".

SPIN-2: a low-energy threshold

"The SPIN-2 protocol adds a simple energy-conservation heuristic to the SPIN-1 protocol." With plenty of energy it behaves as SPIN-1. "When a SPIN-2 node observes that its energy is approaching a low-energy threshold, it adapts by reducing its participation in the protocol. In general, a node will only participate in a stage of the protocol if it believes that it can complete all the other stages of the protocol without going below the low-energy threshold." So a node low on energy does not start a round it cannot finish with all its neighbours, and does not request data it cannot afford to receive; it can still receive ADVs and REQs, but never handles a DATA message below the threshold. This cures resource blindness.

munotes.in386

SPIN: Negotiating Before Sending

The later family

Al-Karaki and Kamal's survey lists four more members, from the journal version of the work:

  • SPIN-PP, for point-to-point communication, hop by hop;
  • SPIN-EC, "similar to SPIN-PP, but with an energy heuristic added to it";
  • SPIN-BC, for broadcast channels;
  • SPIN-RL, for lossy channels, with adjustments to SPIN-PP's protocol.

What negotiation saves

The paper tested five protocols on a 25-node network: SPIN-1, SPIN-2, flooding, gossiping and an ideal protocol with perfect knowledge. Its headline: "SPIN-1 uses negotiation to solve the implosion and overlap problems; it reduces energy consumption by a factor of 3.5 compared to flooding, while disseminating data almost as quickly as theoretically possible." SPIN-2, with limited energy, "disseminates 60% more data per unit energy than flooding".

The program repeats the first comparison with the paper's own settings from its Table 1: 25 nodes, 1 Mbit/s, 600 mW to transmit and 200 mW to receive, data of 500 bytes named by 16 bytes of meta-data, and each node starting with 3 items drawn at random from 25, so that nodes' data overlaps. The network is our own random one of the same size and density. Every message goes to one neighbour and costs its bytes at both ends. Flooding copies each new item to every neighbour except the one it came from; SPIN-1 advertises, and sends data only on request; the ideal protocol delivers each missing copy across one link exactly once.

# SPIN-1 against flooding, with the SPIN paper's own test settings (its Table 1):
# 25 nodes, 1 Mbit/s, 600 mW to transmit and 200 mW to receive, data items of
# 500 bytes with a 16-byte meta-data name, each node starting with 3 items drawn
# from 25, so the nodes' data overlaps. Every message goes to one neighbour and
# costs its bytes at both ends. The network is our own random one, not theirs.
import random
from collections import deque

rnd = random.Random(57)
N, ITEMS, EACH, SIDE, RANGE = 25, 25, 3, 38.0, 10.0
TX, RX = 600 * 8e-6, 200 * 8e-6          # mJ per byte: mW x 8 microseconds a byte
DATA, META = 500 + 16, 16                 # bytes: a DATA message, an ADV or a REQ

def connected_field():
    while True:
        pts = [(rnd.uniform(0, SIDE), rnd.uniform(0, SIDE)) for _ in range(N)]
        nbr = [[j for j in range(N) if j != i and (pts[i][0] - pts[j][0]) ** 2 +
                (pts[i][1] - pts[j][1]) ** 2 <= RANGE ** 2] for i in range(N)]
        seen, q = {0}, deque([0])
        while q:
            for v in nbr[q.popleft()]:
                if v not in seen:
                    seen.add(v)
                    q.append(v)
        if len(seen) == N:
            return nbr

def flooding(nbr, start):
    """A node that gets an item for the first time copies it to every neighbour
    except the one it came from; duplicates are received, and cost, anyway."""
    have = [set(s) for s in start]
    q = deque((u, x, None) for u in range(N) for x in start[u])
    msgs = energy = 0
    while q:
        u, x, came = q.popleft()
        for v in nbr[u]:
            if v == came:
                continue
            msgs += 1
            energy += DATA * (TX + RX)
            if x not in have[v]:
                have[v].add(x)
                q.append((v, x, u))
    return msgs, energy, have

def spin1(nbr, start):
    """ADV to the neighbours; a neighbour that neither has nor has asked for the
    item answers with a REQ; the advertiser sends DATA to each requester."""
    have = [set(s) for s in start]
    asked = [set(s) for s in start]
    q = deque((u, x, None) for u in range(N) for x in start[u])
    count = {"ADV": 0, "REQ": 0, "DATA": 0}
    energy = 0.0
    while q:
        u, x, came = q.popleft()
        for v in nbr[u]:
            if v == came:
                continue
            count["ADV"] += 1
            energy += META * (TX + RX)
            if x not in asked[v]:
                asked[v].add(x)
                count["REQ"] += 1
                count["DATA"] += 1
                energy += META * (TX + RX) + DATA * (TX + RX)
                have[v].add(x)
                q.append((v, x, u))
    return count, energy, have

nbr = connected_field()
edges = sum(len(n) for n in nbr) // 2
start = [set(rnd.sample(range(ITEMS), EACH)) for _ in range(N)]
needed = sum(len(set().union(*start)) - len(s) for s in start)
print("our network: %d nodes, %d edges, average degree %.1f" % (N, edges, 2 * edges / N))
print("item copies still to deliver: %d" % needed)
f_msgs, f_energy, f_have = flooding(nbr, start)
s_count, s_energy, s_have = spin1(nbr, start)
ideal = needed * DATA * (TX + RX)          # each missing copy crosses one link once
assert f_have == s_have                    # both give every node every item
print("flooding : %5d DATA messages               %8.1f mJ" % (f_msgs, f_energy))
print("SPIN-1   : %5d ADV %5d REQ %5d DATA   %8.1f mJ" % (s_count["ADV"], s_count["REQ"],
                                                          s_count["DATA"], s_energy))
print("ideal    : %5d DATA messages               %8.1f mJ" % (needed, ideal))
print("flooding uses %.1f times SPIN-1's energy; SPIN-1 uses %.2f times the ideal"
      % (f_energy / s_energy, s_energy / ideal))
munotes.in387

SPIN: Negotiating Before Sending

our network: 25 nodes, 56 edges, average degree 4.5
item copies still to deliver: 525
flooding :  2163 DATA messages                 7143.1 mJ
SPIN-1   :  2163 ADV   525 REQ   525 DATA     2009.0 mJ
ideal    :   525 DATA messages                 1733.8 mJ
flooding uses 3.6 times SPIN-1's energy; SPIN-1 uses 1.16 times the ideal
munotes.in388

SPIN: Negotiating Before Sending

Reading it. Our network has 56 edges and an average degree of 4.5, close to the paper's 59 and 4.7. Both protocols deliver every item to every node (the program checks), but flooding needs 2,163 DATA messages to deliver 525 missing copies: each copy crosses about four links, most of them to nodes that already have it. SPIN-1 sends exactly one DATA message per missing copy, 525, the same as the ideal protocol; what it pays extra is 2,163 small ADVs and 525 REQs of 16 bytes each.

The factor. Flooding spends 3.6 times SPIN-1's energy here, against the paper's 3.5 on its own network; and SPIN-1 spends only 1.16 times the ideal. Negotiation works because the expensive message is sent only when it is wanted; the cheap ones may be wasted freely. The larger the data compared with its description, the greater the saving: with meta-data as large as the data, SPIN would gain nothing.

SPIN's limits

No guarantee of delivery. A node that wants the data can get it only if a neighbour advertises it, and a neighbour advertises only data it has itself requested. Al-Karaki and Kamal point out the consequence: "SPINs data advertisement mechanism cannot guarantee the delivery of data", for example when the nodes interested in some data are far from the source and the nodes in between are not interested.

Everyone gets everything. SPIN disseminates to all nodes, which suits replication but not a network where one sink wants a few readings; directed diffusion, next chapter, lets the sink ask ([Directed Diffusion and Rumour Routing]).

Meta-data is the application's problem. SPIN defines no format, so every application must design one that satisfies its rules.

Distinctions

FloodingSPIN-1
First messageThe data itselfAn ADV with meta-data
Who gets the dataEvery neighbour, wanted or notOnly neighbours that request it
ImplosionYesNo: a node that has the data sends no REQ
OverlapYesNo: the same data has the same meta-data
Energy (program)3.6 times SPIN-1's1.16 times the ideal
SPIN-1SPIN-2
HandshakeADV, REQ, DATAThe same
EnergyIgnoredA low-energy threshold limits participation
CuresImplosion and overlapAlso resource blindness

What it does not mean

SPIN is not a route to a sink. It spreads data to every node, one hop at a time; it builds no end-to-end routes.

Negotiation is not free. Every new item costs an ADV to every neighbour; it pays only because ADVs are much smaller than the data.

SPIN-2 does not stop a tired node receiving. It still hears ADVs and REQs below its threshold; it only refuses to take part in rounds it cannot finish.

munotes.in389

SPIN: Negotiating Before Sending

A node's silence is not a failure. A node that already has the advertised data simply does not reply; that silence is the saving.

Quick revision

  • SPIN (Heinzelman, Kulik, Balakrishnan, 1999): flat, negotiation-based, disseminates every node's data to all nodes.
  • Two ideas: exchange meta-data before data; be resource-aware.
  • Meta-data: shorter than the data; distinct data, distinct names; identical data, identical names; format chosen by the application.
  • Messages: ADV (meta-data), REQ (request), DATA (data with a meta-data header).
  • SPIN-1: ADV to neighbours; REQ from those who have neither received nor requested it; DATA to the requesters; the receiver advertises in turn. Knows only one-hop neighbours.
  • SPIN-2: SPIN-1 plus a low-energy threshold: take part only in rounds you can finish.
  • Later family (survey): SPIN-PP, SPIN-EC, SPIN-BC, SPIN-RL.
  • Paper: SPIN-1 cuts energy by a factor of 3.5 against flooding; SPIN-2 delivers 60 per cent more data per unit energy. Program: 3.6, and SPIN-1 at 1.16 times the ideal.
  • Limit: no guaranteed delivery.

Test yourself

1. What are the two basic ideas of SPIN? First, nodes should communicate about the data they have and the data they need before exchanging it, by exchanging meta-data, short descriptions of their data, since exchanging data about data is cheap while exchanging the data itself is expensive. Second, nodes must monitor and adapt to their own energy resources to extend the network's lifetime.

2. Name and explain SPIN's three messages. ADV advertises new data a node has, carrying only its meta-data. REQ is sent by a node that wants the advertised data, requesting it. DATA carries the actual sensor data with a meta-data header. ADV and REQ are small and cheap compared with DATA.

3. Explain the SPIN-1 handshake with an example. Node A, with new data, sends an ADV naming it to its neighbour B. B checks whether it has already received or requested that data; if not, it sends a REQ back to A. A sends the DATA to B. B then advertises the data to its own neighbours except A; those that lack it send REQs and receive the data, while a neighbour that already has it stays silent, and the process continues across the network.

4. How does SPIN overcome implosion, overlap and resource blindness? Implosion is avoided because data is sent only to a node that has asked for it, and a node that already has the data does not ask. Overlap is avoided because identical data from different sensors has the same meta-data, so a node that has one copy does not request another. Resource blindness is addressed by SPIN-2, whose nodes reduce their participation as their energy approaches a low-energy threshold.

munotes.in390

SPIN: Negotiating Before Sending

5. What is SPIN-2? SPIN-1 with an energy-conservation heuristic. When energy is plentiful it behaves exactly like SPIN-1; when a node's energy approaches a low-energy threshold, it takes part in a stage of the protocol only if it believes it can complete all the remaining stages without going below the threshold, so it never handles a DATA message below it.

6. Give one limitation of SPIN. Its advertisements cannot guarantee delivery: a node forwards only data it has itself requested, so if the nodes between a source and an interested node are not interested in the data, the interested node may never receive it. SPIN also sends every node's data to every node, which is wasteful when only a sink needs a few readings.

Contents This chapter on its own page

munotes.in391

Chapter Fifty-Eight

Directed Diffusion and Rumour Routing

Syllabus topic Module 1, "Routing in WSN: Routing strategies in WSNs"

In one line

In directed diffusion the sink spreads an interest describing the data it wants, every node remembers which neighbours asked (a gradient), sources send data back down the gradients, and the sink reinforces the best path; in rumour routing, agents leave trails from each event and a query walks straight until it crosses one.

In the wording a student can write in an examination: directed diffusion (Intanagonwiwat, Govindan and Estrin) is a flat, data-centric, query-based protocol. Data is named by attribute-value pairs (such as type = wheeled vehicle, a rectangle, an interval). The sink broadcasts an interest, a named task description, which each node caches and rebroadcasts; each node sets up a gradient (a direction and a data rate) towards each neighbour it received the interest from, so gradients form on every link. A source that detects a matching event sends exploratory data at a low rate along all its gradients; nodes drop duplicates using a data cache. The sink then positively reinforces the neighbour that delivered the first copy, by resending the interest at a higher rate; each reinforced node does the same, so a single low-delay path carries the high-rate data. Other paths time out or are negatively reinforced. Nodes may aggregate data along the way, and all decisions are local interactions between neighbours.

Rumour routing (Braginsky and Estrin) suits many events and few queries: each event launches agents, long-lived packets that walk roughly straight for a set number of hops, leaving event path state in the nodes they pass; a query walks straight in a random direction until it meets a node on a path to its event, then follows that path. Two straight lines in a bounded region are likely to cross, so a few agents per event make delivery very likely, at a cost between query flooding (N transmissions per query) and event flooding (N per event).

Directed diffusion: naming the data

The paper names four elements: "Directed diffusion consists of several elements: interests, data messages, gradients, and reinforcements." It begins with naming. "In directed diffusion, data is named using attribute-value pairs." A vehicle-tracking task might be described by a type (wheeled vehicle), an interval (send an event every 20 ms), a duration (for the next 10 s) and a rect (sensors within a given rectangle). "Intuitively, the task description specifies an interest for data matching the attributes. For this reason, such a task description is called an interest." The data sent back is named the same way: type, instance, location, intensity, confidence and a timestamp.

Naming by content is what makes the protocol data-centric ([Design Principles: Data Centricity, Location, Activity and Heterogeneity]): nobody asks for node 37; the sink asks for wheeled vehicles in a rectangle, and whichever nodes see one answer.

munotes.in392

Directed Diffusion and Rumour Routing

Interests and gradients

"A sensing task (or a subtask thereof) is disseminated throughout the sensor network as an interest for named data." The sink broadcasts the interest periodically, first with a low data rate, "Intuitively, this initial interest may be thought of as exploratory". The interest is soft state: the sink refreshes it with a newer timestamp, because interests are not transmitted reliably, and entries expire unless refreshed.

Each node keeps an interest cache. When an interest arrives, a node that has no matching entry creates one, with a gradient towards the neighbour it came from. "Specifically, a gradient is direction state created in each node that receives an interest. The gradient direction is set toward the neighboring node from which the interest is received." A gradient holds a data rate and a duration. The node may then resend the interest to its neighbours, and "To its neighbors, this interest appears to originate from the sending node, although it might have come from a distant sink. This is an example of a local interaction."

Because a node cannot tell whether an interest from a neighbour is its own coming back or another sink's, flooding the interest makes "every pair of neighboring nodes establishes a gradient toward each other." These two-way gradients cost duplicate low-rate events, but they give the network alternative paths for repair and for choosing better ones. In the paper's summary, "interest propagation sets up state in the network (or parts thereof) to facilitate 'pulling down' data toward the sink."

Four panels on a small grid. (a) Arrows spread outwards from the sink, a square at the bottom left, to every node. (b) Every link drawn bold: a gradient each way. (c) Bold arrows from the source, top right, along one path to the sink. (d) A shaded event with two dashed agent trails, one along a row and one down a column, and a query that walks straight up from its node until it meets a trail

Figure 58.1 Directed diffusion's three stages (after the paper's Fig. 1), and rumour routing (schematic)

Data, exploratory events and reinforcement

A node inside the interest's rectangle turns on its sensors. When it detects a matching target, it sends events at the highest rate any of its gradients requests; at first, the exploratory rate. "This data message is, in effect, unicast individually to the relevant neighbors." A node receiving a data message looks for a matching interest; if there is none, it drops the message. If there is, it checks its data cache: "If a received data message has a matching data cache entry, the data message is silently dropped. Otherwise, the received message is added to the data cache and the data message is resent to the node's neighbors." The data cache is what stops loops, and it also lets a node lower the rate on a gradient that asked for less.

The low-rate events serve a purpose: "We call these exploratory events, since they are intended for path setup and repair." Once the sink starts receiving them, it reinforces one neighbour, "in order to 'draw down' real data", by resending the interest to it with a smaller interval, a higher rate. The rule the paper evaluates is simple: reinforce the neighbour from which the first copy of a new event arrived. That neighbour, seeing a higher rate requested of it, reinforces in turn the neighbour that first gave it the event, and so on back to the source. "The local rule we described above, then, selects an empirically low-delay path".

munotes.in393

Directed Diffusion and Rumour Routing

Negative reinforcement removes paths that are no longer best. "One mechanism for negative reinforcement is soft state, i.e., to time out all data gradients in the network unless they are explicitly reinforced." The alternative the paper evaluates is to send an explicit negative reinforcement to the neighbour whose path should be degraded back to exploratory.

Local repair. Reinforcement can also be started by an intermediate node. If the link from the source to a node C degrades, C can reinforce another neighbour to find a new path, and later negatively reinforce the bad link, without waiting for the sink.

Aggregation. Intermediate nodes may combine data, for example to pinpoint a target more accurately from several sensors' reports, and the data cache lets them suppress duplicates.

Diffusion, counted

The first program runs directed diffusion's set-up on a 10 by 10 grid, each node hearing its four neighbours, with the sink at one corner and a source at the opposite one. The interest floods the grid, leaving gradients on every link in both directions; each exploratory event floods as well, since gradients point every way; and reinforcement follows the neighbours that delivered the first copies back from the sink. The traffic is illustrative: 10 events a second for 60 seconds, with the interest refreshed, and an exploratory event sent, every 10 seconds. It compares the total with flooding every event.

# Directed diffusion on a 10 by 10 grid, each node hearing its four neighbours.
# The sink (0, 0) floods an interest; every node records a gradient towards
# each neighbour it heard the interest from. The source (9, 9) sends exploratory
# events down every gradient; the sink reinforces the neighbour that delivered
# the first copy, and each reinforced node does the same, back to the source.
import random
from collections import deque

W = 10
nodes = [(x, y) for x in range(W) for y in range(W)]
nbr = {(x, y): [(x + dx, y + dy) for dx, dy in ((1, 0), (-1, 0), (0, 1), (0, -1))
                if 0 <= x + dx < W and 0 <= y + dy < W] for x, y in nodes}
SINK, SOURCE = (0, 0), (W - 1, W - 1)
rnd = random.Random(58)

def flood(start):
    """Everyone rebroadcasts once. Returns transmissions and, for each node,
    the neighbour its FIRST copy came from (random delays break ties)."""
    heard, q, sends = {}, [(0.0, start, None)], 0
    while q:
        q.sort()
        when, u, came = q.pop(0)
        if u in heard:
            continue
        heard[u] = came
        sends += 1
        for v in nbr[u]:
            if v not in heard:
                q.append((when + 1 + rnd.random() * 0.5, v, u))
    return sends, heard

interest_sends, _ = flood(SINK)
gradients = sum(len(nbr[u]) for u in nodes)        # two-way: every link, both ways
explore_sends, came_from = flood(SOURCE)           # exploratory events go everywhere
path, u = [SINK], SINK                             # reinforce back towards the source
while u != SOURCE:
    u = came_from[u]
    path.append(u)
hops = len(path) - 1

print("interest flood: %d transmissions; %d gradients set up" % (interest_sends, gradients))
print("each exploratory event floods too: %d transmissions" % explore_sends)
print("reinforced path: %d hops (the fewest possible is %d)"
      % (hops, (W - 1) * 2))
events, refresh = 600, 6            # 10 events a second for 60 s; refreshed every 10 s
diffusion = refresh * (interest_sends + explore_sends) + events * hops
flooding = events * len(nodes)
print("60 s at 10 events a second: flooding every event %d transmissions," % flooding)
print("directed diffusion %d (%d interests and exploratory, %d data): %.1f times fewer"
      % (diffusion, refresh * (interest_sends + explore_sends), events * hops,
         flooding / diffusion))
munotes.in394

Directed Diffusion and Rumour Routing

interest flood: 100 transmissions; 360 gradients set up
each exploratory event floods too: 100 transmissions
reinforced path: 18 hops (the fewest possible is 18)
60 s at 10 events a second: flooding every event 60000 transmissions,
directed diffusion 12000 (1200 interests and exploratory, 10800 data): 5.0 times fewer

Reading it. The interest costs one flood, 100 transmissions, and leaves 360 gradients, one each way on each of the grid's 180 links. Each exploratory event floods too. But reinforcement, following first arrivals, finds an 18-hop path, the fewest possible from corner to corner, with no node knowing the topology: first arrival is shortest delay. Over a minute, the set-up and refresh traffic costs 1,200 transmissions and the data 10,800, five times fewer than flooding all 600 events. The saving grows with the data rate and the length of the task, and shrinks if the sink wants only a few events, when the floods are a larger share.

Rumour routing

Directed diffusion floods queries (as interests). Braginsky and Estrin asked what to do when queries are many but each needs little data, and events are few, or the reverse. Query flooding costs N transmissions per query, whatever the number of events. Event flooding, in which each event floods the network once to set up gradients towards it, costs N per event, after which queries travel cheaply along the gradients. Their proposal, "a logical compromise between flooding queries and flooding event notifications", fills the gap between.

munotes.in395

Directed Diffusion and Rumour Routing

The idea: rumour routing "employs a set of long-lived agents that create paths (in the form of state in nodes) directed towards the events they encounter." A node that witnesses an event records it in its events table at distance zero and, with some probability, launches an agent. "The agent travels the network for some number of hops (La), and then dies." At each node the agent synchronises its own events table with the node's, so that the node learns, and the agent carries, the shortest known way to every event; an agent that crosses another event's trail carries both from then on. And since transmissions are broadcasts, "neighboring nodes can hear the agent as it moves along its path", so "the agent actually leaves a fairly thick path as it travels."

A query from any node that has no route to its event is forwarded in a random direction, and "This continues until the query TTL (Lq) expires, or until the query reaches a node that has observed the target event"; on meeting a node that knows a route, it follows it. Straightness matters: "Simulations show that forwarding queries along a straight path yields better results than random forwarding", and agents too avoid nodes they have just visited, so that each walk covers new ground. A query that fails can be resent or, as a last resort, flooded.

Why should a query find a path? Because two roughly straight lines in a field usually cross. "Monte-Carlo simulations show the probability of two lines intersecting in a bounded rectangular region to be approximately 69%. This means five paths leading to an event will have a 99.7% chance of being encountered by a query." The second figure follows from the first: the chance of missing five independent paths is 0.31 to the power 5, about 0.003. The paper does not say how its lines were drawn; with lines through a random point in a random direction across a square, the crossing chance is about 57 per cent, and five paths still give about 98.5 per cent. The heuristic holds either way.

Rumour routing, counted

The second program builds a 40 by 40 grid (1,600 nodes) with 10 events at random nodes. Each event sends 1, 2, 5 or 10 agents, each walking straight for 60 hops and turning only at the edge of the field, recording the way back in every node it visits and every neighbour that overhears it. Then 1,000 queries go from random nodes to random events, each walking straight for up to 60 hops; a query that finds no trail is flooded instead, at a cost of N. The program reports the share delivered without flooding, the mean transmissions per query, and the totals for 100 and 1,000 queries, against query flooding and event flooding (whose queries, as in the paper, are counted as free).

munotes.in396

Directed Diffusion and Rumour Routing

# Rumour routing on a 40 by 40 grid (1,600 nodes, four neighbours each), with
# 10 events. Each event sends out `agents` agents; an agent walks straight,
# turning only at the edge of the field, for LA hops, and every node it visits,
# and every neighbour that overhears it, records the way back to the event. A
# query from a random node to a random event walks straight in a random
# direction until it meets a node that knows the way, then follows it; after
# LQ hops without success it fails and the query is flooded instead.
import random

W, EVENTS, LA, LQ, QUERIES = 40, 10, 60, 60, 1000
N = W * W
DIRS = ((1, 0), (-1, 0), (0, 1), (0, -1))
rnd = random.Random(58)
inside = lambda x, y: 0 <= x < W and 0 <= y < W

def straight_walk(start, hops):
    """Yield (node, previous node) along a straight walk that turns at edges."""
    (x, y), d, prev = start, rnd.choice(DIRS), None
    for _ in range(hops):
        if not inside(x + d[0], y + d[1]):
            d = rnd.choice([e for e in DIRS if inside(x + e[0], y + e[1]) and e != (-d[0], -d[1])])
        prev, (x, y) = (x, y), (x + d[0], y + d[1])
        yield (x, y), prev

def run(agents):
    events = rnd.sample([(x, y) for x in range(W) for y in range(W)], EVENTS)
    way = {}                                   # (node, event) -> (next hop, hops to event)
    for e in events:
        way[(e, e)] = (None, 0)
        for _ in range(agents):
            dist = 0
            for node, prev in straight_walk(e, LA):
                dist += 1
                if (node, e) not in way or way[(node, e)][1] > dist:
                    way[(node, e)] = (prev, dist)
                for dx, dy in DIRS:            # neighbours overhear the agent
                    nb = (node[0] + dx, node[1] + dy)
                    if inside(*nb) and ((nb, e) not in way or way[(nb, e)][1] > dist + 1):
                        way[(nb, e)] = (node, dist + 1)
    cost, found = 0, 0
    for _ in range(QUERIES):
        src, target = (rnd.randrange(W), rnd.randrange(W)), rnd.choice(events)
        hops, at = 0, src
        if (at, target) not in way:
            for at, _ in straight_walk(src, LQ):
                hops += 1
                if (at, target) in way:
                    break
        if (at, target) in way:
            found += 1
            cost += hops + way[(at, target)][1]
        else:
            cost += hops + N                   # give up and flood the query
    return found / QUERIES, cost / QUERIES, EVENTS * agents * LA

print("transmissions for 10 events and Q queries:  Q = 100    Q = 1,000")
print("query flooding (N per query)              %9d %10d" % (N * 100, N * 1000))
print("event flooding (N per event)              %9d %10d" % (N * EVENTS, N * EVENTS))
for agents in (1, 2, 5, 10):
    ok, per_query, setup = run(agents)
    print("rumour, %2d agents: %5.1f%% delivered, %5.1f hops %6.0f %10.0f"
          % (agents, 100 * ok, per_query, setup + 100 * per_query, setup + 1000 * per_query))
munotes.in397

Directed Diffusion and Rumour Routing

transmissions for 10 events and Q queries:  Q = 100    Q = 1,000
query flooding (N per query)                 160000    1600000
event flooding (N per event)                  16000      16000
rumour,  1 agents:  62.4% delivered, 657.2 hops  66319     657788
rumour,  2 agents:  88.3% delivered, 242.7 hops  25474     243942
rumour,  5 agents:  99.8% delivered,  51.7 hops   8171      54710
rumour, 10 agents:  99.9% delivered,  44.7 hops  10465      50651

Reading it: the agents. With one agent per event, only 62.4 per cent of queries meet a trail, and the rest must be flooded, so the mean cost per query is 657.2 transmissions. With five agents, 99.8 per cent are delivered by their walk, in about 52 hops each; the paper's own simulations, with far larger networks, found the same pattern, a few agents per event being enough. Ten agents add little.

Reading it: where it pays. With 10 queries per event (100 queries), rumour routing with 5 agents costs 8,171 transmissions, against 16,000 for event flooding and 160,000 for query flooding: it wins. With 100 queries per event (1,000 queries), event flooding's 16,000 beats rumour routing's 54,710, because a flood per event is cheap once shared among many queries. This is the paper's region: rumour routing "is only useful if the number of queries compared to the number of events is between the two intersection points".

Distinctions

Directed diffusionRumour routing
What is floodedThe interest (a query)Nothing: agents walk lines
State in nodesInterest cache, gradients, data cacheEvents table with the way to each event
Path chosenBy reinforcement: empirically lowest delayThe trail the query first meets
SuitsContinuous, high-rate data for a taskMany events, few queries, little data per query
DeliveryVery likely (the interest reaches everyone)Probable, tunable by the number of agents
Exploratory dataReinforced data
RateLowThe rate the task asked for
PathsAll gradientsOne (or a few) reinforced paths
PurposePath set-up and repairDelivering the data
Query floodingEvent floodingRumour routing
CostN per queryN per eventAgents per event plus a walk per query
Cheapest whenEvents many, queries fewQueries many per eventIn between
munotes.in398

Directed Diffusion and Rumour Routing

What it does not mean

Directed diffusion does not need addresses. Neighbours must be distinguishable locally, but no node needs a global identifier; data is found by what it is.

Two-way gradients are not a loop. The data cache drops any event a node has already seen, so exploratory events do not circulate.

Reinforcement does not find the shortest path by calculation. It finds the path whose first copy arrived first; on the grid that was also a fewest-hop path, but in a real network it is the lowest-delay one, which may be longer.

Rumour routing does not guarantee delivery. A query can miss every trail; more agents or a longer query lifetime make that rare, and flooding remains as a fallback.

Quick revision

  • Directed diffusion: interests, data, gradients, reinforcement; data named by attribute-value pairs; data-centric; local interactions only.
  • The sink floods an interest (soft state, refreshed); each node caches it and sets a gradient (direction and rate) towards each neighbour it came from: a gradient each way on every link.
  • Sources send exploratory events at a low rate along all gradients; data cache drops duplicates.
  • The sink positively reinforces the neighbour of the first copy (higher rate); repeated back to the source: an empirically low-delay path. Negative reinforcement: time out, or an explicit message. Local repair by intermediate nodes. Aggregation along the way.
  • Program (10 by 10 grid): interest flood 100, 360 gradients, reinforced path 18 hops; one minute of data 5 times fewer transmissions than flooding every event.
  • Rumour routing: events launch agents (live La hops, synchronise events tables, leave thick trails); a query walks straight until it meets a trail (TTL Lq), else is flooded. Two lines cross with about 69 per cent probability (the paper); five paths, about 99.7 per cent.
  • Costs: query flooding N per query; event flooding N per event; rumour routing in between. Program: 5 agents, 99.8 per cent delivered; wins at 10 queries per event (8,171 against 16,000), loses at 100.

Test yourself

1. Explain interests and gradients in directed diffusion. An interest is a named task description, a list of attribute-value pairs such as the type of object, a region and a data rate, which the sink broadcasts periodically and every node caches and rebroadcasts. When a node receives an interest from a neighbour, it creates a gradient towards that neighbour: state giving the direction in which to send matching data and the rate at which to send it. Because interests flood, every pair of neighbours ends up with gradients towards each other.

2. How does directed diffusion choose the path for high-rate data? Sources first send low-rate exploratory events along all gradients, and nodes drop duplicates using their data caches. The sink reinforces the neighbour from which it first received a new event, by resending the interest to it with a higher data rate; that neighbour reinforces its neighbour that first delivered the event, and so on back to the source. The result is a single path of empirically low delay carrying the data at the requested rate; other paths are degraded by timing out or by negative reinforcement.

munotes.in399

Directed Diffusion and Rumour Routing

3. What are the advantages of directed diffusion? It is data-centric, so no global addressing is needed; all communication is between neighbours, so no node needs to know the topology; nodes can cache and aggregate data on the way, saving energy; the reinforced path adapts to delay; and the exploratory gradients allow local repair when a path fails.

4. Explain rumour routing. A node that observes an event may launch agents, long-lived packets that travel roughly straight through the network for a limited number of hops, leaving in each node they pass, and in the neighbours that overhear them, the direction and distance to the event. A query from any node is forwarded along a known route if one exists, or otherwise walks straight in a random direction until it reaches a node on an event path, which it then follows to the event. If the query's lifetime expires first, it can be resent or flooded.

5. When is rumour routing better than query flooding and event flooding? Query flooding costs a network-wide flood per query and event flooding a flood per event. Rumour routing costs a few agent walks per event plus a short walk per query, so it is cheapest when the number of queries per event lies between the two extremes. In the program, with 10 events, it cost 8,171 transmissions for 100 queries, against 16,000 for event flooding, but lost to event flooding at 1,000 queries.

6. Why does rumour routing use straight walks for agents and queries? Two roughly straight lines in a bounded field are likely to cross, which the paper puts at about 69 per cent for two lines, so a query walking straight is likely to meet one of an event's few straight agent trails, and with five trails the chance becomes about 99.7 per cent. Random walks revisit the same area and cover less new ground, so they meet trails less often.

Contents This chapter on its own page

munotes.in400

Chapter Fifty-Nine

LEACH: Clusters That Take Turns

Syllabus topic Module 1, "Routing in WSN: Routing strategies in WSNs"

In one line

LEACH groups the nodes into clusters whose heads collect, fuse and forward their members' data to the base station, and elects new heads at random every round, so that the expensive job of talking to the distant base station is shared by everyone.

In the wording a student can write in an examination: LEACH (Low-Energy Adaptive Clustering Hierarchy, Heinzelman, Chandrakasan and Balakrishnan, 2000) is a hierarchical, cluster-based protocol with three key features: localized coordination for forming and running clusters; randomized rotation of the cluster heads; and local compression (data fusion) to reduce the data sent to the base station. Operation is divided into rounds. Each round begins with a set-up phase: every node n chooses a random number between 0 and 1 and becomes a cluster head if it is below the threshold T(n) = P / (1 - P (r mod 1/P)) for nodes that have not been heads in the current cycle of 1/P rounds (the set G), and 0 otherwise, where P is the desired fraction of heads (say 0.05) and r the round. Heads broadcast an advertisement (using CSMA); every other node joins the head whose advertisement it hears most strongly; each head sends its members a TDMA schedule. In the steady-state phase, members send their data in their own slots and sleep otherwise; the head fuses the data and sends one message to the base station. Neighbouring clusters use different CDMA codes. Rotating the head role spreads the energy load, so nodes die evenly and late; about 5 per cent of nodes as heads is optimal for a 100-node network. LEACH-C lets the base station choose the heads centrally.

Three ideas

The HICSS paper lists LEACH's key features:

  • "Localized coordination and control for cluster set-up and operation."
  • "Randomized rotation of the cluster 'base stations' or 'cluster-heads' and the corresponding clusters."
  • "Local compression to reduce global communication."

Clustering puts most transmissions over short distances, "requiring only a few nodes to transmit far distances to the base station." What LEACH adds to classical clustering is rotation: fixed cluster heads would die first, while "adaptive clusters and rotating cluster-heads" allow "the energy requirements of the system to be distributed among all the sensors."

The radio model it is judged by

LEACH's papers use the first order radio model, which later chapters and most simulations reuse. The radio spends E(elec) = 50 nJ/bit to run the transmitter or receiver electronics, and the transmit amplifier ε(amp) = 100 pJ/bit/m squared. To send a k-bit message over a distance d costs E(Tx)(k, d) = E(elec) × k + ε(amp) × k × d squared, and to receive it E(Rx)(k) = E(elec) × k. The d squared term makes long transmissions expensive: at 100 m, d squared is 10,000 square metres, and at 100 pJ for each the amplifier costs 1,000 nJ per bit, twenty times the electronics' 50. And "For these parameter values, receiving a message is not a low cost operation" either: every relay pays E(elec) per bit to receive as well as to send. ([Single Hop or Multiple Hops: The Energy Argument Worked Out] used the same trade.)

munotes.in401

LEACH: Clusters That Take Turns

A round: the set-up phase

"The operation of LEACH is broken up into rounds, where each round begins with a set-up phase, when the clusters are organized, followed by a steady-state phase, when data transfers to the base station occur." The steady-state phase is made long compared with the set-up, to keep overhead low.

Electing the heads. "each node decides whether or not to become a cluster-head for the current round. This decision is based on the suggested percentage of cluster heads for the network (determined a priori) and the number of times the node has been a cluster-head so far." A node n picks a random number between 0 and 1 and becomes a head if it is below

T(n) = P / (1 - P × (r mod 1/P)) if n is in G, and T(n) = 0 otherwise,

where P is the desired percentage of cluster heads, r the current round, and G the set of nodes that have not been cluster heads in the last 1/P rounds. At the start of each cycle of 1/P rounds every node is eligible and T = P; as the cycle goes on, fewer nodes remain eligible and the threshold rises, so that the expected number of heads stays near P × N; in the last round of the cycle T = 1, and every node that has not yet served must. With this threshold, the paper notes, each node becomes a cluster head at some point within 1/P rounds. The first program walks through one cycle.

# LEACH's cluster-head threshold, round by round: P = 0.05 of the nodes should
# be cluster heads each round, so an epoch lasts 1/P = 20 rounds. A node that
# has not yet been a head in this epoch (the set G) becomes one if a random
# number between 0 and 1 falls below T(n) = P / (1 - P * (r mod 1/P)).
import random

P, N = 0.05, 100
EPOCH = round(1 / P)

def threshold(r):
    return P / (1 - P * (r % EPOCH))

rnd = random.Random(59)
eligible = set(range(N))
print("round  T(n)    eligible  heads elected")
for r in range(EPOCH):
    heads = {n for n in eligible if rnd.random() < threshold(r)}
    print("%5d  %.4f %8d %8d" % (r, threshold(r), len(eligible), len(heads)))
    eligible -= heads
print("after one epoch, %d nodes have never been a head" % len(eligible))
munotes.in402

LEACH: Clusters That Take Turns

round  T(n)    eligible  heads elected
    0  0.0500      100        6
    1  0.0526       94        7
    2  0.0556       87        8
    3  0.0588       79        5
    4  0.0625       74        1
    5  0.0667       73        3
    6  0.0714       70        7
    7  0.0769       63        6
    8  0.0833       57        6
    9  0.0909       51        1
   10  0.1000       50        7
   11  0.1111       43        6
   12  0.1250       37        2
   13  0.1429       35        4
   14  0.1667       31        2
   15  0.2000       29        9
   16  0.2500       20        5
   17  0.3333       15        4
   18  0.5000       11        4
   19  1.0000        7        7
after one epoch, 0 nodes have never been a head

Reading it. The threshold starts at 0.05 and rises as the eligible set shrinks: 0.1 in round 10, 0.5 in round 18 and exactly 1 in round 19. The number of heads varies from round to round (a random election cannot guarantee 5), but the expected number stays near 5, and after 20 rounds every one of the 100 nodes has served exactly once. The worked value for round 10: 0.05 / (1 - 0.05 × 10) = 0.05 / 0.5 = 0.1.

Advertisement. "Each node that has elected itself a cluster-head for the current round broadcasts an advertisement message to the rest of the nodes." Heads use CSMA and the same transmit power; the other nodes keep their receivers on to hear all the advertisements and join the head whose advertisement is strongest, which, with symmetric channels, needs the least energy to reach. A node then tells its head, with a short join message, that it is a member.

The schedule. As [TDMA and Schedule-based MAC] described, "The cluster head node sets up a TDMA schedule and transmits this schedule to the nodes in the cluster", which rules out collisions inside the cluster and lets every member keep its radio off except in its own slot.

A round: the steady-state phase

"The steady-state operation is broken into frames, where nodes send their data to the cluster head at most once per frame during their allocated transmission slot." Members also use power control, setting their transmit power from the strength of the head's advertisement. "The cluster head must be awake to receive all the data from the nodes in the cluster." It then aggregates, "to enhance the common signal and reduce the uncorrelated noise", and sends the result to the base station, a long, expensive transmission made by one node instead of all.

Between clusters. Neighbouring clusters' transmissions would interfere, so each cluster uses its own direct-sequence spreading code, chosen by the head; the heads send to the base station with a fixed code and CSMA. The cost is tight timing synchronisation.

munotes.in403

LEACH: Clusters That Take Turns

Top: a square field of 100 small nodes with six larger ringed cluster heads, from one election, each node joined by a faint line to its nearest head. Bottom: one round as a bar, a short set-up phase (advertise, join, schedule) followed by a long steady state of four frames, each divided into five member slots

Figure 59.1 One round's clusters on the second program's field (heads from one election), and the round's two phases, after LEACH's Fig. 4

Lifetimes: the paper's experiment, repeated

The HICSS paper's Table 2 gives, for 100 nodes each starting with 0.5 J, the round in which the first node dies and the round in which the last does: 109 and 234 with direct transmission to the base station, 932 and 1,312 with LEACH. The second program repeats the experiment on the paper's own set-up: 100 nodes at random in a 50 m by 50 m field, the base station at (0, -100), one 2,000-bit message per node per round, the radio model above, and 5 nJ/bit/message to fuse data at a head. Members join their nearest head; if no node is elected in a round, every node sends straight to the base station.

# Lifetime with direct transmission and with LEACH, on the LEACH paper's own
# set-up (Heinzelman, Chandrakasan and Balakrishnan 2000): 100 nodes at random
# in a 50 m by 50 m field, the base station at (0, -100), 0.5 J per node, one
# 2,000-bit message per node per round, the first order radio model (50 nJ/bit
# for the electronics, 100 pJ/bit/m2 for the amplifier) and 5 nJ/bit/message
# to fuse data at a cluster head. The network is our own random one.
import random

E_ELEC, E_AMP, E_FUSE, K = 50e-9, 100e-12, 5e-9, 2000
BS, START, P = (0.0, -100.0), 0.5, 0.05
EPOCH = round(1 / P)

def tx(d2):                      # J to send K bits over a distance whose square is d2
    return E_ELEC * K + E_AMP * K * d2

def d2(a, b):
    return (a[0] - b[0]) ** 2 + (a[1] - b[1]) ** 2

def lifetime(nodes, leach, rnd):
    energy = [START] * len(nodes)
    alive = set(range(len(nodes)))
    last_head = [-EPOCH] * len(nodes)
    first = r = 0
    while alive:
        heads = set()
        if leach:                                 # G: not yet a head this epoch
            epoch_start = r - r % EPOCH
            for n in alive:
                if last_head[n] < epoch_start and rnd.random() < P / (1 - P * (r % EPOCH)):
                    heads.add(n)
                    last_head[n] = r
        use = {}
        for n in alive:
            if n in heads:
                continue
            if heads:                             # join the nearest cluster head
                h = min(heads, key=lambda h: d2(nodes[n], nodes[h]))
                use[n] = use.get(n, 0.0) + tx(d2(nodes[n], nodes[h]))
                use[h] = use.get(h, 0.0) + E_ELEC * K + E_FUSE * K
            else:                                 # no heads (or direct): straight to BS
                use[n] = use.get(n, 0.0) + tx(d2(nodes[n], BS))
        for h in heads:                           # fuse its own reading, send one message
            use[h] = use.get(h, 0.0) + E_FUSE * K + tx(d2(nodes[h], BS))
        for n, e in use.items():
            energy[n] -= e
        dead = {n for n in alive if energy[n] <= 0}
        if dead and not first:
            first = r + 1
        alive -= dead
        r += 1
    return first, r

rnd = random.Random(59)
nodes = [(rnd.uniform(-25, 25), rnd.uniform(0, 50)) for _ in range(100)]
for name, leach in (("direct transmission", False), ("LEACH, P = 0.05", True)):
    first, last = lifetime(nodes, leach, rnd)
    print("%-20s first node dies in round %4d, last in round %4d" % (name, first, last))
print("(the paper's Table 2, 0.5 J: direct 109 and 234; LEACH 932 and 1312)")
munotes.in404

LEACH: Clusters That Take Turns

direct transmission  first node dies in round  110, last in round  237
LEACH, P = 0.05      first node dies in round 1002, last in round 1232
(the paper's Table 2, 0.5 J: direct 109 and 234; LEACH 932 and 1312)

Reading it: direct transmission. Our network gives 110 and 237 rounds, against the paper's 109 and 234. The arithmetic is simple: the farthest nodes are about 150 m from the base station, where one message costs 2,000 × (50 nJ + 100 pJ × 22,500) = 0.0046 J, so 0.5 J lasts about 108 rounds; the nearest, 100 m away, pay 0.0021 J and last about 238. Nodes die from the far edge inwards, and the region they covered goes unwatched.

Reading it: LEACH. The first node dies in round 1,002 and the last in 1,232, against the paper's 932 and 1,312: roughly nine times the direct lifetime to the first death. Two effects combine. Most nodes now transmit a few metres to a head instead of 100 to 150 m, and fusion means the base station receives about 5 messages a round instead of 100. And because every node takes its turn as head, the energy drains evenly: the first and last deaths are only a few hundred rounds apart, and the nodes that die are scattered, not concentrated at the far edge, so the field stays covered almost to the end.

How many clusters, and LEACH-C

The optimum. Too few clusters, and members must send far to reach a head; too many, and there is little fusion and many long transmissions to the base station. The HICSS paper found "an optimal percent of nodes" as heads, 5 per cent in its network; the journal version derived the optimum analytically and confirmed it by simulating 1 to 11 clusters on a 100-node network: "the optimum number of clusters is around 3" to 5, and "For the rest of the experiments, we set" the number to five.

LEACH-C. The distributed election "offers no guarantee about the placement and/or number of cluster head nodes." In LEACH-C (centralised), each node sends its position (for example from GPS) and energy to the base station, which rules out nodes with less than the average energy and chooses heads by simulated annealing, "by minimizing the total sum of squared distances between all the non-cluster head nodes and the closest cluster head." Its steady state is LEACH's. The journal reports that LEACH-C "delivers about 40% more data per unit energy than" LEACH, at the price of a central controller and position information.

munotes.in405

LEACH: Clusters That Take Turns

LEACH's assumptions and limits

  • Every node can reach the base station in one hop, and heads do: a network larger than a radio's range needs multi-hop routing between heads.
  • Nodes are time-synchronised, for rounds and schedules.
  • Heads are elected by chance, so a round can have too few or too many, badly placed; LEACH-C fixes this centrally.
  • Every node always has data to send in the steady state: event-driven networks waste slots.
  • Data from a cluster is correlated, so fusing it loses little; for independent readings, fusion is less useful.

Distinctions

Direct transmissionLEACH
Who talks to the base stationEvery node, every roundAbout 5 per cent, the heads, taking turns
Distance per message100 to 150 m in the paper's fieldA few metres for members
Messages at the base station per round100About 5, fused
Order of deathsFar nodes firstScattered, and late
Program: first and last deathRounds 110 and 237Rounds 1,002 and 1,232
LEACHLEACH-C
Who chooses headsEach node, with the thresholdThe base station
NeedsNothing but P and the round numberEvery node's position and energy
HeadsRandom number and placementEnergy above average, placed well
Result (journal)The baselineAbout 40 per cent more data per unit energy
Set-up phaseSteady-state phase
HappensOnce per round, brieflyFor most of the round
Consists ofElection, advertisement (CSMA), joining, scheduleFrames of TDMA slots; fusion at the head

What it does not mean

Randomised does not mean unfair. The threshold guarantees that every node serves exactly once per cycle of 1/P rounds; only the order is random.

A cluster head is not a special device. It is an ordinary node taking its turn; LEACH assumes all nodes start with the same energy.

LEACH is not multi-hop. Members reach their head in one hop and heads reach the base station in one hop.

Fewer messages is not the only saving. Shorter distances matter as much: the amplifier's d squared term dominates at 100 m.

Quick revision

  • LEACH (Heinzelman, Chandrakasan, Balakrishnan, 2000): localized coordination, randomized rotation of cluster heads, local compression.
  • Radio model: E(elec) = 50 nJ/bit, ε(amp) = 100 pJ/bit/m squared; E(Tx) = E(elec) k + ε(amp) k d squared; E(Rx) = E(elec) k.
  • Rounds: set-up (election, advertisement by CSMA, join the strongest, TDMA schedule), then a long steady state (frames; members in their slots; the head fuses and sends to the base station). CDMA codes between clusters.
  • T(n) = P / (1 - P (r mod 1/P)) for n in G (not yet a head this cycle), else 0; every node serves once per 1/P rounds; P = 0.05 gives a 20-round cycle, T = 0.1 at round 10, 1 at round 19.
  • Table 2 (0.5 J): direct 109 / 234, LEACH 932 / 1,312; program 110 / 237 and 1,002 / 1,232.
  • Optimum: about 5 per cent heads (3 to 5 clusters for 100 nodes). LEACH-C: base station chooses heads (above-average energy, simulated annealing), about 40 per cent more data per unit energy.
  • Limits: one hop to the base station, synchronisation, random head placement, continuous traffic assumed.
munotes.in406

LEACH: Clusters That Take Turns

Test yourself

1. What are the key features of LEACH? Localized coordination and control for cluster set-up and operation; randomized rotation of the cluster heads and their clusters, so that the high energy cost of being a head is shared; and local compression, data fusion at the cluster heads, to reduce the amount of data sent to the base station.

2. Explain how cluster heads are elected in LEACH. At the start of each round, every node that has not been a cluster head in the current cycle of 1/P rounds chooses a random number between 0 and 1 and becomes a head if it is below T(n) = P / (1 - P (r mod 1/P)), where P is the desired fraction of heads and r the round number; nodes that have already served have T(n) = 0. As the cycle proceeds the threshold rises, so the expected number of heads stays near P times the number of nodes, and in the last round of the cycle it reaches 1, so every node serves once per cycle.

3. With P = 0.05, what is the threshold in round 10 and in round 19 for an eligible node? 1/P = 20. In round 10, T = 0.05 / (1 - 0.05 × 10) = 0.05 / 0.5 = 0.1. In round 19, T = 0.05 / (1 - 0.05 × 19) = 0.05 / 0.05 = 1, so every node that has not yet been a head in this cycle becomes one.

4. Describe the set-up and steady-state phases of a LEACH round. In the set-up phase, nodes elect themselves cluster heads using the threshold; the heads broadcast advertisements using CSMA; every other node joins the head whose advertisement it received most strongly and tells it so; each head sends its members a TDMA schedule. In the steady-state phase, which is much longer, members send their data to the head in their own slots and sleep otherwise; the head receives all the data, fuses it and sends the result to the base station; neighbouring clusters use different CDMA codes to avoid interfering.

munotes.in407

LEACH: Clusters That Take Turns

5. Why does LEACH extend network lifetime compared with direct transmission? Most nodes transmit only a short distance to their head instead of a long distance to the base station, and the amplifier's cost grows with the square of the distance; heads fuse their members' data, so few messages make the long trip; and because the head role rotates, the energy load is spread over all nodes, which therefore die late and at scattered places. In the paper, the first node died in round 932 with LEACH against 109 with direct transmission.

6. What is LEACH-C, and how does it differ from LEACH? LEACH-C is a centralised version: each node reports its position and energy to the base station, which excludes nodes with below-average energy and chooses the cluster heads by simulated annealing to minimise the members' squared distances to their heads. It gives better-placed heads and about 40 per cent more data per unit energy than LEACH, but needs position information and a central controller; its steady state is the same as LEACH's.

Contents This chapter on its own page

munotes.in408

Chapter Sixty

PEGASIS, TEEN and the Other Hierarchical Protocols

Syllabus topic Module 1, "Routing in WSN: Routing strategies in WSNs"

In one line

PEGASIS saves LEACH's cluster overhead by linking all nodes into one chain along which data is fused hop by hop, with one node a round taking it to the base station; TEEN and APTEEN keep clusters but let nodes stay silent unless a reading crosses a hard threshold and changes by a soft one, with APTEEN also reporting at least once every count time.

In the wording a student can write in an examination: PEGASIS (Power-Efficient GAthering in Sensor Information Systems, Lindsey and Raghavendra, 2002) forms a chain through all the nodes, built greedily from the node farthest from the base station, each node adding its nearest unvisited neighbour. In each round every node receives data from one chain neighbour, fuses it with its own and passes it to the other, towards the round's leader, which sends one message to the base station; leadership rotates (node i mod N in round i), and a small token controls the order. It saves the cost of forming clusters and makes each node send only to a close neighbour, giving about twice LEACH's lifetime, but a single leader and long chains add delay. TEEN (Threshold-sensitive Energy Efficient sensor Network protocol, Manjeshwar and Agrawal, 2001) is a cluster-based protocol for reactive networks. At each cluster change, the head broadcasts a hard threshold (HT), the value beyond which a node must report, and a soft threshold (ST), the change that triggers a new report. A node senses continuously but transmits only when the value is at or beyond HT and, after its first report, has changed by at least ST since its last report. It suits time-critical data, but if the thresholds are never reached the user hears nothing. APTEEN adds a count time (CT), the longest a node may go without reporting, combining proactive and reactive reporting.

Proactive and reactive networks

TEEN's paper begins with a classification. Proactive networks "periodically switch on their sensors and transmitters, sense the environment and transmit the data of interest", giving a snapshot at fixed intervals; LEACH is its example. Reactive networks "respond immediately to changes in the relevant parameters of interest", which suits applications where a sudden change matters more than a steady record. PEGASIS improves the first kind; TEEN and APTEEN serve the second.

PEGASIS: one chain instead of clusters

"The key idea in PEGASIS is to form a chain among the sensor nodes so that each node will receive from and transmit to a close neighbor. Gathered data moves from node to node, get fused, and eventually a designated node transmits to the BS. Nodes take turns transmitting to the BS so that the average energy spent by each node per round is reduced."

munotes.in409

PEGASIS, TEEN and the Other Hierarchical Protocols

Building the chain. The shortest chain through all the nodes is a travelling salesman problem, "which is known to be intractable", but "a simple chain built with a greedy approach performs quite well." "To construct the chain, we start with the furthest node from the BS", so that the far nodes get close neighbours, and each step adds the nearest node not yet on the chain. "When a node dies, the chain is reconstructed in the same manner to bypass the dead node."

A round. "each node receives data from one neighbor, fuses with its own data, and transmits to the other neighbor on the chain." Leadership rotates: "we will use node number i mod N (N represents the number of nodes) to transmit to the BS in round i", so the leader sits at a random place on the chain, and nodes die at random places. The leader starts the round with a small token passed to one end of the chain; data flows from that end to the leader; then the token goes to the other end and its data flows back. Every node except the two ends fuses, so "each node will receive and transmit one packet in each round and be the leader once every 100 rounds."

What it saves over LEACH. The paper lists three savings: members send to a close chain neighbour rather than to a cluster head; the leader receives at most two messages instead of a cluster's twenty; and "only one node transmits to the BS in each round". There is also no cluster formation to pay for each round.

A square field of 100 nodes linked by a single winding chain that visits every node, starting from the node farthest from the base station. One node on the chain is ringed as this round's leader, with an arrow down towards the base station below the field

Figure 60.1 The first program's field joined by PEGASIS's greedy chain, with one round's leader

PEGASIS against LEACH, counted

The first program uses the set-up of [LEACH: Clusters That Take Turns]: 100 nodes at random in a 50 m by 50 m field, the base station 100 m below it, 0.5 J per node, 2,000-bit messages, the first order radio model, and 5 nJ/bit to fuse each message received. It runs direct transmission, LEACH with P = 0.05, and PEGASIS with a greedy chain rebuilt whenever a node dies, and records the round by which 1, 20, 50 and 100 per cent of the nodes have died, the measures of the PEGASIS paper.

# PEGASIS against LEACH and direct transmission, on the set-up of the LEACH
# chapter: 100 nodes at random in 50 m by 50 m, the base station 100 m below
# the field, 0.5 J per node, 2,000-bit messages, 50 nJ/bit and 100 pJ/bit/m2,
# and 5 nJ/bit to fuse each message received. PEGASIS builds a greedy chain
# from the node farthest from the base station; each node sends to its chain
# neighbour towards the round's leader, fusing what it received with its own
# reading; the leader (node i mod N in round i) sends one message to the base.
import random

E_ELEC, E_AMP, E_FUSE, K = 50e-9, 100e-12, 5e-9, 2000
BS, START, P = (0.0, -100.0), 0.5, 0.05

def d2(a, b):
    return (a[0] - b[0]) ** 2 + (a[1] - b[1]) ** 2

def tx(a, b):
    return E_ELEC * K + E_AMP * K * d2(a, b)

RX = E_ELEC * K + E_FUSE * K            # receive one message and fuse it

def greedy_chain(nodes, alive):
    left = set(alive)
    chain = [max(left, key=lambda n: d2(nodes[n], BS))]
    left.discard(chain[0])
    while left:
        nxt = min(left, key=lambda n: d2(nodes[n], nodes[chain[-1]]))
        chain.append(nxt)
        left.discard(nxt)
    return chain

def round_pegasis(nodes, alive, r, chain):
    use = dict.fromkeys(alive, 0.0)
    order = sorted(alive)
    leader = chain.index(order[r % len(order)])
    for p, n in enumerate(chain):
        if p < leader:
            use[n] += tx(nodes[n], nodes[chain[p + 1]]) + (RX if p > 0 else 0)
        elif p > leader:
            use[n] += tx(nodes[n], nodes[chain[p - 1]]) + (RX if p < len(chain) - 1 else 0)
        else:
            use[n] += tx(nodes[n], BS) + RX * ((p > 0) + (p < len(chain) - 1))
    return use

def round_leach(nodes, alive, r, last_head, rnd):
    epoch = round(1 / P)
    heads = {n for n in alive if last_head[n] < r - r % epoch
             and rnd.random() < P / (1 - P * (r % epoch))}
    for h in heads:
        last_head[h] = r
    use = dict.fromkeys(alive, 0.0)
    for n in alive:
        if n in heads:
            use[n] += E_FUSE * K + tx(nodes[n], BS)
        elif heads:
            h = min(heads, key=lambda h: d2(nodes[n], nodes[h]))
            use[n] += tx(nodes[n], nodes[h])
            use[h] += RX
        else:
            use[n] += tx(nodes[n], BS)
    return use

def run(nodes, protocol, rnd):
    energy, alive = [START] * len(nodes), set(range(len(nodes)))
    last_head, chain = [-100] * len(nodes), None
    marks, r = {}, 0
    while alive:
        if protocol == "PEGASIS":
            chain = chain or greedy_chain(nodes, alive)
            use = round_pegasis(nodes, alive, r, chain)
        elif protocol == "LEACH":
            use = round_leach(nodes, alive, r, last_head, rnd)
        else:
            use = {n: tx(nodes[n], BS) for n in alive}
        for n, e in use.items():
            energy[n] -= e
        dead = {n for n in alive if energy[n] <= 0}
        if dead:
            alive -= dead
            chain = None                         # rebuild the chain round the dead
        r += 1
        gone = len(nodes) - len(alive)
        for pct in (1, 20, 50, 100):
            if gone >= pct and pct not in marks:
                marks[pct] = r
    return [marks[p] for p in (1, 20, 50, 100)]

rnd = random.Random(60)
nodes = [(rnd.uniform(-25, 25), rnd.uniform(0, 50)) for _ in range(100)]
print("rounds until this share of nodes has died:   1%    20%    50%   100%")
for name in ("direct", "LEACH", "PEGASIS"):
    print("%-40s %6d %6d %6d %6d" % (name, *run(nodes, name, rnd)))
munotes.in410

PEGASIS, TEEN and the Other Hierarchical Protocols

rounds until this share of nodes has died:   1%    20%    50%   100%
direct                                      111    127    150    230
LEACH                                       993   1088   1144   1266
PEGASIS                                    1080   1934   2020   2184
munotes.in411

PEGASIS, TEEN and the Other Hierarchical Protocols

Reading it. From 20 per cent of nodes dead onwards, PEGASIS lasts about 1.7 to 1.8 times as long as LEACH (1,934 rounds against 1,088 at 20 per cent, 2,184 against 1,266 at 100), close to the paper's "approximately 2x the number of rounds compared to LEACH" for a 50 m by 50 m field. Its first death comes only a little later than LEACH's, because the greedy chain leaves a few nodes with a distant neighbour, and such a node pays heavily every round. The paper met this directly: "We improved the performance of PEGASIS by not allowing such nodes to become leaders", by a threshold on the distance to their chain neighbours, a refinement this program leaves out.

PEGASIS's limits

The survey collects them. Nodes must know the positions of all others to build the chain, and must be able to reach the base station directly. "PEGASIS introduces excessive delay for distant node on the chain": data from one end must pass through every node before it. "In addition, the single leader can become a bottleneck." And like LEACH it assumes equal starting energy.

Hierarchical-PEGASIS attacks the delay: with CDMA-capable nodes, the chain is arranged as a tree-like hierarchy, and only spatially separated nodes transmit at the same time, so data moves in parallel; the survey reports that it performs better "than the regular PEGASIS scheme by a factor of about 60".

TEEN: report only what matters

"In this scheme, at every cluster change time, in addition to the attributes, the cluster-head broadcasts to its members":

  • Hard Threshold (HT): "This is a threshold value for the sensed attribute. It is the absolute value of the attribute beyond which, the node sensing this value must switch on its transmitter and report to its cluster head."
  • Soft Threshold (ST): "This is a small change in the value of the sensed attribute which triggers the node to switch on its transmitter and transmit."

"The nodes sense their environment continuously. The first time a parameter from the attribute set reaches its hard threshold value, the node switches on its transmitter and sends the sensed data." The node stores that value as its sensed value (SV), and thereafter transmits in the current cluster period only when both conditions hold: "The current value of the sensed attribute is greater than the hard threshold" and "The current value of the sensed attribute differs from SV by an amount equal to or greater than the soft threshold." Each transmission updates SV.

munotes.in412

PEGASIS, TEEN and the Other Hierarchical Protocols

The two thresholds divide the work: "the hard threshold tries to reduce the number of transmissions by allowing the nodes to transmit only when the sensed attribute is in the range of interest", and "The soft threshold further reduces the number of transmissions" when the value hardly changes. The user can move both at each cluster change, and "A smaller value of the soft threshold gives a more accurate picture of the network, at the expense of increased energy consumption." Sensing continuously costs little; "Message transmission consumes much more energy than data sensing."

TEEN's drawback. "if the thresholds are not reached, the nodes will never communicate": the user cannot tell a quiet network from a dead one.

APTEEN (Adaptive Periodic TEEN) keeps the thresholds and adds, in the survey's words, a count time (CT): "the maximum time period between two successive reports sent by a node". "If a node does not send data for a time period equal to the count time, it is forced to sense and retransmit the data." It "combines both proactive and reactive policies", at the cost of more complexity; in the survey's summary, APTEEN's energy and lifetime lie between LEACH's and TEEN's.

TEEN's thresholds, counted

The second program gives one node a day of temperature readings, one a minute, from an illustrative trace swinging about 85 F with a little noise. The thresholds are TEEN's own experiment's: "The hard threshold is set at the average value of the lowest and the highest possible temperatures, 100 F", and "The soft threshold is set at 2 F". It counts the reports sent by a proactive node, by TEEN with the hard threshold only, by TEEN with both thresholds, and by APTEEN with a count time of 60 minutes, and the worst error the user sees while the temperature is at or above HT.

# TEEN's thresholds on one node's day of temperature readings, one a minute.
# The trace is illustrative: a daily swing about 85 F with a little noise. The
# thresholds are the TEEN paper's own: a hard threshold of 100 F and a soft
# threshold of 2 F. A proactive node reports every minute; TEEN reports the
# first time the reading reaches HT, then only while it stays at or above HT
# and has changed by ST or more since the last report (the stored value SV).
import math
import random

rnd = random.Random(60)
HT, ST, MINUTES = 100.0, 2.0, 24 * 60
noise, trace = 0.0, []
for m in range(MINUTES):
    noise = 0.9 * noise + rnd.gauss(0, 0.4)
    trace.append(85 + 18 * math.sin(2 * math.pi * (m - 6 * 60) / MINUTES) + noise)

def teen(soft, count_time=None):
    """Reports and the worst error the user sees while the reading is at or
    above HT. count_time (APTEEN) forces a report after that many minutes."""
    sent, sv, since, worst = 0, None, 0, 0.0
    for t in trace:
        since += 1
        due = t >= HT and (sv is None or not soft or abs(t - sv) >= ST)
        if count_time and since >= count_time:
            due = True
        if due:
            sent, sv, since = sent + 1, t, 0
        if t >= HT and sv is not None:
            worst = max(worst, abs(t - sv))
    return sent, worst

hot = sum(1 for t in trace if t >= HT)
print("readings: %d; at or above %.0f F: %d minutes, peak %.1f F" % (MINUTES, HT, hot, max(trace)))
print("proactive, every minute       %4d reports" % MINUTES)
for name, soft, ct in (("TEEN, hard threshold only", False, None),
                       ("TEEN, hard and soft", True, None),
                       ("APTEEN, count time 60 min", True, 60)):
    sent, worst = teen(soft, ct)
    print("%-29s %4d reports; worst error above HT %.2f F" % (name, sent, worst))
munotes.in413

PEGASIS, TEEN and the Other Hierarchical Protocols

readings: 1440; at or above 100 F: 278 minutes, peak 104.8 F
proactive, every minute       1440 reports
TEEN, hard threshold only      278 reports; worst error above HT 0.00 F
TEEN, hard and soft              5 reports; worst error above HT 1.99 F
APTEEN, count time 60 min       26 reports; worst error above HT 1.97 F
The day's temperature curve, rising from about 67 F to a peak near 105 F around midday and falling again, with a horizontal line at 100 F. Five dots mark TEEN's reports with both thresholds, all on the stretch above the line

Figure 60.2 The second program's day of readings, the hard threshold, and TEEN's five reports

Reading it. The temperature is at or above 100 F for 278 of the day's 1,440 minutes. A proactive node reports 1,440 times. The hard threshold alone cuts this to 278, one a minute while it is hot, and the user always knows the exact value. Adding the 2 F soft threshold cuts it to 5 reports, and the value the user holds is never more than 1.99 F out: the soft threshold is exactly the accuracy the user chose to give up. For the other 1,162 minutes TEEN says nothing at all, which is its drawback; APTEEN's count time of an hour adds a report at least every 60 minutes, 26 in all, so that silence can be told from failure.

Distinctions

LEACHPEGASIS
StructureClusters, rebuilt every roundOne chain, rebuilt when a node dies
Sends toIts cluster headIts chain neighbour
To the base station per roundAbout 5 heads1 leader
Leader or head chosenBy the threshold, at randomNode i mod N in round i
Lifetime (program, 50 per cent dead)1,144 rounds2,020 rounds
WeaknessCluster overheadDelay along the chain; a bottleneck leader
Proactive (LEACH)Reactive (TEEN)Hybrid (APTEEN)
ReportsPeriodicallyWhen values cross thresholdsBoth, with a count time
SuitsPeriodic monitoringTime-critical eventsBoth kinds of query
In the program1,440 reports526
WeaknessReports whether anything changed or notSilent if thresholds are never reachedComplexity
munotes.in414

PEGASIS, TEEN and the Other Hierarchical Protocols

Hard thresholdSoft threshold
IsAn absolute value of the attributeA change in the value
RuleReport only at or beyond itReport again only after this much change
ControlsWhether the value is of interestHow finely it is tracked

What it does not mean

PEGASIS is not multi-hop to the base station. Data travels along the chain, but the leader still reaches the base station in one hop.

A chain is not a shortest route. The greedy chain is not the shortest possible, and a few nodes end up with distant neighbours.

TEEN does not sense less. Its nodes sense all the time; it is transmissions that the thresholds remove.

A soft threshold is not an error. It is a chosen resolution: the user asks to hear only changes of at least ST.

Quick revision

  • PEGASIS (Lindsey and Raghavendra, 2002): one greedy chain from the node farthest from the base station; each node receives from one neighbour, fuses, passes on; leader node i mod N sends to the base station; token passing; chain rebuilt when a node dies.
  • Saves: short sends, at most two receptions for the leader, one transmission to the base station, no cluster formation. About 2 times LEACH's lifetime (paper); program about 1.8 times from 20 per cent dead. Limits: delay, bottleneck leader, needs positions and one-hop reach. Hierarchical-PEGASIS: parallel transmissions in a tree of chains.
  • TEEN (Manjeshwar and Agrawal, 2001): reactive, cluster-based. Hard threshold (report at or beyond it), soft threshold (report again only after this much change), stored SV. Paper's values: HT 100 F, ST 2 F. Drawback: silent if thresholds are never reached.
  • APTEEN: adds count time (longest gap between reports): proactive and reactive combined.
  • Program (one day, one node): proactive 1,440 reports; hard only 278; hard and soft 5 (error at most 1.99 F); APTEEN 26.

Test yourself

1. How does PEGASIS form its chain and gather data? Using global knowledge of node positions, the chain is built greedily: it starts at the node farthest from the base station and repeatedly adds the nearest node not yet on the chain, so far nodes get close neighbours. In each round, a token from the leader starts data at one end; each node receives its neighbour's data, fuses it with its own, and passes it towards the leader; the same happens from the other end; the leader fuses both and sends one message to the base station. The leader is node i mod N in round i, so the role rotates.

munotes.in415

PEGASIS, TEEN and the Other Hierarchical Protocols

2. Why does PEGASIS outperform LEACH? Each node transmits only to a close chain neighbour instead of a possibly distant cluster head, the leader receives at most two messages instead of a whole cluster's, only one node transmits to the distant base station per round instead of several heads, and there is no cluster formation each round. The paper reports about twice LEACH's lifetime for a 50 m by 50 m network.

3. What are the disadvantages of PEGASIS? Data from the far end of the chain must pass through many nodes, causing long delays; the single leader can become a bottleneck; nodes need the positions of all others to build the chain and must be able to reach the base station directly; and nodes with distant chain neighbours spend much more energy than the rest.

4. Explain the hard and soft thresholds of TEEN. At each cluster change the head broadcasts a hard threshold, an absolute value of the sensed attribute, and a soft threshold, a small change in it. A node senses continuously; the first time the value reaches the hard threshold it transmits and stores the value as SV. After that it transmits only if the value is still beyond the hard threshold and differs from SV by at least the soft threshold, and each transmission updates SV. The hard threshold confines reports to values of interest; the soft threshold suppresses reports when the value barely changes.

5. What is the main drawback of TEEN, and how does APTEEN address it? If the thresholds are never reached, nodes never transmit, so the user receives no data and cannot tell a quiet network from a failed one; TEEN also cannot give periodic reports. APTEEN adds a count time, the longest period a node may go without reporting, after which it is forced to sense and transmit, so the network gives periodic snapshots as well as immediate reports of threshold crossings.

6. In the program, how many reports did each scheme send in a day, and what did the soft threshold cost? A proactive node sent 1,440 reports, TEEN with the hard threshold only 278, TEEN with hard and soft thresholds 5, and APTEEN with a 60-minute count time 26. The soft threshold of 2 F meant that while the temperature was above 100 F the value the user held was at most 1.99 F away from the true one.

Contents This chapter on its own page

munotes.in416

Chapter Sixty-One

Geographic Routing: Greedy Forwarding and GPSR

Syllabus topic Module 1, "Routing in WSN: Routing strategies in WSNs"

In one line

Geographic routing sends each packet to the neighbour closest to the destination's position, which needs no routes at all; where that fails, at the edge of a hole in the network, GPSR walks round the hole along the faces of a planar version of the graph until it can go greedy again.

In the wording a student can write in an examination: geographic (location-based) routing assumes each node knows its own position and its neighbours' (from beacons), and each packet carries the destination's position. In greedy forwarding, a node forwards the packet to the neighbour geographically closest to the destination, if that neighbour is closer than itself; no routing tables or route discovery are needed, only one-hop state, so it scales well. Greedy forwarding fails at a local maximum: a node closer to the destination than all its neighbours, at the edge of a void (a region without nodes). GPSR (Greedy Perimeter Stateless Routing, Karp and Kung, 2000) then switches to perimeter mode: using a planar subgraph of the radio graph (the Gabriel graph or the relative neighbourhood graph, in which no edges cross), it forwards by the right-hand rule (the next edge is the first counterclockwise from the edge the packet arrived on) around the faces crossed by the line from the point of failure to the destination, changing face where that line is crossed, and returns to greedy mode as soon as it reaches a node closer to the destination than the point where greedy failed. The packet carries Lp (where it entered perimeter mode), Lf (where it entered the current face) and e0 (the first edge on that face, to detect an unreachable destination).

GEAR (Geographical and Energy Aware Routing) chooses neighbours by a weighted sum of distance and energy used, and disseminates inside the target region by recursive geographic forwarding.

Greedy forwarding

Geographic routing replaces addresses with positions. Karp and Kung assume each node learns its neighbours' positions from periodic beacons, and each packet carries its destination's location. The forwarding rule then needs nothing else: "Upon receiving a greedy-mode packet for forwarding, a node searches its neighbor table for the neighbor geographically closest to the packet's destination. If this neighbor is closer to the destination, the node forwards the packet to that neighbor."

Its attraction is scale: a node's state is its neighbours' positions, whatever the size of the network, and there is no route discovery to flood and no routing table to keep current.

The void. "there are topologies in which the only route to a destination requires a packet move temporarily farther in geometric distance from the destination." At such a node x, every neighbour is farther from the destination D than x itself: "x is a local maximum in its proximity to D. Some other mechanism must be used to forward packets in these situations." The empty region between x and D is a void. In a sensor network voids are ordinary: a lake, a building, a patch of dead nodes.

munotes.in417

Geographic Routing: Greedy Forwarding and GPSR

The right-hand rule and planar graphs

To get round a void, GPSR uses an old maze-walking rule. "The long-known right-hand rule for traversing a graph" states that when a packet arrives at node x from node y, "the next edge traversed is the next one sequentially counterclockwise about x from edge (x, y)." Applied repeatedly, it walks around the edges of a face, one of the regions into which a drawing of the graph divides the plane, and on the face bordering a void it leads round the void.

The rule works only on a planar graph, one whose edges do not cross; on a radio graph with crossing links it can loop. So GPSR prunes the radio graph to a planar subgraph, using only each node's neighbour list:

  • Relative neighbourhood graph (RNG). Edge (u, v) is kept if no other node w is closer to both u and v than they are to each other: d(u, v) is at most max[d(u, w), d(v, w)] for every w. The region that must be empty of a witness w is the lune between the two circles of radius d(u, v).
  • Gabriel graph (GG). Edge (u, v) is kept if no other node w lies inside the circle whose diameter is uv: d squared (u, v) < d squared (u, w) + d squared (v, w) for every w.

Neither can disconnect a connected network: an edge is removed only when a witness w within range of both ends offers another way round. Karp and Kung note that the RNG is a subset of the GG, so the GG keeps more links. The program uses the GG.

GPSR: greedy, then perimeter, then greedy

"All data packets are marked initially at their originators as greedy-mode." When greedy forwarding fails at x, "the node marks the packet into perimeter mode", and records in the packet the fields of the paper's Table 1:

  • Lp, "Location Packet Entered Perimeter Mode", the point where greedy failed;
  • Lf, "Point on xV Packet Entered Current Face" (the text's xV is the line from x to the destination);
  • e0, "First Edge Traversed on Current Face".

Entering perimeter mode. "x forwards the packet to the first edge counterclockwise about x from the line xD." Thereafter each node forwards by the right-hand rule around that face.

Changing face. At each hop, a node checks whether the edge to its chosen next hop crosses the line from Lp to D at a point y closer to D than Lf. If so, the packet has reached the far side of this face: Lf becomes y, and "The node forwards the packet along the first edge of this next face", "the next edge counterclockwise about itself" from the one that crossed, recording it as the new e0. "This process repeats at successively closer faces to D."

munotes.in418

Geographic Routing: Greedy Forwarding and GPSR

Returning to greedy. A packet goes back to greedy mode at the first node whose distance to D is less than the distance from Lp to D. Perimeter forwarding is only a detour round the obstacle.

An unreachable destination. If D is not connected to the network, the packet ends on a face that does not contain it and tours the whole face; when it is about to traverse e0 a second time, the destination is declared unreachable and the packet dropped.

A square field with an empty circular lake in the middle. A route leaves the source on one side, runs greedily towards the lake until a ringed node on the shore where greedy forwarding stops, then follows the lake's shore round to the far side and on to the destination; a thin dashed line joins the ringed node to the destination across the lake

Figure 61.1 One of the program's routes: greedy until the shore, then perimeter mode round the lake

Greedy and GPSR, counted

The program places 300 nodes at random over 200 m by 200 m, none inside a lake of radius 50 m at the centre, and implements both methods as the paper describes them: greedy forwarding on the full radio graph, and GPSR with the Gabriel graph, the right-hand rule, face changes on the line from Lp to D, the return to greedy mode, and the e0 test. For 400 random pairs of connected nodes, at radio ranges of 25 m and 20 m, it records how often greedy forwarding alone delivers, how often GPSR delivers, how many deliveries needed perimeter mode, and how long GPSR's paths are compared with the shortest path in hops.

# Greedy forwarding and GPSR (Karp and Kung 2000) on a field with a void: 300
# nodes at random over 200 m by 200 m, none inside a lake of radius 50 m at the
# centre, radio range 25 m. Greedy uses the full graph; perimeter mode uses its
# Gabriel graph and the right-hand rule, with the packet fields Lp, Lf and e0.
import math
import random
from collections import deque

rnd = random.Random(61)
N, SIDE, LAKE = 300, 200.0, 50.0
pts = []
while len(pts) < N:
    p = (rnd.uniform(0, SIDE), rnd.uniform(0, SIDE))
    if math.dist(p, (SIDE / 2, SIDE / 2)) > LAKE:
        pts.append(p)

def radio_graph(reach):
    return [[j for j in range(N) if j != i and math.dist(pts[i], pts[j]) <= reach]
            for i in range(N)]

def gabriel(u):
    """Keep edge (u, v) only if no neighbour w lies inside the circle on uv."""
    keep = []
    for v in nbr[u]:
        m = ((pts[u][0] + pts[v][0]) / 2, (pts[u][1] + pts[v][1]) / 2)
        if all(math.dist(m, pts[w]) >= math.dist(pts[u], m) for w in nbr[u] if w != v):
            keep.append(v)
    return keep

def angle(a, b):
    return math.atan2(pts[b][1] - pts[a][1], pts[b][0] - pts[a][0])

def ccw_from(u, ref):
    """The planar neighbour first counterclockwise about u from direction ref."""
    return min(planar[u], key=lambda v: (angle(u, v) - ref) % (2 * math.pi) or 2 * math.pi)

def crossing(a, b, c, d):
    """Where segment ab crosses segment cd, or None."""
    (x1, y1), (x2, y2), (x3, y3), (x4, y4) = a, b, c, d
    den = (x1 - x2) * (y3 - y4) - (y1 - y2) * (x3 - x4)
    if abs(den) < 1e-12:
        return None
    t = ((x1 - x3) * (y3 - y4) - (y1 - y3) * (x3 - x4)) / den
    s = ((x1 - x3) * (y1 - y2) - (y1 - y3) * (x1 - x2)) / den
    return (x1 + t * (x2 - x1), y1 + t * (y2 - y1)) if 0 < t < 1 and 0 < s < 1 else None

def route(src, dst, gpsr=True, limit=4 * N):
    D, u, prev, path = pts[dst], src, None, [src]
    mode, Lp, Lf, e0, perimeter_hops = "greedy", None, None, None, 0
    while u != dst and len(path) < limit:
        if mode == "perimeter" and math.dist(pts[u], D) < math.dist(Lp, D):
            mode = "greedy"                       # closer than where greedy failed
        if mode == "greedy":
            best = min(nbr[u], key=lambda v: math.dist(pts[v], D))
            if math.dist(pts[best], D) < math.dist(pts[u], D):
                prev, u = u, best
                path.append(u)
                continue
            if not gpsr or not planar[u]:
                return None, path, perimeter_hops     # a local maximum: greedy is stuck
            mode, Lp, Lf = "perimeter", pts[u], pts[u]
            nxt = ccw_from(u, math.atan2(D[1] - pts[u][1], D[0] - pts[u][0]))
            e0 = (u, nxt)
        else:
            nxt = ccw_from(u, angle(u, prev))         # the right-hand rule
            changed = False
            for _ in range(len(planar[u])):           # change face where Lp-D is crossed
                y = crossing(pts[u], pts[nxt], Lp, D)
                if y is None or math.dist(y, D) >= math.dist(Lf, D):
                    break
                Lf, changed = y, True
                nxt = ccw_from(u, angle(u, nxt))      # first edge of the next face
            if changed:
                e0 = (u, nxt)
            elif (u, nxt) == e0:
                return None, path, perimeter_hops     # toured the whole face: unreachable
        prev, u = u, nxt
        path.append(u)
        perimeter_hops += 1
    return (u == dst), path, perimeter_hops

def hops(src, dst):
    seen, q = {src: 0}, deque([src])
    while q:
        a = q.popleft()
        for b in nbr[a]:
            if b not in seen:
                seen[b] = seen[a] + 1
                q.append(b)
    return seen.get(dst)

print("range  pairs  greedy alone  GPSR    needed perimeter  path / shortest")
for RANGE in (25.0, 20.0):
    nbr = radio_graph(RANGE)
    planar = [gabriel(u) for u in range(N)]
    tried = greedy_ok = gpsr_ok = needed = 0
    stretch = []
    for _ in range(400):
        s, d = rnd.sample(range(N), 2)
        best = hops(s, d)
        if best is None:
            continue                              # not connected at all: skip
        tried += 1
        g, _, _ = route(s, d, gpsr=False)
        ok, path, per = route(s, d)
        greedy_ok += bool(g)
        if ok:
            gpsr_ok += 1
            needed += per > 0
            stretch.append((len(path) - 1) / best)
    print("%3.0f m %6d %11.1f%% %6.1f%% %10d %13.2f (worst %.2f)"
          % (RANGE, tried, 100 * greedy_ok / tried, 100 * gpsr_ok / tried, needed,
             sum(stretch) / len(stretch), max(stretch)))
munotes.in419

Geographic Routing: Greedy Forwarding and GPSR

range  pairs  greedy alone  GPSR    needed perimeter  path / shortest
 25 m    400        94.5%  100.0%         22          1.04 (worst 2.11)
 20 m    400        76.8%  100.0%         93          1.69 (worst 29.75)
munotes.in420

Geographic Routing: Greedy Forwarding and GPSR

Reading it: greedy alone. With a 25 m range, greedy forwarding delivers 94.5 per cent of packets; the rest stop at a local maximum on the lake's shore. With 20 m, the network is sparser, small voids appear everywhere, and greedy delivers only 76.8 per cent.

Reading it: GPSR. GPSR delivers every one of the 400 connected pairs at both densities, as the paper promises for a connected planar graph: perimeter mode always finds a way round. In the dense field it needed perimeter mode for 22 packets and its paths averaged only 1.04 times the shortest; in the sparse field it needed it for 93, and paths averaged 1.69 times the shortest, one of them almost 30 times. Perimeter mode guarantees delivery, not a short path: a packet that must follow the outside of the network, the exterior face, can wander a long way.

GEAR: energy and regions

Many sensor queries name a region ("what is the average temperature in a region R"), not a node. GEAR, from UCLA and USC, adds two things to geographic forwarding.

Energy-aware neighbour selection. Each node keeps a learned cost h(N, R) of reaching region R, and, where it has none for a neighbour, uses an estimated cost (the paper's equation 1):

c(N, R) = α d(N, R) + (1 - α) e(N),

where α is a tunable weight, d(N, R) the neighbour's distance to the region's centroid normalised by the largest such distance among the node's neighbours, and e(N) its consumed energy, normalised likewise. With equal energy everywhere, "this degenerates to the classical greedy geographic forwarding"; with equal distances, it spreads load towards the neighbours that have used the least energy. After choosing a next hop, a node updates its own learned cost to the chosen neighbour's plus the cost of the hop, and these learned costs, propagated back, teach nodes to route round holes.

Inside the region. Once a packet reaches the target region, GEAR uses recursive geographic forwarding, splitting the region into sub-regions and sending a copy towards the centre of each, recursively, "or Restricted Flooding algorithm to disseminate the packet inside the destination region" where the region is too sparse for recursion to terminate. Its authors found that, "especially for non-uniform traffic distribution, GEAR exhibits noticeably longer network lifetime than non-energy-aware geographic routing algorithms."

munotes.in421

Geographic Routing: Greedy Forwarding and GPSR

Distinctions

Greedy forwardingPerimeter mode (GPSR)
GraphFull radio graphPlanar subgraph (GG or RNG)
Next hopNeighbour closest to D, if closer than selfFirst edge counterclockwise (right-hand rule)
ProgressEvery hop gets closerMay move away from D for a while
Ends whenD is reached, or a local maximumA node closer to D than Lp (back to greedy)
RNGGabriel graph
Edge (u, v) kept ifNo witness in the luneNo witness in the circle on uv
KeepsFewer edges (a subset of the GG)More edges
BothPlanar, and do not disconnect a connected graph
GPSRGEAR
TargetA point (a node's position)A region
Chooses byDistance onlyDistance and energy used (weight α)
HolesPerimeter modeLearned costs
At the targetDeliveryRecursive geographic forwarding or restricted flooding

What it does not mean

Geographic routing does not need GPS on every node. It needs positions, which may come from localisation with a few anchors ([Time Synchronisation and Localisation]).

Stateless does not mean no state at all. Each node keeps its neighbours' positions; the paper's footnote says the word refers to "this small, purely local state".

Perimeter mode is not shortest-path routing. It finds a way round, not the shortest way; in the sparse field one path was almost 30 times the shortest.

Planarising does not change the radio. Nodes still hear all their neighbours and greedy mode uses them all; the planar subgraph is used only in perimeter mode.

Quick revision

  • Greedy forwarding: send to the neighbour closest to D, if closer than self; only neighbours' positions needed; scales well.
  • Fails at a local maximum at the edge of a void.
  • Right-hand rule: next edge = first counterclockwise about x from the edge the packet came in on; walks round a face; needs a planar graph.
  • RNG: no witness in the lune; GG: no witness in the circle on uv; RNG is a subset of the GG; neither disconnects the graph.
  • GPSR: greedy on the full graph; at a local maximum, perimeter mode on the planar graph: first edge counterclockwise from the line xD, then the right-hand rule; change face where the edge crosses Lp-D closer than Lf; back to greedy when closer to D than Lp; e0 again: unreachable.
  • Program (300 nodes, lake of 50 m): 25 m range, greedy 94.5 per cent, GPSR 100 per cent, paths 1.04 times the shortest; 20 m, greedy 76.8, GPSR 100, paths 1.69 times (worst about 30).
  • GEAR: estimated cost c = α d + (1 - α) e, learned costs round holes; recursive geographic forwarding or restricted flooding inside the region.
munotes.in422

Geographic Routing: Greedy Forwarding and GPSR

Test yourself

1. Explain greedy geographic forwarding and when it fails. Each node knows its neighbours' positions, and each packet carries the destination's position. A node forwards the packet to the neighbour geographically closest to the destination, provided that neighbour is closer to it than the node itself. It fails at a local maximum: a node that is closer to the destination than all of its neighbours, typically at the edge of a void, a region with no nodes, even though a path around the void may exist.

2. What is the right-hand rule, and why must the graph be planar? When a packet arrives at node x from node y, it is sent on the next edge counterclockwise about x from edge (x, y). Repeated, this traverses the boundary of a face of the graph, which lets the packet walk round a void. The rule traverses faces correctly only if no edges cross; on a graph with crossing edges it can loop, so GPSR first reduces the radio graph to a planar subgraph.

3. How are the RNG and the Gabriel graph built, and why do they not disconnect the network? In the relative neighbourhood graph, an edge (u, v) is kept only if no other node w is closer to both u and v than they are to each other, that is, no witness lies in the lune between them. In the Gabriel graph, an edge is kept only if no other node lies inside the circle having uv as its diameter. An edge is removed only when such a witness exists within range of both u and v, which provides an alternative path, so a connected network stays connected.

4. Describe GPSR's perimeter mode, including the fields Lp, Lf and e0. When greedy forwarding fails at x, the packet enters perimeter mode and records Lp, the location where it did so. x sends it on the first edge counterclockwise from the line x to D, and subsequent nodes forward by the right-hand rule on the planar graph. When an edge about to be used crosses the line from Lp to D at a point closer to D than Lf, the point where the packet entered its current face, Lf is updated and the packet moves onto the next face, recording that face's first edge as e0. As soon as a node closer to D than Lp is reached, the packet returns to greedy mode. If the packet is about to traverse e0 again, it has toured a whole face without progress, and D is unreachable.

munotes.in423

Geographic Routing: Greedy Forwarding and GPSR

5. In the program, what did GPSR gain over greedy forwarding, and at what cost? Greedy forwarding alone delivered 94.5 per cent of packets with a 25 m range and 76.8 per cent with 20 m; GPSR delivered all of them at both densities. The cost was longer paths: 1.04 times the shortest on average in the dense field, 1.69 times in the sparse one, and once almost 30 times.

6. How does GEAR differ from GPSR? GEAR routes towards a region rather than a point, and chooses its next hop by an estimated cost that weighs the neighbour's distance to the region's centroid against the energy it has already consumed, with a tunable weight α, so as to spread the load; learned costs updated from each choice let it route round holes. Inside the target region it disseminates the packet by recursive geographic forwarding, or by restricted flooding where the region is sparse.

Contents This chapter on its own page

munotes.in424

Chapter Sixty-Two

Routing Tables and What Happens When the Topology Changes

Syllabus topic Module 1, "Routing in WSN: Routing strategies in WSNs" (and the paired practical, "Implement a simple routing mechanism and analyze routing table updates during topology changes")

In one line

A sensor node's routing table lists its neighbours, the cost of the link to each (the expected transmissions, ETX) and the cost each advertises to the root; the node's parent is the neighbour with the lowest total, and when a node dies its children simply choose again from their tables, as the beacons that refresh them arrive.

In the wording a student can write in an examination: in a collection tree, used by TinyOS's CTP (Collection Tree Protocol), every node forwards data towards one or more roots (sinks) through a parent chosen by a routing gradient. CTP's gradient is ETX, the expected number of transmissions: a root has ETX 0, and a node's ETX is its parent's plus the ETX of its link to the parent. A link's ETX is estimated as 1 / (forward quality × backward quality), since both the packet and its acknowledgement must get through. Each node keeps a routing table with, for each neighbour, the link ETX, the neighbour's advertised path ETX and their sum, and chooses as parent the neighbour with the lowest sum, changing parent only if another is better by a margin, for stability. Nodes advertise their path ETX in routing beacons, sent on a Trickle-like timer that grows when the network is stable and resets when something changes. When a node dies, its children notice missing acknowledgements or beacons, pick the next best neighbour, and their changed ETX propagates. Routing loops can form while tables are inconsistent; CTP detects them when a node receives data from a node advertising a lower ETX than its own, and caps the ETX it will accept, since a partitioned loop's ETX grows without end. ETX routes use more hops than hop-count routes but fewer transmissions.

What a sensor node's routing table holds

An Internet router's table lists destinations. A sensor node in a collection network needs only one destination, the root, and its table lists neighbours, with what it knows about each:

FieldMeaning
NeighbourA node within radio range
Link ETXExpected transmissions to deliver one packet to it, acknowledgement included
Advertised path ETXThe neighbour's own cost to the root, from its latest beacon
TotalLink ETX plus advertised path ETX: the cost of routing through it

The node's parent is the neighbour with the lowest total, and the node's own path ETX, which it advertises in turn, is that total. "The minimum cost route has the smallest sum the path ETX from that node and the link ETX of that node. The path ETX is therefore the sum of link ETX values along the entire route."

CTP is address-free: "a node does not send a packet to a particular root; instead, it implicitly chooses a root by choosing a next hop." With several roots, a node simply joins the tree whose root is cheapest to reach.

munotes.in425

Routing Tables and What Happens When the Topology Changes

ETX: why transmissions, not hops

"CTP uses expected transmissions (ETX) as its routing gradient. A root has an ETX of 0. The ETX of a node is the ETX of its parent plus the ETX of its link to its parent." The metric makes sense because nodes retransmit lost packets at the link layer, so the real cost of a hop is how many times it must be sent.

Woo, Tong and Culler explain why hop count misleads on lossy links: "If link quality is not considered in route selection, the real cost of packet delivery can be much larger than the hop count." Shortest-path routing picks long hops, and long hops are the lossy ones ([TOSSIM: Simulating Motes, Radio Gain and Packet Loss] showed the transitional region where they live). "With links of varying quality, a longer path with fewer retransmissions may be better than a shorter path with many retransmissions." They call the transmission count the Minimum Transmission (MT) metric and note that "it is important to determine link quality for both directions since losing an acknowledgment would also trigger a useless retransmission": a link's cost is 1 over the product of its forward and backward qualities.

Worked example. A 10 m link that delivers 95 per cent of packets one way and 90 per cent of acknowledgements the other has ETX 1 / (0.95 × 0.90) = 1 / 0.855, about 1.17. A 20 m link that delivers 40 and 50 per cent has ETX 1 / (0.40 × 0.50) = 1 / 0.2 = 5. Two short hops cost about 2.34 transmissions; one long hop costs 5.

Beacons and link estimates

A node learns its table from routing beacons, in which each neighbour advertises its current path ETX and parent. "When a node hears a routing frame, it MUST update its routing table to reflect the address' new metric." The TinyOS implementation sends beacons "on an exponentially increasing randomized timer", like the Trickle algorithm, so that a stable network beacons rarely, and resets the timer to a short interval when the routing table is empty, when "The node's routing ETX increases by >= 1 transmission", or when it hears a packet asking for routing information.

Link ETX is estimated in two ways and combined: from beacons, which seed the table, and from data traffic, "a direct measure of ETX": the estimator produces a new ETX "every 5 such transmissions, where 0 successes has an ETX of 6." Because data estimates arrive as fast as data is sent, the node "can quickly detect a broken link and switch to another candidate neighbor."

munotes.in426

Routing Tables and What Happens When the Topology Changes

A tree, a table, and a death, computed

The program places 40 nodes at random over 60 m by 60 m, with the root at a corner. Link quality is illustrative: 1 up to 8 m, falling to 0 at 22 m, and a little different in each direction. A link's ETX is 1 / (forward × backward quality); links with ETX above 10 are ignored, and path ETX is capped at 30, as CTP caps routes with "an ETX higher than a reasonable constant". In each beacon round every node advertises its path ETX and then re-chooses its parent from its table, keeping its current parent unless another neighbour is better by a margin of 0.5 ETX, Woo, Tong and Culler's "noise margin". The program builds the tree, prints the deepest node's table, compares the tree's routes with hop-count routes over the same links, and then kills the relay on the most nodes' routes and lets the tree repair itself, watching for loops.

# A collection tree built the way TinyOS's CTP builds one. 40 nodes at random
# over 60 m by 60 m, the root at a corner. A link's quality falls from 1 at
# 8 m to 0 at 22 m, a little differently each way (illustrative). A link's
# ETX is 1 / (forward quality x backward quality), links with ETX over 10 are
# ignored, and a node's path ETX is its parent's plus the link's. Every round
# each node beacons its path ETX, and each node picks the neighbour with the
# lowest total, switching parent only if the new one is better by MARGIN.
import math
import random

rnd = random.Random(62)
N, SIDE, MARGIN, CAP = 40, 60.0, 0.5, 30.0
pts = [(0.0, 0.0)] + [(rnd.uniform(0, SIDE), rnd.uniform(0, SIDE)) for _ in range(N - 1)]

def quality(d):
    return 1.0 if d <= 8 else max(0.0, 1 - (d - 8) / 14)

etx = {}
for i in range(N):
    for j in range(i + 1, N):
        d = math.dist(pts[i], pts[j])
        qf = min(1.0, quality(d) * rnd.uniform(0.85, 1.0))
        qb = min(1.0, quality(d) * rnd.uniform(0.85, 1.0))
        if qf * qb > 0.1:                              # ETX below 10
            etx[i, j] = etx[j, i] = 1 / (qf * qb)
nbrs = {i: [j for j in range(N) if (i, j) in etx] for i in range(N)}

def converge(alive, path, parent):
    """Beacon rounds until nothing changes. Returns the rounds taken and the
    number of rounds in which some nodes' parents formed a loop."""
    rounds = looped = 0
    while True:
        rounds += 1
        adv = dict(path)                               # this round's beacons
        changed = False
        for u in alive:
            if u == 0:
                continue
            options = [(etx[u, v] + adv[v], v) for v in nbrs[u] if v in alive and adv[v] < CAP]
            if not options:
                new, p = CAP, None
            else:
                new, p = min(options)
                cur = parent.get(u)
                if cur in alive and cur is not None and adv[cur] < CAP:
                    keep = etx[u, cur] + adv[cur]
                    if keep <= new + MARGIN:           # not better by the margin: stay
                        new, p = keep, cur
            if p != parent.get(u) or abs(new - path[u]) > 1e-9:
                changed = True
            path[u], parent[u] = new, p
        if any(in_loop(u, parent) for u in alive):
            looped += 1
        if not changed or rounds > 200:
            return rounds, looped

def in_loop(u, parent):
    seen = set()
    while u is not None and u != 0:
        if u in seen:
            return True
        seen.add(u)
        u = parent.get(u)
    return False

alive = set(range(N))
path = {u: (0.0 if u == 0 else CAP) for u in range(N)}
parent = {}
rounds, _ = converge(alive, path, parent)
print("tree built in %d beacon rounds; %d links usable" % (rounds, len(etx) // 2))

def on_route(w, u):
    """Is u on w's route to the root?"""
    while w not in (None, 0):
        if w == u:
            return True
        w = parent.get(w)
    return False

def hops_to_root(u):
    h = 0
    while u != 0:
        u, h = parent[u], h + 1
    return h

reached = [u for u in alive if u and path[u] < CAP]
show = max(reached, key=hops_to_root)
print("%d of %d nodes have a route; the deepest, node %d, has path ETX %.2f in %d hops"
      % (len(reached), N - 1, show, path[show], hops_to_root(show)))
print("its routing table:")
print("  neighbour  link ETX  advertised  total")
for v in sorted(nbrs[show], key=lambda v: etx[show, v] + path[v]):
    print("  %9d %9.2f %11.2f %6.2f%s" % (v, etx[show, v], path[v], etx[show, v] + path[v],
                                          "   <- parent" if v == parent[show] else ""))

# the same links routed by hop count instead: breadth-first from the root
from collections import deque
hop, via, q = {0: 0}, {}, deque([0])
while q:
    a = q.popleft()
    for b in sorted(nbrs[a]):
        if b not in hop:
            hop[b], via[b] = hop[a] + 1, a
            q.append(b)
def hop_route_etx(u):
    total = 0.0
    while u != 0:
        total, u = total + etx[u, via[u]], via[u]
    return total
print("over the %d routed nodes: hop-count routes average %.2f hops and %.2f ETX;"
      % (len(reached), sum(hop[u] for u in reached) / len(reached),
         sum(hop_route_etx(u) for u in reached) / len(reached)))
print("the ETX tree's routes average %.2f hops and %.2f ETX"
      % (sum(hops_to_root(u) for u in reached) / len(reached),
         sum(path[u] for u in reached) / len(reached)))

relay = max(reached, key=lambda u: sum(on_route(w, u) for w in reached if w != u))
affected = [w for w in reached if w != relay and on_route(w, relay)]
before = {w: path[w] for w in affected}
alive.discard(relay)
rounds, looped = converge(alive, path, parent)
print("node %d dies: it was on the route of %d nodes" % (relay, len(affected)))
print("re-converged in %d beacon rounds; a routing loop existed in %d of them" % (rounds, looped))
cut = [w for w in affected if path[w] >= CAP]
worse = [path[w] - before[w] for w in affected if path[w] < CAP]
print("afterwards %d are cut off; the others' path ETX rose by %.2f on average, %.2f at most"
      % (len(cut), sum(worse) / len(worse), max(worse)))
munotes.in427

Routing Tables and What Happens When the Topology Changes

tree built in 9 beacon rounds; 148 links usable
36 of 39 nodes have a route; the deepest, node 6, has path ETX 16.71 in 8 hops
its routing table:
  neighbour  link ETX  advertised  total
         15      1.14       15.39  16.53
         13      1.23       15.48  16.71   <- parent
          2      1.17       15.97  17.15
          1      1.21       16.98  18.19
         23      4.21       14.30  18.51
over the 36 routed nodes: hop-count routes average 3.56 hops and 18.06 ETX;
the ETX tree's routes average 5.08 hops and 11.46 ETX
node 29 dies: it was on the route of 34 nodes
re-converged in 11 beacon rounds; a routing loop existed in 0 of them
afterwards 0 are cut off; the others' path ETX rose by 0.90 on average, 2.88 at most
munotes.in428

Routing Tables and What Happens When the Topology Changes

Two copies of the field side by side, each with the root as a square at the bottom left and every node joined to its parent. Left: the tree as first built, with one node near the root ringed as the relay that will die. Right: the tree after that node's death, the ringed position empty and the nodes that routed through it reattached through other neighbours

Figure 62.1 The program's collection tree before and after its busiest relay dies

Reading it: building the tree. Starting with only the root's ETX known, the tree forms in 9 beacon rounds, each round's beacons carrying the gradient one hop further. 36 of the 39 nodes have a route; the other 3 have no usable link within the ETX limits.

Reading it: one table. The deepest node, 6, reaches the root in 8 hops at a path ETX of 16.71. Its table lists five neighbours. Neighbour 15 now offers the lowest total, 16.53, but node 6 keeps its parent 13, at 16.71: the difference, 0.18, is smaller than the 0.5 margin. Without the margin, small fluctuations in link estimates would make nodes change parent constantly, and each change would ripple through the tree below them. Neighbour 23 shows why ETX matters: it advertises the best path of all (14.30) but its link to node 6 costs 4.21 transmissions.

Reading it: ETX against hop count. Over the same usable links, hop-count routes average 3.56 hops but 18.06 expected transmissions; the ETX tree's routes average 5.08 hops and 11.46 transmissions, about 37 per cent fewer. Hop count chooses long, lossy hops; ETX chooses more, shorter, reliable ones.

Reading it: a death. The busiest relay, node 29, next to the root, was on the route of 34 nodes. After it dies, the tree re-forms in 11 beacon rounds, no node is cut off, and path ETX rises by 0.90 on average and 2.88 at most. No loop formed in this run; the next section is about the runs where one does.

munotes.in429

Routing Tables and What Happens When the Topology Changes

Loops, and CTP's two defences

"Routing loops generally occur when a node choose a new route that has a significantly higher ETX than its old one, perhaps in response to losing connectivity with a candidate parent. If the new route includes a node which was a descendant, then a loop occurs." A child whose parent has died may see its own child still advertising the old, cheap path through the dead node, choose it, and so point at a node that points back at it. Each then advertises a slightly higher ETX than the other, and the numbers climb, the count to infinity of distance-vector routing, which the paired practical demonstrates.

CTP has two defences:

  1. Every data frame carries the sender's ETX. A node that receives data from a node whose ETX is lower than its own knows the tree is inconsistent (data should always flow downhill in ETX), and "CTP tries to resolve the inconsistency by broadcasting a beacon frame, with the hope that the node which sent the data frame will hear it and adjust its routes accordingly." The data traffic itself detects loops, quickly, wherever there is traffic.
  2. A cap on ETX. "If a collection of nodes is separated from the rest of the network, then they will form a loop whose ETX increases forever. CTP's second mechanism is to not consider routes with an ETX higher than a reasonable constant." The program's cap of 30 plays this part.

Loops also complicate duplicate suppression: a looping packet revisits nodes, so CTP's data frames carry a time-has-lived count that distinguishes a looping packet from a duplicate sent twice.

Distinctions

Hop-count routingETX routing
Cost of a link11 / (forward × backward quality)
PrefersFew, long hopsMore, short, reliable hops
Program: routes3.56 hops, 18.06 transmissions5.08 hops, 11.46 transmissions
NeedsNothing but connectivityLink estimation in both directions
Beacon-based estimateData-based estimate
FromPeriodic routing beaconsAcknowledged data transmissions
RoleSeeds the table, covers idle linksMeasures the links actually used, fast
RateSlows in a stable network (Trickle)As fast as data is sent
Loop detectionETX cap
CatchesInconsistency seen on the data pathA partitioned loop counting upwards
ActionBroadcast a beacon to fix the senderIgnore routes costing more than the limit

What it does not mean

A routing table is not a list of destinations here. A collection node has one destination; its table is a list of neighbours and their costs to it.

munotes.in430

Routing Tables and What Happens When the Topology Changes

The lowest total is not always chosen. A parent is kept unless another neighbour is better by the margin; stability is worth a little cost.

More hops is not worse. On lossy links, ETX's longer routes need fewer transmissions, and so less energy and delay.

A dead node is not reported by anyone. Its children notice missing acknowledgements or beacons and choose again from their tables; the change spreads through ordinary beacons.

Quick revision

  • Collection tree (CTP): data flows to a root through a parent; address-free: choosing a next hop chooses the root.
  • ETX gradient: root 0; node ETX = parent's ETX + link ETX; link ETX = 1 / (forward × backward quality) (the MT metric of Woo, Tong and Culler).
  • Routing table: neighbour, link ETX, advertised path ETX, total; parent = lowest total, changed only if better by a margin.
  • Beacons advertise path ETX on a Trickle-like timer, reset when the table is empty, ETX rises by 1 or more, or a neighbour asks. Link estimates from beacons and from data (every 5 transmissions; none acknowledged counts as ETX 6).
  • Loops: a node choosing a former descendant; CTP detects data from a lower-ETX sender and beacons; caps ETX.
  • Program: tree in 9 rounds; hop-count routes 3.56 hops, 18.06 ETX against ETX routes 5.08 hops, 11.46 ETX; the busiest relay (on 34 routes) dies: re-formed in 11 rounds, none cut off, ETX up 0.90 on average.

Test yourself

1. What does a node's routing table contain in a collection tree protocol such as CTP? For each neighbour: the ETX of the link to it, the path ETX it last advertised to the root, and their sum, the cost of routing through it. The node chooses as its parent the neighbour with the lowest sum, and its own path ETX, which it advertises in its beacons, is that sum.

2. What is ETX, and why is it a better routing metric than hop count in a sensor network? ETX is the expected number of transmissions, including retransmissions, needed to deliver a packet and receive its acknowledgement over a link, estimated as 1 over the product of the forward and backward delivery ratios; a path's ETX is the sum of its links' ETX. Hop count ignores link quality and so prefers long, lossy links that need many retransmissions. In the program, hop-count routes averaged 3.56 hops but 18.06 transmissions, ETX routes 5.08 hops but 11.46 transmissions.

3. Compute the ETX of a link with 80 per cent forward and 75 per cent backward delivery, and compare it with a route of two links of ETX 1.2 each. 1 / (0.80 × 0.75) = 1 / 0.6, about 1.67. The two-link route costs 1.2 + 1.2 = 2.4, so the single link is cheaper despite being less reliable than either of the two.

munotes.in431

Routing Tables and What Happens When the Topology Changes

4. Why does CTP use a margin when choosing a parent? Link estimates fluctuate, and if a node switched to whichever neighbour looked slightly cheaper each time, parents would change constantly and every change would disturb the tree below. Keeping the current parent unless another is better by a margin makes the tree stable at a small cost in optimality; in the program, a node kept a parent costing 0.18 more than the best, within the 0.5 margin.

5. What happens in a collection tree when a relay node dies? Its children stop receiving acknowledgements or beacons from it, remove it from their tables, and choose the neighbour with the next lowest total as their new parent. Their path ETX changes, their beacons carry the new value, and the change spreads down their subtrees until the tree is consistent again. In the program, a relay on 34 nodes' routes died and the tree re-formed in 11 beacon rounds with every node still connected.

6. How do routing loops form, and how does CTP deal with them? When a node loses its parent it may choose a neighbour that was its own descendant and still advertises an old, cheap route through the lost parent; the two then point at each other and their advertised ETX climbs. CTP detects loops on the data path: every data frame carries the sender's ETX, and a node that receives data from a node with a lower ETX than its own broadcasts a beacon to correct it. It also ignores routes whose ETX exceeds a set limit, which stops a partitioned loop counting upwards for ever.

Contents This chapter on its own page

munotes.in432

Chapter Sixty-Three

IEEE 802.15.4: The Standard, Its Devices and Its Topologies

Syllabus topic Module 1, "Routing in WSN: IEEE 802.15.4 LR-WPAN standard (case study)"

In one line

IEEE 802.15.4 defines the radio and the MAC for low-rate, low-power, low-cost networks: full-function devices can coordinate and relay, reduced-function devices only talk to one full-function device, and the network is a star around a PAN coordinator, a peer-to-peer mesh, or a cluster tree, with everything above the MAC left to other standards such as Zigbee and 6LoWPAN.

In the wording a student can write in an examination: IEEE 802.15.4 specifies the physical layer (PHY) and the medium access control (MAC) sublayer for low-rate wireless personal area networks (LR-WPANs), whose objectives are ease of installation, reliable data transfer, short range, extremely low cost and reasonable battery life. Its characteristics include data rates of 250, 100, 40 and 20 kb/s, star or peer-to-peer operation, 16-bit short or 64-bit extended addresses, optional guaranteed time slots (GTSs), CSMA-CA channel access, acknowledged transfer, low power consumption, energy detection and link quality indication, and channels in the 868 MHz, 915 MHz and 2450 MHz bands. There are two device types: a full-function device (FFD), which can act as a PAN coordinator, a coordinator or a device, and a reduced-function device (RFD), which is simple (a light switch, a passive infrared sensor), talks only to an FFD and cannot relay. Every network has exactly one PAN coordinator. In the star topology all devices communicate with the PAN coordinator; in the peer-to-peer topology any device may talk to any other in range, allowing mesh networks; the cluster tree is a peer-to-peer network of coordinators with RFDs as leaves, which extends coverage at the cost of latency. Network formation beyond this and multi-hop routing belong to higher layers: Zigbee adds a network layer with star, tree and mesh routing, and 6LoWPAN carries IPv6 over 802.15.4.

What 802.15.4 is for

The standard opens with its purpose. "An LR-WPAN is a simple, low-cost communication network that allows wireless connectivity in applications with limited power and relaxed throughput requirements. The main objectives of an LR-WPAN are ease of installation, reliable data transfer, short-range operation, extremely low cost, and a reasonable battery life, while maintaining a simple and flexible protocol."

Its list of characteristics, in the 2006 edition:

  • over-the-air data rates of 250 kb/s, 100 kb/s, 40 kb/s and 20 kb/s;
  • star or peer-to-peer operation;
  • allocated 16-bit short or 64-bit extended addresses;
  • optional allocation of guaranteed time slots (GTSs);
  • CSMA-CA channel access;
  • a fully acknowledged protocol for transfer reliability;
  • low power consumption;
  • energy detection (ED) and link quality indication (LQI);
  • 16 channels in the 2450 MHz band, 30 in the 915 MHz band and 3 in the 868 MHz band.

These are the choices of a sensor radio: a modest rate, simple access, and features (energy detection, link quality) that the layers above use to pick channels and routes. The CC2420 of the Telos motes is an 802.15.4 radio ([The Radio, the Sensors and the Power Supply of a Node]).

munotes.in433

IEEE 802.15.4: The Standard, Its Devices and Its Topologies

Two kinds of device, three roles

"Two different device types can participate in an IEEE 802.15.4 network; a full-function device (FFD) and a reduced-function device (RFD). The FFD can operate in three modes serving as a personal area network (PAN) coordinator, a coordinator, or a device. An FFD can talk to RFDs or other FFDs, while an RFD can talk only to an FFD."

The definitions in clause 3 make the roles precise. A coordinator is "A full-function device (FFD) capable of relaying messages. If a coordinator is the principal controller of a personal area network (PAN), it is called the PAN coordinator." And "An IEEE 802.15.4 network has exactly one PAN coordinator."

The RFD is the cheap end of the range: it "is intended for applications that are extremely simple, such as a light switch or a passive infrared sensor; they do not have the need to send large amounts of data and may only associate with a single FFD at a time. Consequently, the RFD can be implemented using minimal resources and memory capacity."

Addresses. Every device has a unique 64-bit extended address, and on joining may be given a 16-bit short address by the PAN coordinator, saving bytes in every frame. Each network chooses a PAN identifier, which lets devices use short addresses within the network and still talk across networks.

The smallest network. "Two or more devices within a POS communicating on the same physical channel constitute a WPAN. However, this WPAN shall include at least one FFD, operating as the PAN coordinator." (POS, the personal operating space, is the area around a person or object a WPAN typically covers; the standard adds that "A well-defined coverage area does not exist for wireless media because propagation characteristics are dynamic and uncertain".)

Three topologies

Three panels. Star: a PAN coordinator at the centre with spokes to full-function and reduced-function devices around it. Peer-to-peer: full-function devices joined in a mesh, one of them the PAN coordinator, with a few reduced-function devices attached by dashed lines to single full-function devices. Cluster tree: a PAN coordinator at the top, three full-function devices below it, and reduced-function devices as leaves under each

Figure 63.1 Star, peer-to-peer and cluster tree, after the standard's Figs. 1 and 2

Star. "In the star topology the communication is established between devices and a single central controller, called the PAN coordinator." The coordinator "might often be mains powered, while the devices will most likely be battery powered", and the standard's examples are home automation, PC peripherals, toys and games, and personal health care. To form one, an FFD that is switched on can "establish its own network and become the PAN coordinator", choosing a PAN identifier not used nearby, and then allows other devices to join.

Peer-to-peer. "The peer-to-peer topology also has a PAN coordinator; however, it differs from the star topology in that any device may communicate with any other device as long as they are in range of one another." It allows mesh networks, "ad hoc, self-organizing, and self-healing", and multi-hop routing, but "Such functions can be added at the higher layer, but are not part of this standard." Its applications include "industrial control and monitoring, wireless sensor networks, asset and inventory tracking, intelligent agriculture, and security".

munotes.in434

IEEE 802.15.4: The Standard, Its Devices and Its Topologies

Cluster tree. "The cluster tree network is a special case of a peer-to-peer network in which most devices are FFDs. An RFD connects to a cluster tree network as a leaf device at the end of a branch because RFDs do not allow other devices to associate." It grows by beacons: the PAN coordinator "forms the first cluster by choosing an unused PAN identifier and broadcasting beacon frames to neighboring devices"; a device that joins adds the coordinator as its parent and, if it is an FFD, "begins transmitting periodic beacons; other candidate devices may then join the network at that device." Several clusters can form a larger network, and the standard names the trade: "The advantage of a multicluster structure is increased coverage area, while the disadvantage is an increase in message latency."

Who can join: a star and a cluster tree, formed

The program scatters 80 devices at random over 120 m by 120 m with the PAN coordinator at the centre and a range of 20 m, both illustrative. In a star, a device joins only if it can hear the PAN coordinator. In a cluster tree, a device joins through any FFD that has already joined and is beaconing; an FFD that joins becomes a coordinator in turn, while an RFD stays a leaf. The same devices are given FFD shares of 20, 50 and 100 per cent.

# Who can join an 802.15.4 network: a star against a cluster tree. 80 devices
# at random over 120 m by 120 m, the PAN coordinator at the centre, a range of
# 20 m (illustrative). In a star, a device joins only if it can hear the PAN
# coordinator. In a cluster tree, a device joins through any FFD that has
# already joined and beacons; an FFD that joins becomes a coordinator in turn,
# while an RFD stays a leaf, since RFDs "do not allow other devices to associate".
import math
import random
from collections import deque

rnd = random.Random(63)
N, SIDE, RANGE = 80, 120.0, 20.0
PAN = (SIDE / 2, SIDE / 2)
devices = [(rnd.uniform(0, SIDE), rnd.uniform(0, SIDE)) for _ in range(N)]
draws = [rnd.random() for _ in range(N)]           # fixed, so every share is comparable

def cluster_tree(ffd_share):
    ffd = [d < ffd_share for d in draws]
    depth = {"PAN": 0}
    queue, coordinators = deque([("PAN", PAN)]), 1
    while queue:
        name, where = queue.popleft()
        for i, p in enumerate(devices):
            if i not in depth and math.dist(p, where) <= RANGE:
                depth[i] = depth[name] + 1          # associates with this coordinator
                if ffd[i]:
                    coordinators += 1
                    queue.append((i, p))            # beacons: others may join here
    joined = [d for k, d in depth.items() if k != "PAN"]
    return len(joined), max(joined), sum(joined) / len(joined), coordinators

star = sum(1 for p in devices if math.dist(p, PAN) <= RANGE)
print("star: %d of %d devices can hear the PAN coordinator" % (star, N))
print("cluster tree    joined  deepest  mean depth  coordinators")
for share in (0.2, 0.5, 1.0):
    joined, deepest, mean, coords = cluster_tree(share)
    print("%3.0f%% FFDs %11d %8d %11.2f %13d" % (100 * share, joined, deepest, mean, coords))
munotes.in435

IEEE 802.15.4: The Standard, Its Devices and Its Topologies

star: 11 of 80 devices can hear the PAN coordinator
cluster tree    joined  deepest  mean depth  coordinators
 20% FFDs          29        4        2.00             8
 50% FFDs          62       10        3.47            36
100% FFDs          77        7        3.25            78

Reading it. A star reaches only the 11 devices within one hop of the PAN coordinator: its coverage is one radio range. A cluster tree in which only a fifth of the devices are FFDs reaches 29, because RFDs cannot pass the network on; half FFDs reach 62; all FFDs reach 77, the rest being out of range of everyone. Coverage costs depth: with half FFDs the deepest device is 10 hops from the PAN coordinator, 3.47 on average, and every hop adds latency, most of all in a beacon-enabled tree where each coordinator transmits on its own schedule. With all FFDs, the tree is shallower (7 hops at most, 3.25 on average), since every device offers a way on. The design choice is the standard's: cheap RFDs as leaves, enough FFDs to hold the network together.

Two layers, and what lies above

"An LR-WPAN device comprises a PHY, which contains the radio frequency (RF) transceiver along with its low-level control mechanism, and a MAC sublayer that provides access to the physical channel for all types of transfer."

  • PHY. "The features of the PHY are activation and deactivation of the radio transceiver, ED, LQI, channel selection, clear channel assessment (CCA), and transmitting as well as receiving packets across the physical medium." It uses the unlicensed bands at 868 to 868.6 MHz (Europe), 902 to 928 MHz (North America) and 2400 to 2483.5 MHz (worldwide). [The 802.15.4 Physical Layer] takes it up.
  • MAC. "The features of the MAC sublayer are beacon management, channel access, GTS management, frame validation, acknowledged frame delivery, association, and disassociation. In addition, the MAC sublayer provides hooks for implementing application-appropriate security mechanisms." [The 802.15.4 Superframe and Guaranteed Time Slots] and [CSMA-CA, Data Transfer and Frames in 802.15.4] take it up.
munotes.in436

IEEE 802.15.4: The Standard, Its Devices and Its Topologies

Everything else is left above: "The network formation is performed by the higher layer, which is not part of this standard." That is why MU's placing of 802.15.4 under routing needs care: the standard makes routing possible (addresses, peer-to-peer links, cluster trees) but does not route.

Power. The standard was written for batteries. "Battery-powered devices will require duty-cycling to reduce power consumption. These devices will spend most of their operational life in a sleep state; however, each device periodically listens to the RF channel in order to determine whether a message is pending. This mechanism allows the application designer to decide on the balance between battery consumption and message latency." ([Duty Cycling: Preamble Sampling, B-MAC and X-MAC] discussed the same balance.)

Zigbee is the best-known stack above it. "The IEEE 802.15.4 standard defines the two lower layers: the physical (PHY) layer and the medium access control (MAC) sub-layer. The ZigBee Alliance builds on this foundation by providing the network (NWK) layer and the framework for the application layer." "The ZigBee network layer (NWK) supports star, tree, and mesh topologies", with its own names for the roles: the Zigbee coordinator (the PAN coordinator), Zigbee routers (FFDs that route) and end devices. [Built on 802.15.4: Zigbee Routing, Security and the Later Amendments] covers its routing.

6LoWPAN carries IPv6 over 802.15.4 instead: RFC 4944 describes the "Transmission of IPv6 Packets over IEEE 802.15.4 Networks", the basis of [Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP].

Distinctions

Full-function device (FFD)Reduced-function device (RFD)
RolesPAN coordinator, coordinator or deviceDevice only
Talks toFFDs and RFDsOne FFD
Relays, lets others joinYesNo: a leaf
ExamplesA mains-powered controller, a routerA light switch, a passive infrared sensor
ResourcesMoreMinimal
StarPeer-to-peerCluster tree
CommunicationEvery device with the PAN coordinatorAny device with any other in rangeAlong parent and child links
Multi-hopNoPossible (mesh), by higher layersYes, through coordinators
CoverageOne radio rangeWideWide, grows with clusters
WeaknessRange; the coordinator carries everythingRouting is left to higher layersLatency grows with depth
Program11 of 80 devices29 to 77 of 80, by FFD share
PAN coordinatorCoordinator
NumberExactly one per networkAny number
JobStarts the network, chooses the PAN identifier, principal controllerRelays, beacons, lets devices join

What it does not mean

802.15.4 is not Zigbee. It defines only the PHY and MAC; Zigbee, 6LoWPAN and others are built on it.

Peer-to-peer does not mean routed. Devices in range can talk to each other; forwarding across several hops is a higher layer's job.

munotes.in437

IEEE 802.15.4: The Standard, Its Devices and Its Topologies

An RFD is not a broken FFD. It is deliberately simple and cheap, for devices that only report or obey.

One PAN coordinator does not mean a star. Every topology has one; in a peer-to-peer or cluster-tree network it is simply the device that started the network.

Quick revision

  • 802.15.4: PHY and MAC for LR-WPANs: easy to install, reliable, short range, extremely low cost, reasonable battery life.
  • Characteristics: 250/100/40/20 kb/s; star or peer-to-peer; 16-bit short / 64-bit extended addresses; optional GTSs; CSMA-CA; acknowledgements; low power; ED and LQI; 16 channels at 2450 MHz, 30 at 915, 3 at 868.
  • FFD: PAN coordinator, coordinator or device; talks to anyone. RFD: talks to one FFD only; a leaf; minimal resources. Exactly one PAN coordinator per network; at least one FFD.
  • Star (all to the PAN coordinator), peer-to-peer (any to any in range; mesh by higher layers), cluster tree (FFD coordinators, RFD leaves; coverage against latency).
  • Program: star 11 of 80; cluster tree 29 / 62 / 77 with 20 / 50 / 100 per cent FFDs; deepest 4 to 10 hops.
  • PHY: transceiver on and off, ED, LQI, channel selection, CCA, send and receive. MAC: beacons, channel access, GTSs, frame validation, acknowledgements, association, security hooks.
  • Above it: Zigbee (network and application layers; coordinator, router, end device; star, tree, mesh) and 6LoWPAN (IPv6, RFC 4944).

Test yourself

1. What is an LR-WPAN, and what are its objectives? A low-rate wireless personal area network is a simple, low-cost wireless network for applications with limited power and relaxed throughput requirements. IEEE 802.15.4 gives its objectives as ease of installation, reliable data transfer, short-range operation, extremely low cost and a reasonable battery life, while keeping the protocol simple and flexible.

2. Distinguish an FFD and an RFD. A full-function device can operate as a PAN coordinator, a coordinator or an ordinary device, can talk to both FFDs and RFDs, and can relay messages and let other devices join through it. A reduced-function device is intended for very simple applications such as a light switch or a passive infrared sensor; it can talk only to an FFD, associates with a single FFD at a time, cannot relay or accept other devices, and so needs minimal resources.

3. Describe the three topologies of 802.15.4. In the star topology every device communicates only with a single central PAN coordinator. In the peer-to-peer topology there is still a PAN coordinator, but any device may communicate with any other within range, which allows mesh networks, with multi-hop routing added by higher layers. The cluster tree is a peer-to-peer network in which most devices are FFDs acting as coordinators arranged in a tree under the PAN coordinator, with RFDs as leaves; it covers a larger area at the cost of more latency.

munotes.in438

IEEE 802.15.4: The Standard, Its Devices and Its Topologies

4. How is a cluster tree formed? The PAN coordinator chooses an unused PAN identifier and broadcasts beacons. A device that hears a beacon asks to join; if accepted, it records the coordinator as its parent and, if it is an FFD, starts sending its own beacons so that other devices can join through it. RFDs join as leaves. The first PAN coordinator may also instruct a device to become the head of a new, adjacent cluster.

5. Which features belong to the 802.15.4 PHY and which to its MAC? The PHY activates and deactivates the transceiver, performs energy detection and link quality indication, selects channels, performs clear channel assessment, and transmits and receives packets. The MAC manages beacons, channel access and guaranteed time slots, validates frames, delivers acknowledged frames, handles association and disassociation, and provides hooks for security.

6. In the program, why did a cluster tree with 20 per cent FFDs reach far fewer devices than one with 100 per cent? Only FFDs can let other devices join through them; RFDs remain leaves. With few FFDs, most joined devices could not extend the network, so devices beyond their reach could not join: 29 of 80 joined, against 77 when every device was an FFD and could act as a coordinator. A star reached only the 11 devices within one hop of the PAN coordinator.

Contents This chapter on its own page

munotes.in439

Chapter Sixty-Four

The 802.15.4 Physical Layer

Syllabus topic Module 1, "Routing in WSN: IEEE 802.15.4 LR-WPAN standard (case study)"

In one line

The 802.15.4 physical layer carries the MAC's frames over the air: in the 2450 MHz band it sends 250 kb/s on one of 16 channels 5 MHz apart, turning every 4 bits into one of 16 sequences of 32 chips sent by O-QPSK at 2 Mchip/s, inside a packet made of a synchronisation header, a one-octet length field and at most 127 octets of payload; and it measures channel energy, grades each received packet's link quality and tells the MAC whether the channel is clear.

In the wording a student can write in an examination: the PHY of IEEE 802.15.4 is responsible for activating and deactivating the radio transceiver, energy detection (ED) within the current channel, the link quality indicator (LQI) for received packets, clear channel assessment (CCA) for CSMA-CA, channel frequency selection, and data transmission and reception. The 2006 edition defines four PHYs: an 868/915 MHz DSSS PHY with BPSK (20 kb/s at 868 MHz, 40 kb/s at 915 MHz); two optional 868/915 MHz PHYs, one with ASK and parallel sequence spread spectrum (250 kb/s in both bands) and one with O-QPSK (100 kb/s at 868 MHz, 250 kb/s at 915 MHz); and a 2450 MHz DSSS PHY with O-QPSK at 250 kb/s, the band available worldwide. Channel page 0 has 27 channels: channel 0 at 868.3 MHz, channels 1 to 10 at 906 + 2(k - 1) MHz, and channels 11 to 26 at 2405 + 5(k - 11) MHz.

At 2450 MHz every octet becomes two 4-bit symbols, and every symbol one of 16 PN sequences of 32 chips, sent at 62.5 ksymbol/s and 2 Mchip/s by O-QPSK with half-sine pulses (even-indexed chips on the I carrier, odd-indexed on the Q carrier, Q delayed by one chip period). The PPDU is a synchronisation header (SHR) of a 4-octet preamble of zeros and a 1-octet start-of-frame delimiter (SFD), a PHY header (PHR) holding a 7-bit frame length, and the PSDU, which carries the MAC frame and is at most 127 octets (aMaxPHYPacketSize). ED estimates the received power in a channel, for choosing channels; LQI characterises the strength and/or quality of each received packet; CCA reports the channel busy by energy above a threshold (mode 1), by carrier sense of an 802.15.4 signal (mode 2), or by a combination of the two (mode 3).

What the physical layer is responsible for

Clause 6 of the standard opens with the PHY's duties. "The PHY is responsible for the following tasks:"

  • "Activation and deactivation of the radio transceiver"
  • "Energy detection (ED) within the current channel"
  • "Link quality indicator (LQI) for received packets"
  • "Clear channel assessment (CCA) for carrier sense multiple access with collision avoidance (CSMA-CA)"
  • "Channel frequency selection"
  • "Data transmission and reception"
munotes.in440

The 802.15.4 Physical Layer

It offers them to the MAC as two services. "The PHY data service enables the transmission and reception of PHY protocol data units (PPDUs) across the physical radio channel." The management service is how the MAC asks for the rest: an energy measurement, a clear channel assessment, a change of channel. In the standard's names, constants start with a small a (aMaxPHYPacketSize, aTurnaroundTime) and the PHY's managed attributes with phy (phyCurrentChannel, phyCCAMode).

Turning the transceiver off is the first duty for a reason: a radio that is off draws almost nothing, and every MAC in [Duty Cycling: Preamble Sampling, B-MAC and X-MAC] depends on the PHY switching it on and off on command.

Four PHYs in three bands

"The standard specifies the following four PHYs:"

  • "An 868/915 MHz direct sequence spread spectrum (DSSS) PHY employing binary phase-shift keying (BPSK) modulation"
  • "An 868/915 MHz DSSS PHY employing offset quadrature phase-shift keying (O-QPSK) modulation"
  • "An 868/915 MHz parallel sequence spread spectrum (PSSS) PHY employing BPSK and amplitude shift keying (ASK) modulation"
  • "A 2450 MHz DSSS PHY employing O-QPSK modulation"

The three bands are unlicensed: 868 to 868.6 MHz (the standard's example is Europe), 902 to 928 MHz (North America) and 2400 to 2483.5 MHz (worldwide). The BPSK PHY is the original of the 2003 edition; the two optional PHYs were added to raise the rate in the lower bands. The standard's Table 1, set out here:

PHYBand (MHz)ModulationChip rate (kchip/s)Bit rate (kb/s)Symbol rate (ksymbol/s)Symbols
868/915 MHz868 to 868.6BPSK3002020Binary
902 to 928BPSK6004040Binary
868/915 MHz (optional)868 to 868.6ASK40025012.520-bit PSSS
902 to 928ASK1600250505-bit PSSS
868/915 MHz (optional)868 to 868.6O-QPSK4001002516-ary orthogonal
902 to 928O-QPSK100025062.516-ary orthogonal
2450 MHz2400 to 2483.5O-QPSK200025062.516-ary orthogonal

The 2450 MHz PHY is the one met in sensor networks: the TI CC2420, the radio of the Telos motes in [The Radio, the Sensors and the Power Supply of a Node], is a "2.4 GHz IEEE 802.15.4 compliant RF transceiver". The rest of this chapter follows that PHY.

The channels, computed

"For channel page 0, 27 channels numbered 0 to 26 are available across the three frequency bands. Sixteen channels are available in the 2450 MHz band, 10 in the 915 MHz band, and 1 in the 868 MHz band." The centre frequency of channel k, in MHz, is:

  • channel 0: 868.3;
  • channels 1 to 10: 906 + 2(k - 1), so 2 MHz apart from 906 to 924;
  • channels 11 to 26: 2405 + 5(k - 11), so 5 MHz apart from 2405 to 2480.
munotes.in441

The 802.15.4 Physical Layer

Worked: channel 15 is at 2405 + 5 × 4 = 2425 MHz, and channel 26 at 2405 + 5 × 15 = 2480 MHz; channel 10 is at 906 + 2 × 9 = 924 MHz.

Channel pages. The optional PHYs reuse the channel numbers. "A total of 32 channel pages are available with channel pages 3 to 31 being reserved for future use." Page 0 holds the 27 channels of the 2003 edition; on pages 1 and 2, channels 0 to 10 sit at the same centre frequencies but belong to the ASK and the O-QPSK PHYs respectively. So a channel number means something only with its page. The count in the standard's list of characteristics, 30 channels at 915 MHz and 3 at 868 MHz, matches ten and one on each of the three pages.

Living beside Wi-Fi

The 2450 MHz band is shared with Wi-Fi, Bluetooth and microwave ovens, and the standard's informative Annex E studies the neighbours. Its Figure E.1 draws the three non-overlapping 802.11b channels, 22 MHz wide, over the sixteen 802.15.4 channels, 2 MHz wide: 802.11b channels 1, 6 and 11 (2412, 2437 and 2462 MHz) in North America, and 1, 7 and 13 (2412, 2442 and 2472 MHz) in Europe.

"There are four IEEE 802.15.4 channels that fall in the guard bands between (or above) the three IEEE 802.11b channels": channels 15, 20, 25 and 26 in North America, and 15, 16, 21 and 22 in Europe. The energy there is not zero, the annex adds, but it is lower than within the Wi-Fi channels, and "operating an IEEE 802.15.4 network on one of these channels will minimize interference between systems." The program below finds the same four channels from the widths alone.

A network does not have to guess. "When performing dynamic channel selection, either at network initialization or in response to an outage, an IEEE 802.15.4 device will scan a set of channels specified by the ChannelList parameter." Where Wi-Fi is known to be busy, that list can be set to the four clear channels.

From octets to chips

The 2450 MHz PHY sends data in three steps, each in its own subclause.

Bits to symbols. "The 4 LSBs (b0, b1, b2, b3) of each octet shall map into one data symbol, and the 4 MSBs (b4, b5, b6, b7) of each octet shall map into the next data symbol." A symbol is therefore a number from 0 to 15, and an octet is two symbols, low half first.

Symbols to chips. "Each data symbol shall be mapped into a 32-chip PN sequence as specified in Table 24. The PN sequences are related to each other through cyclic shifts and/or conjugation (i.e., inversion of odd-indexed chip values)." Table 24 lists the sixteen sequences; the program holds them and checks that relation.

munotes.in442

The 802.15.4 Physical Layer

Chips to a signal. "Even-indexed chips are modulated onto the in-phase (I) carrier and odd-indexed chips are modulated onto the quadrature-phase (Q) carrier. Because each data symbol is represented by a 32-chip sequence, the chip rate (nominally 2.0 Mchip/s) is 32 times the symbol rate." To make the offset of offset QPSK, the Q-phase chips are delayed by Tc, "where Tc is the inverse of the chip rate", and each chip is a half-sine pulse lasting 2Tc. A new chip therefore starts every Tc, in turn on I and on Q.

The rates follow. 250 kb/s divided into 4-bit symbols is 62.5 ksymbol/s, a symbol every 16 µs, and 32 chips per symbol is 2 Mchip/s: 8 chips for every bit.

Why spend 32 chips on 4 bits? Annex E answers. The 2450 MHz PHY "uses a quasi-orthogonal modulation scheme, where each symbol is represented by one of 16 nearly orthogonal PN sequences. This is a power-efficient modulation method that achieves low signal-to-noise ratio (SNR) and signal-to-interference ratio (SIR) requirements at the expense of a signal bandwidth that is significantly larger than the symbol rate." And "A typical low-cost detector implementation is expected to meet the 1% packet error rate (PER) requirement at SNR values of 5 dB to 6 dB." A wide signal is bought in exchange for a receiver that still decodes when the signal is barely above the noise, which is what a low-power radio at the edge of its range needs. The general theory is in [Spread Spectrum and Direct Sequence], and QPSK with its offset form in [Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM].

The PHY packet

"Each PPDU packet consists of the following basic components:"

  • "A synchronization header (SHR), which allows a receiving device to synchronize and lock onto the bit stream"
  • "A PHY header (PHR), which contains frame length information"
  • "A variable length payload, which carries the MAC sublayer frame"
Top: a PHY packet at 2450 MHz as four boxes, a 4-octet preamble, a 1-octet SFD, a 1-octet PHR holding the length, and the PSDU carrying the MAC frame of up to 127 octets, with the preamble and SFD bracketed as the synchronisation header; at 250 kb/s an octet lasts 32 microseconds and the longest packet, 133 octets, 4.256 ms. Bottom: an axis from 2400 to 2483.5 MHz with sixteen narrow channel bars, channels 11 to 26, 5 MHz apart, centre 2405 + 5(k - 11) MHz

Figure 64.1 The 2450 MHz PPDU, after the standard's Fig. 16 and Tables 19 and 20, and the band's sixteen channels

Preamble. "The Preamble field is used by the transceiver to obtain chip and symbol synchronization with an incoming message." "For all PHYs except the ASK PHY, the bits in the Preamble field shall be binary zeros." At 2450 MHz it is 4 octets: 8 symbols, all symbol 0, lasting 128 µs (Table 19).

SFD. "The SFD is a field indicating the end of the SHR and the start of the packet data." It is one octet, whose bits b0 to b7 are 1 1 1 0 0 1 0 1 (Figure 17). The preamble and the SFD together form the SHR.

munotes.in443

The 802.15.4 Physical Layer

PHR. "The Frame Length field is 7 bits in length and specifies the total number of octets contained in the PSDU (i.e., PHY payload)." A reserved bit completes the octet. Of the values, 0 to 4 and 6 to 8 are reserved, 5 means an acknowledgement frame, and 9 to 127 a MAC frame (Table 21).

PSDU. The MAC frame itself. Seven bits count to 127 at most, and the constant aMaxPHYPacketSize, "The maximum PSDU size (in octets) the PHY shall be able to receive", is 127.

Order on the air. "All multiple octet fields shall be transmitted or received least significant octet first and each octet shall be transmitted or received least significant bit (LSB) first."

Why 127 octets matters. The MAC frame inside must carry its own header and checksum, and the layers above feel the squeeze. RFC 4944, which carries IPv6 over 802.15.4, starts "from a maximum physical layer packet size of 127 octets (aMaxPHYPacketSize) and a maximum frame overhead of 25 (aMaxFrameOverhead)", leaving 102 octets, and with the heaviest link-layer security "only 81 octets available". IPv6 needs 1280, so 6LoWPAN must fragment ([Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP]).

The PHY, computed

The program computes the channel plan and the channels clear of Wi-Fi, derives the rates, checks the two relations the standard states for Table 24, measures how far apart the sixteen sequences are, and then puts them to work: it flips chips at random in sent sequences and decodes by the nearest sequence, spells out the header of a frame as symbols, and times frames of three sizes.

# The 2450 MHz PHY of IEEE 802.15.4-2006, checked and put to work. CHIPS is
# the standard's Table 24: each 4-bit data symbol (0 to 15) is sent as one of
# 16 sequences of 32 chips, at 62.5 ksymbol/s.
import random

CHIPS = ["11011001110000110101001000101110", "11101101100111000011010100100010",
         "00101110110110011100001101010010", "00100010111011011001110000110101",
         "01010010001011101101100111000011", "00110101001000101110110110011100",
         "11000011010100100010111011011001", "10011100001101010010001011101101",
         "10001100100101100000011101111011", "10111000110010010110000001110111",
         "01111011100011001001011000000111", "01110111101110001100100101100000",
         "00000111011110111000110010010110", "01100000011101111011100011001001",
         "10010110000001110111101110001100", "11001001011000000111011110111000"]

def centre(k):                       # 6.1.2.1, channel page 0, in MHz
    if k == 0:
        return 868.3
    return 906 + 2 * (k - 1) if k <= 10 else 2405 + 5 * (k - 11)

print("868 MHz band: channel 0 at %.1f MHz" % centre(0))
print("915 MHz band: channels 1 to 10 at %d to %d MHz" % (centre(1), centre(10)))
print("2450 MHz band: channels 11 to 26 at %d to %d MHz, 5 MHz apart" % (centre(11), centre(26)))
# living beside Wi-Fi, after the standard's Figure E.1: which of channels 11 to
# 26 (2 MHz wide) reach into none of three 802.11b channels (22 MHz wide)?
for region, wifi in (("North America", (2412, 2437, 2462)), ("Europe", (2412, 2442, 2472))):
    clear = [k for k in range(11, 27)
             if all(centre(k) + 1 <= f - 11 or centre(k) - 1 >= f + 11 for f in wifi)]
    print("%s, 802.11b at %d, %d, %d MHz: clear channels %s" % ((region,) + wifi + (clear,)))
print("rates: 62.5 ksymbol/s x 4 bits = %.0f kb/s; x 32 chips = %.1f Mchip/s"
      % (62.5 * 4, 62.5 * 32 / 1000))

# the table's structure: 1 to 7 are 0 rotated by 4, 8, ... chips; 8 to 15 are
# 0 to 7 with every odd-indexed chip inverted
rot = lambda s, n: s[-n:] + s[:-n]
flip_odd = lambda s: "".join(c if i % 2 == 0 else "10"[int(c)] for i, c in enumerate(s))
print("symbols 1 to 7 are symbol 0 rotated:",
      all(CHIPS[k] == rot(CHIPS[0], 4 * k) for k in range(1, 8)))
print("symbols 8 to 15 are 0 to 7 with odd chips inverted:",
      all(CHIPS[k + 8] == flip_odd(CHIPS[k]) for k in range(8)))

distance = lambda a, b: sum(x != y for x, y in zip(a, b))
pairs = [distance(CHIPS[i], CHIPS[j]) for i in range(16) for j in range(i + 1, 16)]
print("chips differing between any two sequences: at least %d, at most %d" % (min(pairs), max(pairs)))

# despreading: flip k chips of a sent sequence at random; decide by the nearest sequence
rnd = random.Random(64)
print("chip errors in a symbol   symbol decoded wrongly")
for k in (2, 4, 6, 8, 10):
    wrong = 0
    for _ in range(10000):
        sym = rnd.randrange(16)
        rx = list(CHIPS[sym])
        for i in rnd.sample(range(32), k):
            rx[i] = "10"[int(rx[i])]
        rx = "".join(rx)
        best = min(range(16), key=lambda s: (distance(CHIPS[s], rx), rnd.random()))
        wrong += best != sym
    print("%12d of 32 %22.2f%%" % (k, 100 * wrong / 10000))

# the SHR and PHR as symbols: four zero octets of preamble, the SFD of Figure 17
# (bits b0 to b7 are 1 1 1 0 0 1 0 1), then the PHR holding the PSDU length in
# 7 bits with the reserved bit 0. Each octet gives two symbols, its 4 LSBs first.
sfd = sum(bit << i for i, bit in enumerate([1, 1, 1, 0, 0, 1, 0, 1]))
header = [0, 0, 0, 0, sfd, 20]
print("SFD octet 0x%02X; the SHR and PHR before a 20-octet PSDU, as symbols:" % sfd)
print([s for octet in header for s in (octet & 15, octet >> 4)])

# frame time: 4 octets of preamble, 1 of SFD, 1 of PHY header, then up to 127
for psdu in (20, 50, 127):
    octets = 6 + psdu
    print("PSDU %3d octets: %3d octets on air, %.3f ms at 250 kb/s; header share %.1f%%"
          % (psdu, octets, octets * 8 / 250_000 * 1000, 100 * 6 / octets))
munotes.in444

The 802.15.4 Physical Layer

868 MHz band: channel 0 at 868.3 MHz
915 MHz band: channels 1 to 10 at 906 to 924 MHz
2450 MHz band: channels 11 to 26 at 2405 to 2480 MHz, 5 MHz apart
North America, 802.11b at 2412, 2437, 2462 MHz: clear channels [15, 20, 25, 26]
Europe, 802.11b at 2412, 2442, 2472 MHz: clear channels [15, 16, 21, 22]
rates: 62.5 ksymbol/s x 4 bits = 250 kb/s; x 32 chips = 2.0 Mchip/s
symbols 1 to 7 are symbol 0 rotated: True
symbols 8 to 15 are 0 to 7 with odd chips inverted: True
chips differing between any two sequences: at least 12, at most 20
chip errors in a symbol   symbol decoded wrongly
           2 of 32                   0.00%
           4 of 32                   0.00%
           6 of 32                   0.09%
           8 of 32                   2.65%
          10 of 32                  20.96%
SFD octet 0xA7; the SHR and PHR before a 20-octet PSDU, as symbols:
[0, 0, 0, 0, 0, 0, 0, 0, 7, 10, 4, 1]
PSDU  20 octets:  26 octets on air, 0.832 ms at 250 kb/s; header share 23.1%
PSDU  50 octets:  56 octets on air, 1.792 ms at 250 kb/s; header share 10.7%
PSDU 127 octets: 133 octets on air, 4.256 ms at 250 kb/s; header share 4.5%
munotes.in445

The 802.15.4 Physical Layer

The channels and Wi-Fi. The formulas give 868.3 MHz, 906 to 924 MHz and 2405 to 2480 MHz. Keeping only the 2 MHz channels that reach into none of the three 22 MHz Wi-Fi channels leaves exactly the standard's lists: 15, 20, 25 and 26 in North America, and 15, 16, 21 and 22 in Europe. A channel whose edge only touches a Wi-Fi channel's edge counts as clear here, as Figure E.1 draws it.

The chip table's structure. Both relations hold: symbols 1 to 7 are symbol 0 rotated by 4, 8, and so on up to 28 chips, and symbols 8 to 15 are symbols 0 to 7 with every odd-indexed chip inverted, the "conjugation" of the standard. A cheap radio needs to store one sequence and two simple operations, not sixteen unrelated patterns.

How far apart the sequences are. Any two of the sixteen differ in at least 12 and at most 20 of their 32 chips. That distance is what spreading buys. A received sequence with 5 or fewer wrong chips is still nearer to the sequence sent than to any other: it is within 5 of the right one and at least 12 - 5 = 7 from every wrong one.

munotes.in446

The 802.15.4 Physical Layer

Decoding through chip errors. With 2 or 4 wrong chips in 32, not one of 10,000 symbols was decoded wrongly, as the distance promises. With 6 the first mistakes appear (0.09 per cent), with 8 about 2.65 per cent of symbols are wrong, and with 10, about a third of the chips, about 21 per cent. The 4 bits of a symbol survive chip errors that would have ruined 4 bits sent plainly. A real receiver does better than this one: it correlates the analogue signal instead of first deciding each chip, and the CC2420 "does not do chip decision" at all.

The header as symbols. The SFD's bits make the octet 0xA7 (1 + 2 + 4 + 32 + 128 = 167). The whole header before a 20-octet PSDU is twelve symbols: eight 0s of preamble, then 7 and 10 (the SFD's low half first), then 4 and 1 (the length, 20, is 0x14). Every receiver knows the preamble in advance, which is how it locks on.

Time on the air. At 250 kb/s an octet lasts 32 µs. A PSDU of 20 octets goes out as 26 octets in 0.832 ms, and 23.1 per cent of that time is PHY header; the longest PSDU, 127 octets, takes 133 octets and 4.256 ms, the header then only 4.5 per cent. The MAC's own header, checksum and acknowledgement come on top ([CSMA-CA, Data Transfer and Frames in 802.15.4]).

What the radio must meet, and what a real radio does

The standard fixes minimums for the 2450 MHz radio.

  • Sensitivity: minus 85 dBm or better (6.5.3.3), where sensitivity is the "Threshold input signal power that yields a specified PER", measured with 20-octet PSDUs, a packet error rate below 1 per cent and no interference (Table 4).
  • Jamming resistance: adjacent channel rejection 0 dB and alternate channel rejection 30 dB (Table 26). "For example, when channel 13 is the desired channel, channel 12 and channel 14 are the adjacent channels, and channel 11 and channel 15 are the alternate channels."
  • Transmit power: at least minus 3 dBm. "Devices should transmit lower power when possible in order to reduce interference to other devices and systems. The maximum transmit power is limited by local regulatory bodies."
  • Frequency tolerance: within 40 ppm of the centre frequency (6.9.4); at 2480 MHz that is about 0.1 MHz. The symbol rate, 62.5 ksymbol/s, has the same 40 ppm tolerance.
  • Turnaround: switching from receiving to transmitting, or back, within aTurnaroundTime, 12 symbol periods, which is 12 × 16 = 192 µs. The RX-to-TX time is measured "from the trailing edge of the last chip (of the last symbol) of a received PPDU to the leading edge of the first chip (of the first symbol) of the next transmitted PPDU."
  • Largest input: a receiver must still meet the error rate with a desired signal as strong as minus 20 dBm (6.9.6).
  • Modulation accuracy: "A transmitter shall have EVM values of less than 35% when measured for 1000 chips."
munotes.in447

The 802.15.4 Physical Layer

A real radio clears these with room to spare. The CC2420 data sheet prints the standard's requirement beside its own figures:

IEEE 802.15.4-2006 requiresCC2420, typical
Sensitivity (PER 1 per cent)minus 85 dBmminus 95 dBm (minus 90 in its minimum column)
Largest input (saturation)minus 20 dBm10 dBm
Adjacent channel rejection0 dB45 dB at +5 MHz, 30 dB at minus 5 MHz
Alternate channel rejection30 dB54 dB at +10 MHz, 53 dB at minus 10 MHz
Output powerat least minus 3 dBm0 dBm nominal, set in 8 steps from about minus 24 dBm
EVMbelow 35 per cent11 per cent
Power measurement rangeat least 40 dB (ED)100 dB (RSSI)

Ten decibels of extra sensitivity is ten times weaker a signal still heard, which in practice means more range or fewer relays.

Energy detection, link quality and clear channel assessment

Three measurements pass from the PHY to the layers above, and each has a different job.

Energy detection. "The receiver ED measurement is intended for use by a network layer as part of a channel selection algorithm. It is an estimate of the received signal power within the bandwidth of the channel. No attempt is made to identify or decode signals on the channel." It averages over 8 symbol periods (128 µs) and reports an 8-bit value from 0x00 to 0xff. Zero means a power less than 10 dB above the sensitivity; the values must span at least 40 dB, mapped linearly to within 6 dB. A coordinator choosing a channel for a new network measures each channel this way and picks a quiet one.

Link quality indicator. "The LQI measurement is a characterization of the strength and/or quality of a received packet." It is taken for every received packet, is reported as 0x00 to 0xff, and "At least eight unique values of LQI shall be used." The standard does not say how to compute it (received energy, an estimate of the signal-to-noise ratio, or both) and "The use of the LQI result by the network or application layers is not specified in this standard."

The CC2420 shows why the method matters. "Using the RSSI value directly to calculate the LQI value has the disadvantage that e.g. a narrowband interferer inside the channel bandwidth will increase the LQI value although it actually reduces the true link quality. CC2420 therefore also provides an average correlation value for each incoming packet, based on the 8 first symbols following the SFD." Software then scales that correlation to the range 0 to 255 with constants found by measuring packet error rates. A strong signal is not necessarily a good link; a clean one is.

munotes.in448

The 802.15.4 Physical Layer

Routing is where LQI is spent, which is part of why MU files this standard under routing. Zigbee sums a link cost along each path, and "Even if some other method is used, the initial cost estimates shall be based on average LQI." Its specification adds that counting lost frames by their sequence numbers "is generally regarded as the most accurate measure of reception probability", the approach of the ETX trees in [Routing Tables and What Happens When the Topology Changes]. Zigbee's costs are worked in [Built on 802.15.4: Zigbee Routing, Security and the Later Amendments].

Clear channel assessment. The CSMA-CA of [CSMA/CA Worked Step by Step] asks the PHY one question before each transmission: is the channel clear? The PHY answers by at least one of three methods.

  • "CCA Mode 1: Energy above threshold. CCA shall report a busy medium upon detecting any energy above the ED threshold."
  • "CCA Mode 2: Carrier sense only. CCA shall report a busy medium only upon the detection of a signal compliant with this standard with the same modulation and spreading characteristics of the PHY that is currently in use by the device. This signal may be above or below the ED threshold."
  • "CCA Mode 3: Carrier sense with energy above threshold." The medium is busy by a logical combination of an 802.15.4 signal detected and energy above the threshold, "where the logical operator may be AND or OR."

The threshold must "correspond to a received signal power of at most 10 dB above the specified receiver sensitivity", and "The CCA detection time shall be equal to 8 symbol periods." While a packet is being received, CCA reports busy whatever the mode.

The modes differ in what they defer to. Mode 2 defers only to other 802.15.4 signals and would transmit over a Wi-Fi burst. Mode 1 defers to any energy, and Annex E prefers it for that reason: "Use of the ED option improves coexistence by allowing transmission backoff if the channel is occupied by any device, regardless of the communication protocol it may use."

Distinctions

At 2450 MHzBitSymbolChip
What it isOne binary digit of the frame4 bits, a number from 0 to 15One of the 32 elements of a symbol's PN sequence
Rate250 kb/s62.5 ksymbol/s2 Mchip/s
Duration4 µs16 µs0.5 µs
Program20-octet PSDU: 208 bits52 symbols1,664 chips
munotes.in449

The 802.15.4 Physical Layer

Energy detection (ED)Link quality indicator (LQI)Clear channel assessment (CCA)
MeasuresPower in a channel, signal or notStrength and/or quality of one received packetWhether the channel is busy now
WhenOn request, over 8 symbolsFor every received packetBefore transmitting, over 8 symbols
Result0x00 to 0xff0x00 to 0xff, at least 8 valuesBusy or idle
Used forChoosing a channelRoute and link choice (Zigbee's initial link cost)CSMA-CA
Mode 1Mode 2Mode 3
Busy whenAny energy above the ED thresholdAn 802.15.4 signal is detected, at any powerBoth (AND) or either (OR)
Defers to Wi-FiYesNoWith OR, yes
PreambleSFDPHRPSDU
Size at 2450 MHz4 octets (8 symbols, 128 µs)1 octet (2 symbols)1 octet: 7-bit length, 1 reserved0 to 127 octets
JobChip and symbol synchronisationMarks the end of the SHRSays how long the PSDU isCarries the MAC frame

What it does not mean

250 kb/s is not what an application gets. It is the rate on the air while a frame is being sent. PHY headers, MAC headers, backoffs, turnarounds and acknowledgements all take their share: the program's 20-octet PSDU spends 23.1 per cent of its airtime on the PHY header alone.

A channel number is not a frequency by itself. Channels 0 to 10 appear on three channel pages for three different PHYs; the page says which.

Spreading is not encryption. The sixteen sequences are printed in Table 24 for every radio to use. Confidentiality is the MAC's security, in [Built on 802.15.4: Zigbee Routing, Security and the Later Amendments].

LQI is not a standard unit. The standard fixes only its range and its minimum of eight values; two radios can give different LQI for the same link.

CCA is not collision detection. It looks before sending; it cannot hear a collision during a transmission. The MAC learns of a loss from a missing acknowledgement.

Sixteen channels are not sixteen free channels. Where Wi-Fi is busy, only four of them lie in its guard bands.

Quick revision

  • PHY duties: transceiver on and off, ED, LQI, CCA for CSMA-CA, channel selection, send and receive. Two services: data and management.
  • Four PHYs: 868/915 MHz BPSK (20 and 40 kb/s); optional 868/915 ASK (250 kb/s) and O-QPSK (100 and 250 kb/s); 2450 MHz O-QPSK (250 kb/s, worldwide).
  • Channels on page 0: 0 at 868.3 MHz; 1 to 10 at 906 + 2(k - 1); 11 to 26 at 2405 + 5(k - 11). Pages 1 and 2 reuse 0 to 10 for the ASK and O-QPSK PHYs.
  • Beside Wi-Fi: 15, 20, 25, 26 (North America), 15, 16, 21, 22 (Europe) lie in the 802.11b guard bands.
  • 2450 MHz: octet to 2 symbols (low 4 bits first); symbol to one of 16 PN sequences of 32 chips; half-sine O-QPSK, even chips on I, odd on Q, Q delayed by Tc. 62.5 ksymbol/s, 2 Mchip/s, 8 chips per bit, a symbol every 16 µs.
  • Program: sequences differ in 12 to 20 chips; up to 5 chip errors always corrected; 0.09 / 2.65 / 21 per cent symbol errors at 6 / 8 / 10 chip errors; SFD 0xA7.
  • PPDU: SHR (preamble, 4 zero octets; SFD, 1 octet) + PHR (7-bit length) + PSDU (at most 127 octets). Longest packet 133 octets, 4.256 ms.
  • Radio: sensitivity minus 85 dBm or better; adjacent 0 dB, alternate 30 dB; transmit at least minus 3 dBm; 40 ppm; turnaround 12 symbols (192 µs); EVM below 35 per cent. CC2420: minus 95 dBm typical.
  • ED: channel power, 8 symbols, 0x00 to 0xff, at least 40 dB. LQI: per packet, at least 8 values, use left to higher layers (Zigbee: initial link cost from average LQI). CCA modes 1 (energy), 2 (carrier sense), 3 (both, AND or OR); 8 symbols; threshold at most 10 dB above sensitivity.
munotes.in450

The 802.15.4 Physical Layer

Test yourself

1. List the responsibilities of the IEEE 802.15.4 PHY. The PHY activates and deactivates the radio transceiver, performs energy detection within the current channel, gives a link quality indicator for received packets, performs clear channel assessment for CSMA-CA, selects the channel frequency, and transmits and receives data. It offers these through a data service, which carries PPDUs across the radio channel, and a management service, through which the MAC requests measurements and settings.

2. Give the frequency bands and data rates of IEEE 802.15.4-2006. The 868 to 868.6 MHz band (Europe) and the 902 to 928 MHz band (North America) are served by a mandatory BPSK PHY at 20 kb/s and 40 kb/s respectively, and by two optional PHYs: ASK with parallel sequence spread spectrum at 250 kb/s in both bands, and O-QPSK at 100 kb/s at 868 MHz and 250 kb/s at 915 MHz. The 2400 to 2483.5 MHz band, available worldwide, has an O-QPSK PHY at 250 kb/s, 62.5 ksymbol/s and 2 Mchip/s.

3. Compute the centre frequencies of channels 0, 5, 11 and 20. Channel 0 is at 868.3 MHz. Channels 1 to 10 are at 906 + 2(k - 1) MHz, so channel 5 is at 906 + 2 × 4 = 914 MHz. Channels 11 to 26 are at 2405 + 5(k - 11) MHz, so channel 11 is at 2405 MHz and channel 20 at 2405 + 5 × 9 = 2450 MHz.

munotes.in451

The 802.15.4 Physical Layer

4. Explain how the 2450 MHz PHY transmits an octet. The octet is split into two 4-bit data symbols, the 4 least significant bits first. Each symbol, a value from 0 to 15, is replaced by its 32-chip pseudo-noise sequence from the standard's table; the sixteen sequences are cyclic shifts of one another or those shifts with the odd-indexed chips inverted. The chips are sent by O-QPSK with half-sine pulses: even-indexed chips on the in-phase carrier and odd-indexed chips on the quadrature carrier, the quadrature chips delayed by one chip period. The symbol rate is 62.5 ksymbol/s and the chip rate 2 Mchip/s, so an octet takes 32 µs.

5. Describe the PPDU format of 802.15.4 and explain the 127-octet limit. A PPDU begins with the synchronisation header: a preamble (at 2450 MHz, 4 octets of zeros, 8 symbols) that lets the receiver obtain chip and symbol synchronisation, then the 1-octet start-of-frame delimiter marking the end of the header. Next comes the PHY header, whose 7-bit frame length field gives the PSDU length, with one reserved bit. Last is the PSDU, which carries the MAC frame. Because the length field has 7 bits, the PSDU can be at most 127 octets, the constant aMaxPHYPacketSize; the MAC's header and checksum and any security must fit inside it, which is why IPv6 over 802.15.4 needs fragmentation.

6. Distinguish energy detection, link quality indication and clear channel assessment. Energy detection estimates the received power within a channel, averaged over 8 symbol periods, without trying to decode it; network layers use it to choose a channel. The link quality indicator characterises the strength and/or quality of each received packet on a scale of 0x00 to 0xff with at least eight values; its use is left to higher layers, for example Zigbee's link costs. Clear channel assessment tells the MAC, before it transmits, whether the channel is busy, by energy above a threshold (mode 1), by detecting an 802.15.4 signal (mode 2), or by combining the two with AND or OR (mode 3).

7. Why can CCA mode 1 coexist with Wi-Fi better than mode 2? Mode 2 reports the channel busy only when it detects a signal with 802.15.4's own modulation and spreading, so it would start transmitting over a Wi-Fi transmission and both would suffer. Mode 1 reports busy whenever the energy exceeds the threshold, whatever sent it, so the node backs off while Wi-Fi is on the air. The standard's Annex E notes that energy detection improves coexistence for this reason.

munotes.in452

The 802.15.4 Physical Layer

8. In the program, why were no symbols decoded wrongly with 4 chip errors, while some were with 6? Any two of the sixteen chip sequences differ in at least 12 chips. With 4 errors, the received sequence is 4 chips from the one sent and at least 12 - 4 = 8 from any other, so the nearest sequence is always the right one; the same holds up to 5 errors. With 6 errors the received sequence can be as close to, or closer to, another sequence that differs from the sent one in exactly 12 chips, so occasional mistakes appear: 0.09 per cent of symbols.

Contents This chapter on its own page

munotes.in453

Chapter Sixty-Five

The 802.15.4 Superframe and Guaranteed Time Slots

Syllabus topic Module 1, "Routing in WSN: IEEE 802.15.4 LR-WPAN standard (case study)"

In one line

In a beacon-enabled 802.15.4 PAN the coordinator's beacons divide time into superframes: a beacon every aBaseSuperframeDuration × 2^BO symbols, an active portion of aBaseSuperframeDuration × 2^SO symbols cut into 16 slots (the beacon, a contention access period for slotted CSMA-CA and a contention-free period of up to seven guaranteed time slots), and an inactive portion in which everyone may sleep.

In the wording a student can write in an examination: IEEE 802.15.4 lets a PAN run beacon-enabled, with a superframe structure, or nonbeacon-enabled (beacon order and superframe order both 15), where devices use unslotted CSMA-CA and there are no GTSs. The superframe is bounded by beacons from the coordinator, which synchronise the devices, identify the PAN and describe the superframe. Its active portion is divided into 16 equal slots and contains the beacon (sent in slot 0 without CSMA), the contention access period (CAP), in which devices use slotted CSMA-CA, and an optional contention-free period (CFP) made of guaranteed time slots (GTSs); an optional inactive portion follows, in which the coordinator and devices may enter a low-power mode. With aBaseSuperframeDuration = 960 symbols (60 symbols per slot × 16 slots), the beacon interval is BI = 960 × 2^BO symbols and the superframe duration is SD = 960 × 2^SO symbols, for SO from 0 up to BO, and BO up to 14; at 2450 MHz a symbol lasts 16 µs, so BI runs from 15.36 ms to about 252 s, and the fraction of time active is 2^(SO - BO).

The CAP must last at least aMinCAPLength = 440 symbols, and MAC commands always go in it. GTSs are allocated only by the PAN coordinator, first come, first served, at most seven at a time, placed contiguously at the end of the active portion; each is a number of slots, in a transmit or receive direction relative to the device, used without CSMA-CA and with short addresses only. A device asks with a GTS request command (length, direction, allocation or deallocation) and learns the answer from a GTS descriptor (short address, starting slot, length) in the beacon. Gaps left by released GTSs are closed by moving GTSs toward the end, and a GTS unused for 2n superframes (n = 2^(8 - BO) for BO up to 8, otherwise 1) is taken back.

Two ways to run a PAN

The superframe is optional, and a PAN chooses when it starts. "PANs that wish to use the superframe structure (referred to as a beacon-enabled PAN) shall set macBeaconOrder to a value between 0 and 14, both inclusive, and macSuperframeOrder to a value between 0 and the value of macBeaconOrder, both inclusive."

munotes.in454

The 802.15.4 Superframe and Guaranteed Time Slots

"PANs that do not wish to use the superframe structure (referred to as a nonbeacon-enabled PAN) shall set both macBeaconOrder and macSuperframeOrder to 15." In such a PAN the coordinator sends a beacon only when a device asks for one with a beacon request command, frames go by unslotted CSMA-CA, and "In addition, GTSs shall not be permitted."

The choice is a trade. A beacon-enabled PAN gives every device a common clock, so all can sleep together in the inactive portion and wake together for the beacon, and it can guarantee slots. A nonbeacon-enabled PAN is simpler: "When a device wishes to transfer data in a nonbeacon-enabled PAN, it simply transmits its data frame, using unslotted CSMA-CA, to the coordinator." No synchronisation is needed, but a coordinator that may be sent a frame at any moment has to be listening at every moment. That fits the standard's expectation for a star: "The PAN coordinator might often be mains powered, while the devices will most likely be battery powered." [CSMA-CA, Data Transfer and Frames in 802.15.4] follows data through both kinds of PAN.

The superframe

"The superframe is bounded by network beacons sent by the coordinator (see Figure 4a) and is divided into 16 equally sized slots. Optionally, the superframe can have an active and an inactive portion (see Figure 4b). During the inactive portion, the coordinator may enter a low-power mode." The beacons do three jobs: they synchronise the attached devices, identify the PAN, and "describe the structure of the superframes."

Top: one superframe drawn as sixteen slots, slot 0 holding the beacon, slots 1 to 10 the contention access period, slots 11 to 15 a contention-free period of two guaranteed time slots, then an inactive portion as long again in which all may sleep, then the next beacon; brackets mark the active portion SD, 960 × 2 to the power SO symbols, and the beacon interval BI, 960 × 2 to the power BO symbols. Bottom: the standard's Figure 73 in three rows of sixteen slots: GTSs A, B and C at slots 14, 10 and 8; B given back, leaving four empty slots; C moved up to slots 12 and 13 so the contention access period grows

Figure 65.1 The superframe, after the standard's Fig. 66, and its Fig. 73: a GTS given back and the gap closed

The active portion "is composed of three parts: a beacon, a CAP and a CFP. The beacon shall be transmitted, without the use of CSMA, at the start of slot 0, and the CAP shall commence immediately following the beacon." The CFP, if there is one, runs from the end of the CAP to the end of the active portion, and "Any allocated GTSs shall be located within the CFP."

Two orders, two durations

Two MAC attributes set the timing, each as a power of two.

  • Beacon order (BO), macBeaconOrder, sets how often beacons come: the beacon interval is BI = aBaseSuperframeDuration × 2^BO symbols, for BO from 0 to 14. BO = 15 means no periodic beacons.
  • Superframe order (SO), macSuperframeOrder, sets how long the active portion lasts, beacon included: the superframe duration is SD = aBaseSuperframeDuration × 2^SO symbols, for SO from 0 up to BO.
  • Each of the 16 slots lasts aBaseSlotDuration × 2^SO symbols.

The constants come from Table 85: aBaseSlotDuration is 60 symbols, aNumSuperframeSlots is 16, so aBaseSuperframeDuration is 60 × 16 = 960 symbols. At 2450 MHz a symbol is 16 µs, so the shortest superframe lasts 960 × 16 = 15,360 µs, which is 15.36 ms.

munotes.in455

The 802.15.4 Superframe and Guaranteed Time Slots

Worked. Take BO = 6 and SO = 3. The beacon interval is 960 × 64 = 61,440 symbols, which is 983.04 ms. The active portion is 960 × 8 = 7,680 symbols, 122.88 ms, and each slot 480 symbols, 7.68 ms. The PAN is active for 7,680 of every 61,440 symbols: one eighth, 12.5 per cent, and asleep for the rest. In general the active fraction is 2^SO divided by 2^BO, that is 2^(SO - BO): each step apart halves the time awake.

"The beacon order and superframe order shall be equal for all superframes on a PAN. All devices shall interact with the PAN only during the active portion of a superframe."

The contention access period

"The CAP shall start immediately following the beacon and complete before the beginning of the CFP on a superframe slot boundary." It shrinks and grows as GTSs come and go, but never below a floor: it must be at least aMinCAPLength, 440 symbols, which "ensures that MAC commands can still be transferred to devices when GTSs are being used." (The one exception is a beacon made temporarily longer by GTS descriptors.)

Frames in the CAP, apart from acknowledgements and a data frame that quickly follows the acknowledgement of a data request, use slotted CSMA-CA, whose backoff periods line up with the slots ([CSMA-CA, Data Transfer and Frames in 802.15.4]). A device must finish its whole transaction, acknowledgement included, one interframe spacing (IFS) before the CAP ends; if it cannot, it waits for the next superframe's CAP. And "MAC command frames shall always be transmitted in the CAP."

The contention-free period and guaranteed time slots

"For low-latency applications or applications requiring specific data bandwidth, the PAN coordinator may dedicate portions of the active superframe to that application. These portions are called guaranteed time slots (GTSs)." They form the CFP, and "No transmissions within the CFP shall use a CSMA-CA mechanism to access the channel."

The rules, from clause 7.5.7:

  • Who allocates. "A GTS shall be allocated only by the PAN coordinator, and it shall be used only for communications between the PAN coordinator and a device associated with the PAN through the PAN coordinator."
  • How many. "The PAN coordinator may allocate up to seven GTSs at the same time, provided there is sufficient capacity in the superframe." A GTS may span several slots.
  • In what order and where. "GTSs shall be allocated on a first-come-first-served basis, and all GTSs shall be placed contiguously at the end of the superframe and after the CAP."
  • Direction. "The GTS direction, which is relative to the data flow from the device that owns the GTS, is specified as either transmit or receive." "Each device may request one transmit GTS and/or one receive GTS." For a receive GTS the device keeps its receiver on for the whole GTS; for a transmit GTS the PAN coordinator does.
  • Addressing. "A data frame transmitted in an allocated GTS shall use only short addressing."
  • Synchronisation. A device may ask for and use a GTS only while it tracks the beacons, and "If a device loses synchronization with the PAN coordinator, all its GTS allocations shall be lost."
  • Fitting. Before transmitting in a GTS, a device must be sure that the frame, its acknowledgement if one is requested, and the IFS suited to the frame's size all end before the GTS does; otherwise it waits for its GTS in the next superframe.
munotes.in456

The 802.15.4 Superframe and Guaranteed Time Slots

A device with a GTS may still use the CAP. A GTS is a reservation, not a prison.

Asking for a GTS, and the answer in the beacon

A device asks with a GTS request command, sent in the CAP to the PAN coordinator and acknowledged. Its GTS Characteristics field holds the GTS length (4 bits: the number of slots), the GTS direction (1 bit: one for receive, zero for transmit) and the characteristics type (1 bit: one to allocate, zero to deallocate).

The coordinator then checks capacity. "The superframe shall have available capacity if the maximum number of GTSs has not been reached and allocating a GTS of the desired length would not reduce the length of the CAP to less than aMinCAPLength."

The answer comes in the beacon, as a GTS descriptor: the device's 16-bit short address, the 4-bit starting slot and the 4-bit length. A descriptor stays in the beacon for aGTSDescPersistenceTime, 4 superframes, and a device that sees none for itself within that time reports failure. A refusal is a descriptor with starting slot zero and, as its length, the largest GTS the coordinator could give now.

The beacon carries the rest of the superframe's description in two fields:

  • the Superframe Specification field (16 bits): beacon order (4 bits), superframe order (4 bits), Final CAP Slot (4 bits, the last slot the CAP uses), Battery Life Extension (1 bit), a reserved bit, PAN Coordinator (1 bit) and Association Permit (1 bit, whether the coordinator accepts devices joining);
  • the GTS fields: a GTS Specification octet (a 3-bit count of descriptors and a GTS Permit bit, whether requests are accepted), a GTS Directions mask (7 bits, one per descriptor, one for receive) and the list of 3-octet descriptors.
munotes.in457

The 802.15.4 Superframe and Guaranteed Time Slots

Giving a GTS back, closing the gap, and expiry

Deallocation. A device gives a GTS back with the same command, characteristics type zero; the coordinator may also take one back, because the higher layer asks, because the GTS has expired, or to keep the CAP at its minimum.

Closing the gap. Releasing a GTS in the middle leaves a hole. The standard's own example: "In stage 1, three GTSs are allocated starting at slots 14, 10, and 8, respectively. If GTS 2 is now deallocated (stage 2), there will be a gap in the superframe during which nothing can happen. To solve this, GTS 3 will have to be shifted to fill the gap, thus increasing the size of the CAP (stage 3)." The coordinator gives each GTS nearer the CAP a new starting slot, announces it in a descriptor, and the device "shall adjust the starting slot of the GTS corresponding to the GTS descriptor and start using it immediately." The figure's lower half draws the three stages; the program replays them.

Expiry. A device may stop using its GTS without saying so. For a transmit GTS, the coordinator assumes the device has finished if no data frame arrives from it in the GTS at least once every 2n superframes; for a receive GTS, if no acknowledgement comes back from it in that time. Here n is 2^(8 - BO) for a beacon order from 0 to 8, and 1 for a beacon order from 9 to 14. The program shows what this choice achieves.

Missing a beacon. "If a device misses the beacon at the beginning of a superframe, it shall not use its GTSs until it receives a subsequent beacon correctly." After aMaxLostBeacons, 4 lost in a row, the device declares a loss of synchronisation, and with it loses its GTSs.

Superframes in a cluster tree

In a cluster tree ([IEEE 802.15.4: The Standard, Its Devices and Its Topologies]) every coordinator beacons, so superframes must share time. "On a beacon-enabled PAN, a coordinator that is not the PAN coordinator shall maintain the timing of both the superframe in which its coordinator transmits a beacon (the incoming superframe) and the superframe in which it transmits its own beacon (the outgoing superframe)." Its own beacon is offset from its parent's by a start time, so that its active portion falls in its parent's inactive portion.

This is where BO - SO pays twice. Since all superframes on a PAN share BO and SO, one beacon interval has room for 2^(BO - SO) active portions end to end: with BO = 6 and SO = 3, eight. Coordinators whose active portions would otherwise overlap can be given different offsets, so the gap between the orders sets both how long a device sleeps and how many coordinators can take turns.

munotes.in458

The 802.15.4 Superframe and Guaranteed Time Slots

The superframe, computed

The program tabulates the timing for pairs of orders; replays Figure 73 at SO = 0, the smallest slots, with the capacity rule enforced; asks how many of the 16 slots could be GTSs at each superframe order; counts the acknowledged frames of three sizes that fit in a one-slot GTS; and computes the GTS expiry time for several beacon orders.

# The 802.15.4 superframe, worked at 2450 MHz, where a symbol lasts 16 us. The
# constants are the standard's Table 85; the orders BO and SO run from 0 to 14.
aBaseSlotDuration, aNumSuperframeSlots, aMinCAPLength = 60, 16, 440
aBaseSuperframeDuration = aBaseSlotDuration * aNumSuperframeSlots
ms = lambda symbols: symbols * 16 / 1000

print("aBaseSuperframeDuration: %d symbols, %.2f ms" % (aBaseSuperframeDuration, ms(aBaseSuperframeDuration)))
print(" BO  SO  beacon interval  active portion      slot   awake")
for bo, so in ((0, 0), (3, 3), (6, 3), (6, 0), (8, 4), (10, 2), (14, 0)):
    bi, sd = aBaseSuperframeDuration * 2 ** bo, aBaseSuperframeDuration * 2 ** so
    print("%3d %3d %13.2f ms %12.2f ms %7.2f ms %6.2f%%"
          % (bo, so, ms(bi), ms(sd), ms(sd / aNumSuperframeSlots), 100 * sd / bi))

# A frame on the air, in symbols: 2 per octet, with 6 octets of PHY header. The
# smallest beacon (short addresses, no security, no GTS descriptors, no pending
# addresses, no payload) is a 13-octet MAC frame (Figure 44).
ppdu = lambda mpdu: 2 * (6 + mpdu)
BEACON = ppdu(13)

class PANCoordinator:
    def __init__(self, so):
        self.slot = aBaseSlotDuration * 2 ** so
        self.gts = []                                   # [device, start slot, length]
    def final_cap_slot(self):
        return min([start for _, start, _ in self.gts], default=aNumSuperframeSlots) - 1
    def cap_after(self, first_gts_slot):                # CAP runs from the beacon's end
        return first_gts_slot * self.slot - BEACON
    def request(self, device, length):
        start = self.final_cap_slot() + 1 - length
        if len(self.gts) == 7 or self.cap_after(start) < aMinCAPLength:
            largest = max([n for n in range(16) if self.cap_after(self.final_cap_slot() + 1 - n)
                           >= aMinCAPLength], default=0)
            print("  %s asks for %d slot(s): DENIED, the CAP would be %d symbols; largest now %d"
                  % (device, length, self.cap_after(start), largest))
            return
        self.gts.append([device, start, length])
        print("  %s asks for %d slot(s): starts at slot %2d  %s" % (device, length, start, self.map()))
    def release(self, device):
        i = [d for d, _, _ in self.gts].index(device)
        _, start, length = self.gts.pop(i)
        for g in self.gts:                              # 7.5.7.5: close the gap
            if g[1] < start:
                g[1] += length
        print("  %s gives its GTS back; the CFP closes up    %s" % (device, self.map()))
    def map(self):
        cells = ["b"] + ["c"] * 15
        for device, start, length in self.gts:
            cells[start:start + length] = [device] * length
        return "".join(cells) + "  CAP %d symbols" % self.cap_after(self.final_cap_slot() + 1)

print("\nThe standard's Figure 73, at SO = 0 (b beacon, c CAP, letters GTSs):")
pan = PANCoordinator(0)
for device, length in (("A", 2), ("B", 4), ("C", 2), ("D", 1)):
    pan.request(device, length)
pan.release("B")
pan.request("D", 1)

print("\nSlots that can be GTSs at most, keeping the CAP at %d symbols:" % aMinCAPLength)
for so in range(5):
    pan = PANCoordinator(so)
    print("  SO = %d, slot %4d symbols: %2d of 16" % (so, pan.slot,
          max(n for n in range(16) if pan.cap_after(16 - n) >= aMinCAPLength)))

# What one GTS slot carries: each acknowledged frame costs its PPDU,
# aTurnaroundTime (12 symbols), the 5-octet acknowledgement, and then an IFS
# suited to the frame: SIFS (12) for an MPDU of up to 18 octets, else LIFS (40).
def transaction(mpdu):
    return ppdu(mpdu) + 12 + ppdu(5) + (12 if mpdu <= 18 else 40)

print("\nAcknowledged frames that fit in a one-slot GTS:")
print("  SO   slot    18-octet   50-octet  127-octet   (symbols per frame: %d, %d, %d)"
      % (transaction(18), transaction(50), transaction(127)))
for so in (0, 1, 3, 6):
    slot = aBaseSlotDuration * 2 ** so
    print("  %2d %6d %10d %10d %10d" % ((so, slot) + tuple(slot // transaction(m) for m in (18, 50, 127))))

# 7.5.7.6: a GTS not used for 2n superframes is taken back, n = 2^(8 - BO) up to BO = 8
print("\nGTS expiry: 2n superframes")
for bo in (0, 4, 8, 9, 12, 14):
    n = 2 ** (8 - bo) if bo <= 8 else 1
    print("  BO = %2d: n = %3d, after %3d superframes = %7.2f s"
          % (bo, n, 2 * n, ms(2 * n * aBaseSuperframeDuration * 2 ** bo) / 1000))
munotes.in459

The 802.15.4 Superframe and Guaranteed Time Slots

aBaseSuperframeDuration: 960 symbols, 15.36 ms
 BO  SO  beacon interval  active portion      slot   awake
  0   0         15.36 ms        15.36 ms    0.96 ms 100.00%
  3   3        122.88 ms       122.88 ms    7.68 ms 100.00%
  6   3        983.04 ms       122.88 ms    7.68 ms  12.50%
  6   0        983.04 ms        15.36 ms    0.96 ms   1.56%
  8   4       3932.16 ms       245.76 ms   15.36 ms   6.25%
 10   2      15728.64 ms        61.44 ms    3.84 ms   0.39%
 14   0     251658.24 ms        15.36 ms    0.96 ms   0.01%

The standard's Figure 73, at SO = 0 (b beacon, c CAP, letters GTSs):
  A asks for 2 slot(s): starts at slot 14  bcccccccccccccAA  CAP 802 symbols
  B asks for 4 slot(s): starts at slot 10  bcccccccccBBBBAA  CAP 562 symbols
  C asks for 2 slot(s): starts at slot  8  bcccccccCCBBBBAA  CAP 442 symbols
  D asks for 1 slot(s): DENIED, the CAP would be 382 symbols; largest now 0
  B gives its GTS back; the CFP closes up    bcccccccccccCCAA  CAP 682 symbols
  D asks for 1 slot(s): starts at slot 11  bccccccccccDCCAA  CAP 622 symbols

Slots that can be GTSs at most, keeping the CAP at 440 symbols:
  SO = 0, slot   60 symbols:  8 of 16
  SO = 1, slot  120 symbols: 12 of 16
  SO = 2, slot  240 symbols: 14 of 16
  SO = 3, slot  480 symbols: 15 of 16
  SO = 4, slot  960 symbols: 15 of 16

Acknowledged frames that fit in a one-slot GTS:
  SO   slot    18-octet   50-octet  127-octet   (symbols per frame: 94, 186, 340)
   0     60          0          0          0
   1    120          1          0          0
   3    480          5          2          1
   6   3840         40         20         11

GTS expiry: 2n superframes
  BO =  0: n = 256, after 512 superframes =    7.86 s
  BO =  4: n =  16, after  32 superframes =    7.86 s
  BO =  8: n =   1, after   2 superframes =    7.86 s
  BO =  9: n =   1, after   2 superframes =   15.73 s
  BO = 12: n =   1, after   2 superframes =  125.83 s
  BO = 14: n =   1, after   2 superframes =  503.32 s
munotes.in460

The 802.15.4 Superframe and Guaranteed Time Slots

The timing. The base superframe is 15.36 ms. Equal orders keep the PAN awake all the time (0 and 0, 3 and 3); BO = 6 with SO = 3 is the worked example, 983.04 ms between beacons and 12.5 per cent awake; BO = 6 with SO = 0 is awake 1.56 per cent; BO = 14 with SO = 0 sends a beacon every 251.66 s and is awake 0.01 per cent. A coordinator chooses between low latency (a short BI) and low energy (a large BO - SO).

Figure 73, replayed. At SO = 0 each slot is only 60 symbols. A, B and C receive 2, 4 and 2 slots, starting at 14, 10 and 8, exactly the standard's stage 1, and the CAP falls from 802 to 442 symbols, just above the 440 minimum. D's one-slot request is therefore denied: the CAP would drop to 382 symbols, and the largest GTS the coordinator could give is 0 slots. When B gives its 4 slots back, C moves from slots 8 and 9 to 12 and 13 (stage 3), the CAP grows to 682 symbols, and D's request now succeeds at slot 11.

How much can be guaranteed. Keeping 440 symbols of CAP after the beacon, at SO = 0 only 8 of the 16 slots can be GTSs, at SO = 1 twelve, at SO = 2 fourteen, and from SO = 3 fifteen: slot 0 alone then holds the beacon and a full minimum CAP.

What a slot carries. An acknowledged frame costs its airtime, a 12-symbol turnaround, the acknowledgement and an IFS: 94 symbols for an 18-octet MAC frame (with the short IFS), 186 for 50 octets and 340 for 127 (with the long one). A one-slot GTS at SO = 0 (60 symbols) cannot carry even the smallest; at SO = 1 it carries one 18-octet frame; at SO = 3 five small frames, two of 50 octets or one of 127; at SO = 6 (3,840 symbols) eleven full-size frames. At small superframe orders a useful GTS must be several slots long.

munotes.in461

The 802.15.4 Superframe and Guaranteed Time Slots

Expiry is a fixed time. For BO from 0 to 8, n = 2^(8 - BO) and a superframe lasts 960 × 2^BO symbols, so 2n superframes last 2 × 2^(8 - BO) × 960 × 2^BO = 2 × 256 × 960 = 491,520 symbols, about 7.86 s, whatever the beacon order. The formula is built so that a silent device loses its GTS after the same time at any of those orders. From BO = 9, n = 1 and the time grows with the beacon interval: two beacon intervals, 15.73 s at BO = 9 and 503.32 s at BO = 14.

Distinctions

Beacon-enabled PANNonbeacon-enabled PAN
OrdersBO from 0 to 14, SO from 0 to BOBO = SO = 15
BeaconsPeriodic, in slot 0 of every superframeOnly when a device asks
Channel accessSlotted CSMA-CA in the CAP; GTSs in the CFPUnslotted CSMA-CA
GTSsUp to sevenNot permitted
SleepingEveryone in the inactive portionDevices sleep; the coordinator listens
Contention access period (CAP)Contention-free period (CFP)
PlaceFrom the end of the beaconAt the end of the active portion, after the CAP
AccessSlotted CSMA-CAGTSs, no CSMA-CA
WhoAny device; all MAC commandsThe owner of each GTS and the PAN coordinator
SizeAt least aMinCAPLength (440 symbols)Whatever the GTSs need, within that limit
Beacon order (BO)Superframe order (SO)
SetsThe beacon interval, BI = 960 × 2^BO symbolsThe active portion, SD = 960 × 2^SO symbols
Range0 to 14; 15 means no beacons0 to BO; 15 means inactive after the beacon
Worked (BO 6, SO 3)983.04 ms122.88 ms, slots of 7.68 ms
Transmit GTSReceive GTS
Data flowsFrom the device to the PAN coordinatorFrom the PAN coordinator to the device
Receiver on for the whole GTSThe PAN coordinator'sThe device's
Expiry if missingData frames from the deviceAcknowledgements from the device

What it does not mean

A GTS is not a slot of the whole network. GTSs exist only between the PAN coordinator and its devices; two ordinary devices cannot share one.

Sixteen slots are not sixteen GTSs. At most seven GTSs exist at once, a GTS may be several slots, and the CAP must keep at least 440 symbols; the program found as few as 8 usable slots at SO = 0.

munotes.in462

The 802.15.4 Superframe and Guaranteed Time Slots

Beacon-enabled is not always better. The superframe saves energy only when everyone's traffic fits the active portion; a device with data just after the active portion ends waits a whole inactive portion.

The inactive portion is not wasted time. It is the point: the radios are off, and in a cluster tree other coordinators can use it.

Guaranteed does not mean delivered. A GTS removes contention, not noise; a frame in a GTS can still be lost and is retried like any other.

Quick revision

  • Beacon-enabled (BO 0 to 14, SO 0 to BO): superframes, slotted CSMA-CA, GTSs. Nonbeacon-enabled (BO = SO = 15): unslotted CSMA-CA, no GTSs.
  • Superframe: 16 slots; beacon in slot 0 without CSMA; CAP (slotted CSMA-CA, MAC commands, at least 440 symbols); CFP of GTSs; optional inactive portion for sleep.
  • BI = 960 × 2^BO, SD = 960 × 2^SO, slot = 60 × 2^SO symbols; at 2450 MHz a symbol is 16 µs, so the base superframe is 15.36 ms; active fraction 2^(SO - BO). Worked: BO 6, SO 3: 983.04 ms, 122.88 ms, 12.5 per cent.
  • GTSs: PAN coordinator only, first come first served, up to 7, contiguous at the end, transmit or receive, one of each per device, short addresses, no CSMA-CA, lost with synchronisation.
  • Request: GTS request command (length 4 bits, direction, type); answer: descriptor (short address, start slot, length) in the beacon for 4 superframes; start slot 0 = denied.
  • Beacon: Superframe Specification (BO, SO, Final CAP Slot, BLE, PAN Coordinator, Association Permit); GTS Specification (count, GTS Permit), directions mask, descriptors.
  • Release: the CFP is defragmented (Figure 73: 14, 10, 8; release 10; 8 moves to 12). Expiry after 2n superframes, n = 2^(8 - BO) up to BO 8, else 1: a fixed 7.86 s up to BO 8.
  • Cluster tree: each coordinator keeps an incoming and an outgoing superframe; 2^(BO - SO) active portions fit in one beacon interval.

Test yourself

1. Distinguish a beacon-enabled and a nonbeacon-enabled 802.15.4 PAN. In a beacon-enabled PAN the coordinator sets the beacon order between 0 and 14 and the superframe order between 0 and the beacon order, and sends beacons that bound superframes; devices use slotted CSMA-CA in the contention access period, may be given guaranteed time slots, and can all sleep in the inactive portion. In a nonbeacon-enabled PAN both orders are 15, the coordinator sends beacons only on request, devices use unslotted CSMA-CA, and GTSs are not permitted.

munotes.in463

The 802.15.4 Superframe and Guaranteed Time Slots

2. Describe the structure of the 802.15.4 superframe. The superframe begins with a beacon from the coordinator and runs to the next beacon. Its active portion is divided into 16 equal slots: the beacon is sent at the start of slot 0 without CSMA; the contention access period follows immediately, in which devices compete with slotted CSMA-CA and all MAC commands are sent; the contention-free period, if any, occupies the last slots and consists of guaranteed time slots. An optional inactive portion follows the active portion, during which the coordinator and devices may enter a low-power mode.

3. With BO = 8 and SO = 4 at 2450 MHz, find the beacon interval, the superframe duration, the slot length and the fraction of time active. aBaseSuperframeDuration is 960 symbols and a symbol lasts 16 µs. The beacon interval is 960 × 256 = 245,760 symbols, which is 3,932.16 ms. The superframe duration is 960 × 16 = 15,360 symbols, 245.76 ms, and each of the 16 slots lasts 960 symbols, 15.36 ms. The active fraction is 2^(4 - 8), one sixteenth, 6.25 per cent.

4. What is a guaranteed time slot, and what rules govern its allocation? A GTS is a portion of the active superframe, one or more slots, dedicated to one device for communication with the PAN coordinator without contention, for low-latency or fixed-bandwidth traffic. Only the PAN coordinator allocates GTSs, first come first served, at most seven at a time, contiguously at the end of the active portion in the contention-free period; each is transmit or receive relative to the device, a device may hold one of each, and frames in it use short addresses and no CSMA-CA. An allocation is refused if seven exist already or if it would shrink the contention access period below aMinCAPLength, 440 symbols.

5. How does a device obtain a GTS and learn the result? It sends a GTS request command to the PAN coordinator in the contention access period, with the GTS length in slots, the direction and the characteristics type set to allocation. The coordinator acknowledges the command, checks the number of GTSs and the remaining CAP length, and puts a GTS descriptor (the device's short address, the starting slot and the length) in its beacon for four superframes. A starting slot greater than zero means success; zero means refusal, with the length set to the largest GTS now available.

6. Explain GTS deallocation and the reallocation that may follow. A device can release a GTS with a GTS request command whose characteristics type is zero, and the coordinator can take one back when a higher layer asks, when the GTS expires or to keep the CAP at its minimum. Releasing a GTS in the middle of the contention-free period leaves a gap in which nothing can happen, so the coordinator moves each GTS nearer the CAP toward the end to close it and announces the new starting slots in beacon descriptors; the CAP grows accordingly. In the standard's example, GTSs start at slots 14, 10 and 8; when the one at 10 is released, the one at 8 moves to 12.

munotes.in464

The 802.15.4 Superframe and Guaranteed Time Slots

7. Show that GTS expiry takes the same time for every beacon order from 0 to 8. The coordinator takes back a GTS unused for 2n superframes, with n equal to 2^(8 - BO) for BO up to 8. A superframe (a beacon interval) lasts 960 × 2^BO symbols, so 2n superframes last 2 × 2^(8 - BO) × 960 × 2^BO = 2 × 256 × 960 = 491,520 symbols, which at 16 µs per symbol is about 7.86 s, independent of BO. For BO from 9 to 14, n is 1 and the time is two beacon intervals.

8. In the program, why was D's one-slot request denied at first but granted after B released its GTS? At superframe order 0 each slot is 60 symbols. With GTSs at slots 8 to 15, the CAP ran from the end of the 38-symbol beacon to the end of slot 7: 8 × 60 - 38 = 442 symbols, just above the 440 minimum. A one-slot GTS for D would have ended the CAP at slot 6, leaving 382 symbols, so the coordinator refused. When B released four slots, C moved up and the CAP grew to 682 symbols, so taking one more slot for D left 622, above the minimum, and D was given slot 11.

Contents This chapter on its own page

munotes.in465

Chapter Sixty-Six

CSMA-CA, Data Transfer and Frames in 802.15.4

Syllabus topic Module 1, "Routing in WSN: IEEE 802.15.4 LR-WPAN standard (case study)"

In one line

802.15.4's MAC sends every frame in the CAP after CSMA-CA (slotted, on backoff boundaries with two clear assessments, when there are beacons; unslotted, with one, when there are none), confirms it with an acknowledgement sent a fixed turnaround later, retries up to three times, carries data to, from or between devices by three transfer models (data from a coordinator waiting until the device asks for it), and packs everything into four frame types sharing one format: frame control, sequence number, addresses, payload and a 16-bit CRC.

In the wording a student can write in an examination: IEEE 802.15.4 uses CSMA-CA for data and command frames in the contention access period, but not for beacons, acknowledgements or frames in guaranteed time slots. Each attempt keeps three variables: NB, the number of backoffs so far (starts at 0); CW, the contention window, the number of backoff periods that must be clear before sending (starts at 2, slotted only); and BE, the backoff exponent (starts at macMinBE, 3 by default, or the lesser of 2 and macMinBE under battery life extension). In slotted CSMA-CA, used in beacon-enabled PANs, a device aligns to the backoff period boundaries (20 symbols each, counted from the start of the beacon), waits a random 0 to 2^BE - 1 backoff periods, and performs a clear channel assessment (CCA) on a boundary. If the channel is busy, it sets CW back to 2, increases NB and BE (BE at most macMaxBE, 5), and either backs off again or, once NB exceeds macMaxCSMABackoffs (4), reports a channel access failure. If it is idle, CW is decreased; when CW reaches zero, that is after two consecutive idle CCAs, the frame is sent on the next boundary, provided the whole transaction fits in the remaining CAP. Unslotted CSMA-CA, used without beacons, has no CW and no boundaries: one idle CCA and the frame goes.

A frame with the acknowledgement request bit set is acknowledged by the recipient aTurnaroundTime (12 symbols) after it ends, or in the CAP on a backoff boundary within one further backoff period; the sender waits up to macAckWaitDuration and retries up to macMaxFrameRetries (3) times. The three data transfer models are data to a coordinator (the device sends when it wishes), data from a coordinator (indirect transmission: the coordinator holds the frame and announces it in the beacon's pending address list, or waits to be polled; the device sends a data request command, and the coordinator sends the frame), and peer-to-peer transfer. The four frame types are beacon, data, acknowledgement and MAC command. Every frame is a MAC header (Frame Control, 2 octets; Sequence Number, 1; addressing fields, 0 to 20; optional auxiliary security header), a payload, and a MAC footer holding the 16-bit FCS, an ITU-T CRC with generator x^16 + x^12 + x^5 + 1.

munotes.in466

CSMA-CA, Data Transfer and Frames in 802.15.4

When CSMA-CA is used, and when it is not

"The CSMA-CA algorithm shall be used before the transmission of data or MAC command frames transmitted within the CAP", with one exception, a data frame sent quickly after the acknowledgement of a data request. And "The CSMA-CA algorithm shall not be used for the transmission of beacon frames in a beacon-enabled PAN, acknowledgment frames, or data frames transmitted in the CFP."

Each exclusion has a reason in the protocol. A beacon owns slot 0 by rule; a GTS is owned by one device ([The 802.15.4 Superframe and Guaranteed Time Slots]); an acknowledgement answers a frame at a fixed moment, and contending for the channel would only delay it.

Which version is used depends on the beacons. "If periodic beacons are being used in the PAN, the MAC sublayer shall employ the slotted version of the CSMA-CA algorithm for transmissions in the CAP of the superframe." Otherwise, or if a device cannot find the beacon, it uses the unslotted version. Both count time in backoff periods of aUnitBackoffPeriod, 20 symbols, which is 320 µs at 2450 MHz.

Three variables

"Each device shall maintain three variables for each transmission attempt: NB, CW and BE."

  • NB. "NB is the number of times the CSMA-CA algorithm was required to backoff while attempting the current transmission; this value shall be initialized to zero before each new transmission attempt."
  • CW. "CW is the contention window length, defining the number of backoff periods that need to be clear of channel activity before the transmission can commence; this value shall be initialized to two before each transmission attempt and reset to two each time the channel is assessed to be busy. The CW variable is only used for slotted CSMA-CA."
  • BE. "BE is the backoff exponent, which is related to how many backoff periods a device shall wait before attempting to assess a channel." It starts at macMinBE; "In slotted systems with the received BLE subfield set to one, this value shall be initialized to the lesser of two and the value of macMinBE."

Slotted CSMA-CA, step by step

A flowchart of slotted CSMA-CA. Start: NB = 0, CW = 2, BE = macMinBE, or with battery life extension the lesser of 2 and macMinBE. Locate the next backoff boundary. Wait a random 0 to 2 to the power BE minus 1 whole backoff periods. CCA on a backoff boundary. Channel idle? If no: CW = 2, NB = NB + 1, BE = the smaller of BE + 1 and macMaxBE; then if NB is greater than macMaxCSMABackoffs, channel access failure, otherwise back to the random wait. If yes: CW = CW - 1; if CW is 0, send on the next boundary, otherwise perform another CCA on the next boundary. A note: unslotted CSMA-CA has no CW and no boundaries, and after one idle CCA the frame goes at once

Figure 66.1 Slotted CSMA-CA, after the standard's Fig. 69

  1. Initialise and align. NB = 0, CW = 2, BE = macMinBE (or the lesser of 2 and macMinBE with battery life extension), then find the next backoff period boundary. "In slotted CSMA-CA, the backoff period boundaries of every device in the PAN shall be aligned with the superframe slot boundaries of the PAN coordinator, i.e., the start of the first backoff period of each device is aligned with the start of the beacon transmission."
  2. Random backoff. Wait a random whole number of backoff periods from 0 to 2^BE - 1.
  3. Assess. Ask the PHY for a CCA, starting on a backoff boundary ([The 802.15.4 Physical Layer] gives its three modes and its 8-symbol duration).
  4. Busy. Increase NB and BE by one, BE no higher than macMaxBE, and reset CW to 2. If NB is still at most macMaxCSMABackoffs, return to step 2; otherwise "the CSMA-CA algorithm shall terminate with a channel access failure status."
  5. Idle. Decrease CW by one. If it is not yet zero, assess again (step 3) on the next boundary. "If it is equal to zero, the MAC sublayer shall begin transmission of the frame on the boundary of the next backoff period."
munotes.in467

CSMA-CA, Data Transfer and Frames in 802.15.4

Fitting in the CAP. After the random backoff, the device must check that the two CCAs, the frame and any acknowledgement can all be completed before the CAP ends. If the backoff itself runs past the end of the CAP, the countdown pauses and resumes in the next superframe's CAP; if the backoff fits but the transaction does not, the device waits for the next CAP and draws a fresh random backoff there.

Battery life extension. A coordinator can set the BLE bit in its beacon. Devices then start with the smaller exponent (at most 2), and "The backoff countdown shall only occur during the first macBattLifeExtPeriods full backoff periods after the end of the IFS period following the beacon." The coordinator's receiver needs to stay on only for that short stretch after its beacon. At 2450 MHz macBattLifeExtPeriods defaults to 6 backoff periods, the sum of three terms in Table 86: 3 for the largest backoff when BE is 2, then 2 for the contention window, then 1 for the preamble and SFD (10 symbols, rounded up to one period).

Unslotted CSMA-CA, by contrast

Without beacons there is no common clock, so "the backoff periods of one device are not related in time to the backoff periods of any other device in the PAN." The device keeps only NB and BE, waits its random backoff, makes one CCA immediately, and if the channel is idle sends at once; if busy, it backs off as in step 4. [CSMA/CA Worked Step by Step] follows it through a burst of reports and sets it beside 802.11's version.

Why two clear assessments

A single idle CCA can be fooled by the gap between a frame and its acknowledgement. The recipient waits at least a turnaround time before it answers, and in the CAP may wait up to one more backoff period for a boundary, so for up to 32 symbols after a frame the channel is quiet although the exchange is not over. A device whose one CCA fell in that gap would send straight into the acknowledgement. Two assessments on consecutive boundaries span more than that gap, so at least one of them falls on the frame or on the acknowledgement. The program shows the difference.

munotes.in468

CSMA-CA, Data Transfer and Frames in 802.15.4

Acknowledgements and retries

"A frame transmitted with the Acknowledgment Request subfield of its Frame Control field set to one shall be acknowledged by the recipient." The acknowledgement carries the same sequence number (the DSN) as the frame it confirms.

When it is sent. "The transmission of an acknowledgment frame in a nonbeacon-enabled PAN or in the CFP shall commence aTurnaroundTime symbols after the reception of the last symbol of the data or MAC command frame." In the CAP it may instead wait for a backoff boundary, and then "the transmission of an acknowledgment frame shall commence between aTurnaroundTime and (aTurnaroundTime + aUnitBackoffPeriod) symbols after the reception of the last symbol of the data or MAC command frame."

How long the sender waits. macAckWaitDuration is Equation (13): aUnitBackoffPeriod + aTurnaroundTime + phySHRDuration + 6 × phySymbolsPerOctet, where 6 is the PHY header octet plus the 5 octets of an acknowledgement. At 2450 MHz that is 20 + 12 + 10 + 12 = 54 symbols, 864 µs. Read as a timeline, it is the latest moment an acknowledgement may start (a turnaround plus one backoff period) plus the time the acknowledgement takes (22 symbols).

Retries. A frame not acknowledged in time is sent again, each time with a fresh CSMA-CA, up to macMaxFrameRetries times, 3 by default. "If a data transfer attempt fails a total of (1 + macMaxFrameRetries) times, the originator MAC sublayer will issue a failure confirmation to the next higher layer." A frame sent without the request bit is assumed delivered and never retried.

Three data transfer models

"Three types of data transfer transactions exist. The first one is the data transfer to a coordinator in which a device transmits the data. The second transaction is the data transfer from a coordinator in which the device receives the data. The third transaction is the data transfer between two peer devices." A star uses only the first two; a peer-to-peer network may use all three.

To a coordinator. In a beacon-enabled PAN, the device "first listens for the network beacon. When the beacon is found, the device synchronizes to the superframe structure. At the appropriate time, the device transmits its data frame, using slotted CSMA-CA, to the coordinator." The coordinator may acknowledge. In a nonbeacon-enabled PAN the device "simply transmits its data frame, using unslotted CSMA-CA, to the coordinator."

munotes.in469

CSMA-CA, Data Transfer and Frames in 802.15.4

From a coordinator: indirect transmission. A coordinator cannot simply send to a device, which may be asleep. So in a beacon-enabled PAN, "it indicates in the network beacon that the data message is pending. The device periodically listens to the network beacon and, if a message is pending, transmits a MAC command requesting the data, using slotted CSMA-CA." Then:

  1. The coordinator acknowledges the data request command. If it knows in time, it sets the Frame Pending bit of that acknowledgement to say whether a frame is really waiting; if it cannot tell in time, it sets the bit to one.
  2. The coordinator sends the data frame, either at once (between aTurnaroundTime and one backoff period later, on a boundary, if it fits in the CAP) or by slotted CSMA-CA.
  3. The device may acknowledge it. The address then leaves the pending list.

If nothing was pending after all, the coordinator sends a data frame with a zero-length payload. In a nonbeacon-enabled PAN there is no beacon to announce anything: the coordinator stores the frame, and the device polls with a data request "at an application-defined rate".

The pending list is small: "The maximum number of addresses pending shall be limited to seven and may comprise both short and extended addresses." And a frame does not wait forever: the coordinator keeps it for macTransactionPersistenceTime, by default 0x01f4 (500) unit periods, where a unit period is a beacon interval in a beacon-enabled PAN.

Indirect transmission is the MAC's answer to the energy problem of [MAC Protocols for Sensor Networks: The Job and Where the Energy Goes]: the device, not the coordinator, decides when its receiver is on, and the always-on job falls to the coordinator.

Peer to peer. "In a peer-to-peer PAN, every device may communicate with every other device in its radio sphere of influence. In order to do this effectively, the devices wishing to communicate will need to either receive constantly or synchronize with each other." Receiving constantly allows plain unslotted CSMA-CA; synchronising is left open: "Such measures are beyond the scope of this standard."

Four frame types

"The frame structures have been designed to keep the complexity to a minimum while at the same time making them sufficiently robust for transmission on a noisy channel." The standard defines four:

  • "A beacon frame, used by a coordinator to transmit beacons"
  • "A data frame, used for all transfers of data"
  • "An acknowledgment frame, used for confirming successful frame reception"
  • "A MAC command frame, used for handling all MAC peer entity control transfers"

The nine MAC commands of the 2006 edition are association request and response, disassociation notification, data request, PAN ID conflict notification, orphan notification, beacon request, coordinator realignment and GTS request (Table 82).

munotes.in470

CSMA-CA, Data Transfer and Frames in 802.15.4

The general MAC frame format

Top: the general MAC frame as a row of fields with their sizes in octets: frame control 2, sequence number 1, addressing fields (PAN identifiers and addresses) 0 to 20, security header 0 to 14, payload variable, FCS 2; the first four form the MAC header, the payload follows, and the FCS forms the MAC footer. Bottom: the 16 bits of the Frame Control field from b0: frame type bits 0 to 2, security enabled bit 3, frame pending bit 4, acknowledgement request bit 5, PAN ID compression bit 6, reserved bits 7 to 9, destination addressing mode bits 10 and 11, frame version bits 12 and 13, source addressing mode bits 14 and 15; frame type 0 is beacon, 1 data, 2 acknowledgement, 3 command

Figure 66.2 The general MAC frame and its Frame Control field, after the standard's Figs. 41 and 42

Every frame is a MAC header (MHR), a MAC payload and a MAC footer (MFR):

  • Frame Control (2 octets): the frame type (beacon 0, data 1, acknowledgement 2, MAC command 3), Security Enabled, Frame Pending, Acknowledgement Request, PAN ID Compression, the destination and source addressing modes, and the frame version (0 for a frame compatible with the 2003 edition, 1 for the 2006 format).
  • Sequence Number (1 octet): a beacon sequence number (BSN) in beacons; otherwise a data sequence number (DSN) "that is used to match an acknowledgment frame to the data or MAC command frame."
  • Addressing fields (0 to 20 octets): destination PAN identifier and address, source PAN identifier and address. Each addressing mode is 0 (no PAN identifier and no address), 2 (a 16-bit short address) or 3 (a 64-bit extended address). The PAN identifier 0xffff and the short address 0xffff are the broadcast values.
  • Auxiliary Security Header (0, 5, 6, 10 or 14 octets), present only when Security Enabled is set ([Built on 802.15.4: Zigbee Routing, Security and the Later Amendments]).
  • Payload, whose contents depend on the frame type.
  • FCS (2 octets). "The FCS field is 2 octets in length and contains a 16-bit ITU-T CRC. The FCS is calculated over the MHR and MAC payload parts of the frame." The generator is x^16 + x^12 + x^5 + 1.

PAN ID compression. "If this subfield is set to one and both the source and destination addresses are present, the frame shall contain only the Destination PAN Identifier field, and the Source PAN Identifier field shall be assumed equal to that of the destination." Within one PAN this saves two octets in every frame.

The acknowledgement is the smallest frame: Frame Control, Sequence Number and FCS, 5 octets, and no addresses at all. It is matched to its frame by the sequence number alone.

The MAC, computed

The program builds frame overheads from the field sizes of Figure 41; reproduces the FCS example of 7.2.1.9 and decodes its header; simulates slotted CSMA-CA for 5 and for 20 devices that each have one 20-octet frame when a beacon ends, as an event seen by all would give them, with the default constants, with the exponent starting at 2, and with a single assessment; and times the persistence of a pending frame.

# IEEE 802.15.4-2006 MAC frames and slotted CSMA-CA, at 2450 MHz (2 symbols an
# octet, 16 us a symbol). Field sizes are from Figure 41; the FCS example and the
# CSMA-CA rules from 7.2.1.9 and 7.5.1.4.
import random

# 1. The general MAC frame: frame control 2, sequence number 1, addressing
#    fields, (no security here), payload, FCS 2.
def overhead(dst, src, same_pan):
    size = {"none": 0, "short": 2, "extended": 8}
    pans = (2 if dst != "none" else 0) + (2 if src != "none" and not (same_pan and dst != "none") else 0)
    return 2 + 1 + pans + size[dst] + size[src] + 2

print("MAC overhead and room for payload in a 127-octet PSDU:")
for label, dst, src, same in (("short destination only", "short", "none", True),
                              ("short to short, one PAN ID", "short", "short", True),
                              ("extended to extended, one PAN ID", "extended", "extended", True),
                              ("extended to extended, two PAN IDs", "extended", "extended", False)):
    o = overhead(dst, src, same)
    print("  %-34s %2d octets, payload up to %3d" % (label, o, 127 - o))

# 2. The FCS: the ITU-T CRC with G(x) = x^16 + x^12 + x^5 + 1, the bits taken in
#    the order they are sent. The standard's example: an acknowledgement's MHR.
def fcs(bits):
    reg = [int(b) for b in bits] + [0] * 16
    gen = [int(b) for b in "10001000000100001"]
    for i in range(len(bits)):
        if reg[i]:
            for j, g in enumerate(gen):
                reg[i + j] ^= g
    return "".join(map(str, reg[-16:]))

mhr = "0100 0000 0000 0000 0101 0110".replace(" ", "")
value = lambda bits: sum(int(b) << i for i, b in enumerate(bits))   # b0 is the LSB
print("\nFCS of the example MHR:", fcs(mhr), "(the standard prints 0010011110011110)")
print("frame type %d (2 = acknowledgement), sequence number 0x%02X"
      % (value(mhr[0:3]), value(mhr[16:24])))

# 3. Slotted CSMA-CA after a beacon. n devices each have one frame when the
#    beacon ends (an event seen by all); time is counted in backoff periods of
#    20 symbols from the start of the beacon (38 symbols long). A CCA looks at
#    the first 8 symbols of a period; a frame is lost if it overlaps any other
#    transmission. The acknowledgement of a clean frame starts on the first
#    backoff boundary at least aTurnaroundTime (12 symbols) after it, as 7.5.6.4.2
#    allows in the CAP, and is lost if a frame overlaps it. No reply within
#    macAckWaitDuration (54 symbols) means a retry, up to macMaxFrameRetries (3).
#    The CAP is taken as long enough.
def burst(n, cw0=2, be0=3, mpdu=20, trials=600, seed=66):
    rnd = random.Random(seed)
    frame, stats = 2 * (6 + mpdu), dict(ok=0, access=0, retries=0, acks_hit=0, first=0, time=0)
    for _ in range(trials):
        on_air = []                     # [start, end, kind, device, outcome]
        dev = [dict(nb=0, cw=cw0, be=be0, tries=0, cca=2 + rnd.randrange(2 ** be0), tx=None)
               for _ in range(n)]
        done, k = 0, 0
        def overlaps(a, b, me):
            return any(t[0] < b and a < t[1] for t in on_air if (t[2], t[3]) != me)
        def retry(d, end):
            dv = dev[d]
            dv["tries"] += 1
            if dv["tries"] > 3:
                return "retries"
            dv.update(nb=0, cw=cw0, be=be0)
            dv["cca"] = -(-(end + 54) // 20) + rnd.randrange(2 ** be0)
            return None
        while done < n:
            now = 20 * k
            for d, dv in enumerate(dev):                     # frames start on a boundary
                if dv["tx"] == k:
                    on_air.append([now, now + frame, "frame", d, None])
                    dv["tx"] = None
            for t in sorted((t for t in on_air if t[4] is None and t[1] <= now), key=lambda t: t[1]):
                s, e, kind, d = t[:4]
                t[4] = "lost" if overlaps(s, e, (kind, d)) else "clean"
                if kind == "frame" and t[4] == "clean":
                    a = -(-(e + 12) // 20) * 20
                    on_air.append([a, a + 22, "ack", d, None, e])
                    continue
                if kind == "ack" and t[4] == "clean":
                    stats["ok"] += 1
                    stats["first"] += dev[d]["tries"] == 0
                    stats["time"] += e
                    done += 1
                    continue
                if kind == "ack":
                    stats["acks_hit"] += 1
                failed = retry(d, e if kind == "frame" else t[5])
                if failed:
                    stats[failed] += 1
                    done += 1
            for d, dv in enumerate(dev):                     # clear channel assessments
                if dv["cca"] != k:
                    continue
                dv["cca"] = None
                if overlaps(now, now + 8, None):
                    dv["nb"] += 1
                    dv["be"] = min(dv["be"] + 1, 5)
                    dv["cw"] = cw0
                    if dv["nb"] > 4:
                        stats["access"] += 1
                        done += 1
                    else:
                        dv["cca"] = k + 1 + rnd.randrange(2 ** dv["be"])
                else:
                    dv["cw"] -= 1
                    if dv["cw"] == 0:
                        dv["tx"] = k + 1
                    else:
                        dv["cca"] = k + 1
            on_air = [t for t in on_air if t[1] > now - 200]
            k += 1
    frames = n * trials
    return [100 * stats[x] / frames for x in ("ok", "first", "access", "retries")] + \
           [stats["acks_hit"] / trials, stats["time"] * 16 / 1000 / max(stats["ok"], 1)]

print("\nOne 20-octet frame each, sent after the beacon (600 bursts):")
print("                        delivered  first try  gave up:  busy  retries   ACKs hit   mean ms")
for label, n, cw0, be0 in (("5 devices, standard", 5, 2, 3), ("20 devices, standard", 20, 2, 3),
                           ("20, BE starts at 2", 20, 2, 2), ("20, one CCA (CW = 1)", 20, 1, 3)):
    ok, first, access, retries, acks, ms = burst(n, cw0, be0)
    print("  %-21s %8.1f%% %9.1f%% %12.1f%% %7.1f%% %10.2f %9.2f" % (label, ok, first, access, retries, acks, ms))

# 4. Waiting for mail: a pending frame is kept for macTransactionPersistenceTime,
#    0x01f4 unit periods of aBaseSuperframeDuration x 2^BO symbols.
print("\nA pending frame waits at most 0x01f4 = %d beacon intervals:" % 0x01f4)
for bo in (0, 6, 10):
    print("  BO = %2d: %9.2f s" % (bo, 0x01f4 * 960 * 2 ** bo * 16e-6))
print("macBattLifeExtPeriods at 2450 MHz: %d backoff periods" % ((2 ** 2 - 1) + 2 + -(-10 // 20)))
munotes.in471

CSMA-CA, Data Transfer and Frames in 802.15.4

MAC overhead and room for payload in a 127-octet PSDU:
  short destination only              9 octets, payload up to 118
  short to short, one PAN ID         11 octets, payload up to 116
  extended to extended, one PAN ID   23 octets, payload up to 104
  extended to extended, two PAN IDs  25 octets, payload up to 102

FCS of the example MHR: 0010011110011110 (the standard prints 0010011110011110)
frame type 2 (2 = acknowledgement), sequence number 0x6A

One 20-octet frame each, sent after the beacon (600 bursts):
                        delivered  first try  gave up:  busy  retries   ACKs hit   mean ms
  5 devices, standard       98.5%      74.3%          1.4%     0.0%       0.00     12.07
  20 devices, standard      50.8%      16.5%         46.6%     2.6%       0.00     25.65
  20, BE starts at 2        38.8%       8.6%         55.5%     5.8%       0.00     23.62
  20, one CCA (CW = 1)      29.8%       2.1%         39.1%    31.1%       6.72     37.04

A pending frame waits at most 0x01f4 = 500 beacon intervals:
  BO =  0:      7.68 s
  BO =  6:    491.52 s
  BO = 10:   7864.32 s
macBattLifeExtPeriods at 2450 MHz: 6 backoff periods
munotes.in472

CSMA-CA, Data Transfer and Frames in 802.15.4

Overheads. With only a short destination address (a frame from the PAN coordinator) the MAC adds 9 octets, which is the standard's aMinMPDUOverhead, and leaves 118 for payload, its aMaxMACPayloadSize. Short to short within one PAN costs 11 octets and leaves 116. Extended addresses with two PAN identifiers cost 25 octets, the standard's aMaxMPDUUnsecuredOverhead, leaving 102: the figure RFC 4944 starts from for IPv6 ([The 802.15.4 Physical Layer]). The field sizes alone reproduce the standard's constants.

munotes.in473

CSMA-CA, Data Transfer and Frames in 802.15.4

The FCS. Dividing the example header, bits in the order they are sent, by the generator gives 0010011110011110, exactly the standard's result. The header decodes as frame type 2, an acknowledgement, with sequence number 0x6A.

Five devices. With the defaults, 98.5 per cent of reports got through, three quarters at the first try, about 12 ms after the beacon on average; 1.4 per cent were given up after five busy assessments.

Twenty devices. Only 50.8 per cent got through. The losses were mostly channel access failures, 46.6 per cent: with twenty devices waking together, a device often finds the channel busy five times in a row and gives up, the same weakness [CSMA/CA Worked Step by Step] found for the unslotted version. Starting the exponent at 2, as battery life extension does, made it worse (38.8 per cent): smaller windows put more devices on the same boundaries.

Why CW is 2. With a single assessment, 6.72 acknowledgements per burst were destroyed by frames sent into the gap before them, 31.1 per cent of reports ran out of retries, and only 29.8 per cent were delivered. With the standard's two assessments, no acknowledgement was hit in 600 bursts. That second CCA is cheap insurance.

munotes.in474

CSMA-CA, Data Transfer and Frames in 802.15.4

Waiting for mail. A pending frame is kept for 500 beacon intervals: 7.68 s at BO = 0, 491.52 s at BO = 6, over two hours at BO = 10. A device that sleeps through more beacons than that loses its mail. The last line confirms the default macBattLifeExtPeriods of 6.

Distinctions

Slotted CSMA-CAUnslotted CSMA-CA
Used inThe CAP of a beacon-enabled PANA nonbeacon-enabled PAN, or when the beacon is lost
TimingBackoff boundaries aligned to the beaconEach device's own clock
VariablesNB, CW, BENB, BE
Sends afterTwo idle CCAs on consecutive boundaries (CW from 2 to 0)One idle CCA, at once
Extra ruleThe whole transaction must fit in the CAPNone
Direct transmissionIndirect transmission
DirectionDevice to coordinator (or peer)Coordinator to device
Who starts itThe sender, when it has dataThe device, by a data request command
How the device learnsNot neededThe beacon's pending address list, or by polling
WhyThe coordinator is listeningThe device may be asleep
FrameSent byCarriesCSMA-CA
BeaconA coordinatorSuperframe specification, GTS fields, pending addresses, beacon payloadNo (slot 0)
DataAny deviceHigher-layer dataIn the CAP, or none in a GTS
AcknowledgementThe recipientOnly frame control, sequence number, FCSNever
MAC commandAny deviceOne of nine commands (association, data request, GTS request and others)Yes, always in the CAP

What it does not mean

CSMA-CA does not detect collisions. It avoids some; a collision is noticed only as a missing acknowledgement.

Two CCAs do not guarantee a clear channel. Two devices can pass both on the same boundaries and send together; the program's twenty devices did.

A channel access failure is not a collision. The frame was never sent: the device found the channel busy too often and gave up.

An acknowledgement does not prove the data is right for the application. It confirms that a frame with that sequence number arrived with a correct FCS at the next hop.

Indirect transmission is not slower by accident. It trades delay for the device's sleep: data for a device waits until the device asks.

A broadcast is not acknowledged. "A beacon or acknowledgment frame shall always be sent with the Acknowledgment Request subfield set to zero. Similarly, any frame that is broadcast shall be sent with its Acknowledgment Request subfield set to zero." Many recipients cannot all answer, so reliability for a broadcast is left to the layers above.

Quick revision

  • CSMA-CA for data and command frames in the CAP; not for beacons, acknowledgements or GTS frames.
  • Variables: NB (backoffs, from 0), CW (idle periods still needed, from 2, slotted only), BE (from macMinBE, 3; BLE: the lesser of 2 and macMinBE).
  • Slotted: align to backoff boundaries (20 symbols, from the beacon's start); wait random 0 to 2^BE - 1 periods; CCA; busy: CW = 2, NB + 1, BE + 1 (max 5), give up when NB exceeds 4; idle: CW - 1, send when CW = 0 on the next boundary; the transaction must fit in the CAP.
  • Unslotted: no CW, no alignment; send after one idle CCA.
  • Why CW = 2: the gap before an acknowledgement (12 to 32 symbols); program: CW 1 lost 6.72 ACKs per burst, CW 2 none.
  • ACK: after aTurnaroundTime (12 symbols), in the CAP possibly on the next boundary; macAckWaitDuration = 20 + 12 + 10 + 12 = 54 symbols; 3 retries, then failure.
  • Transfers: to a coordinator (slotted or unslotted CSMA-CA), from a coordinator (indirect: pending list in the beacon, at most 7 addresses, or polling; data request; frame pending bit), peer to peer (receive constantly or synchronise).
  • Frames: beacon, data, acknowledgement, MAC command (9 commands). Format: Frame Control 2, Sequence Number 1, addressing 0 to 20, security header, payload, FCS 2 (CRC-16, x^16 + x^12 + x^5 + 1).
  • Program: overheads 9 / 11 / 23 / 25 octets (payload 118 / 116 / 104 / 102); the FCS example reproduced; 20 devices after a beacon: 50.8 per cent delivered, 46.6 per cent channel access failures.
munotes.in475

CSMA-CA, Data Transfer and Frames in 802.15.4

Test yourself

1. When does IEEE 802.15.4 use CSMA-CA, and which version? CSMA-CA is used before sending data and MAC command frames in the contention access period, except a data frame sent quickly after acknowledging a data request. It is not used for beacons in a beacon-enabled PAN, for acknowledgement frames, or for data frames in the contention-free period. The slotted version is used in the CAP when periodic beacons are in use; the unslotted version is used when there are no beacons or a device cannot find the beacon.

2. Explain the slotted CSMA-CA algorithm with its variables. Each attempt keeps NB, the number of backoffs, initialised to 0; CW, the contention window, the number of backoff periods that must be clear, initialised to 2; and BE, the backoff exponent, initialised to macMinBE (or the lesser of 2 and macMinBE under battery life extension). The device locates the next backoff boundary, waits a random 0 to 2^BE - 1 backoff periods, and performs a CCA on a boundary. If the channel is busy, it resets CW to 2, increments NB and BE (BE at most macMaxBE) and, if NB is at most macMaxCSMABackoffs, backs off again; otherwise it reports a channel access failure. If the channel is idle, it decrements CW and, if CW is not yet zero, performs another CCA on the next boundary; when CW reaches zero it transmits on the next boundary. It proceeds only if the CCAs, the frame and the acknowledgement fit before the CAP ends, otherwise it waits for the next CAP.

munotes.in476

CSMA-CA, Data Transfer and Frames in 802.15.4

3. Why does slotted CSMA-CA require two clear channel assessments? After a frame the recipient waits a turnaround time before acknowledging, and in the CAP may wait up to a further backoff period for a boundary, so the channel is quiet for a short gap in the middle of an exchange. A device relying on one assessment that fell in that gap would transmit into the acknowledgement. Two assessments on consecutive backoff boundaries span more than the gap, so a transmission in progress is detected. In the program, one assessment let 6.72 acknowledgements per burst be destroyed; two let none be.

4. Describe the three data transfer models of 802.15.4. Data to a coordinator: in a beacon-enabled PAN the device synchronises to the beacon and sends its frame with slotted CSMA-CA; in a nonbeacon-enabled PAN it simply sends with unslotted CSMA-CA; the coordinator may acknowledge. Data from a coordinator (indirect transmission): the coordinator stores the frame and, in a beacon-enabled PAN, lists the device's address in the beacon's pending address list; the device sends a data request command, the coordinator acknowledges and then sends the data, which the device may acknowledge; in a nonbeacon-enabled PAN the device polls with data requests at an application-defined rate. Peer-to-peer: any device may send to any device in range, but they must either receive constantly or synchronise, which the standard leaves to higher layers.

5. Why does 802.15.4 use indirect transmission for data from a coordinator? A device, especially a reduced-function one, spends most of its time with its receiver off to save energy, so a frame sent to it at a random moment would be lost. Indirect transmission lets the coordinator hold the frame, announce it in the beacon or wait to be polled, and send it only when the device has asked, so the device alone decides when to listen and the always-on burden falls on the coordinator.

6. Name the four frame types and draw the general MAC frame format. The frame types are beacon, data, acknowledgement and MAC command. The general frame is a MAC header, a payload and a MAC footer: Frame Control (2 octets), Sequence Number (1 octet), the addressing fields (destination PAN identifier, destination address, source PAN identifier, source address; 0 to 20 octets), an optional auxiliary security header (0, 5, 6, 10 or 14 octets), the frame payload (variable), and the frame check sequence (2 octets). The Frame Control field holds the frame type, security enabled, frame pending, acknowledgement request and PAN ID compression bits, the destination and source addressing modes and the frame version.

munotes.in477

CSMA-CA, Data Transfer and Frames in 802.15.4

7. What is the FCS of an 802.15.4 frame, and how is it computed? It is a 16-bit ITU-T CRC over the MAC header and payload. The bits of the header and payload, taken in the order they are transmitted, are treated as a polynomial, multiplied by x^16 and divided modulo 2 by the generator x^16 + x^12 + x^5 + 1; the 16-bit remainder is the FCS. The standard's example, an acknowledgement header of three octets, gives 0010 0111 1001 1110, which the program reproduces.

8. Compute macAckWaitDuration at 2450 MHz and explain its terms. macAckWaitDuration is aUnitBackoffPeriod + aTurnaroundTime + phySHRDuration + 6 × phySymbolsPerOctet. At 2450 MHz these are 20, 12, 10 (8 symbols of preamble and 2 of SFD) and 6 × 2 = 12, a total of 54 symbols, 864 µs. The first two terms are the latest an acknowledgement may start after the frame (a turnaround plus up to one backoff period for a boundary); the last two are the time the acknowledgement itself takes, its synchronisation header plus one PHY header octet and five MAC octets.

Contents This chapter on its own page

munotes.in478

Chapter Sixty-Seven

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

Syllabus topic Module 1, "Routing in WSN: IEEE 802.15.4 LR-WPAN standard (case study)"

In one line

802.15.4 stops at the MAC; Zigbee builds a network layer on it that routes in two ways (along a tree whose addresses are handed out in blocks sized by Cskip, so a router can route by arithmetic alone, or along mesh routes found by route discovery and chosen by summed link costs), secures frames with a network key and link keys distributed by a Trust Center, while 802.15.4 itself moved on: its 2012 amendment added time-slotted channel hopping, now part of the 2015 standard.

In the wording a student can write in an examination: IEEE 802.15.4 defines only the PHY and MAC; multi-hop routing is provided by higher layers such as Zigbee, whose network (NWK) layer sits above the 802.15.4 MAC. Zigbee has three device types: the Zigbee coordinator (an 802.15.4 PAN coordinator), Zigbee routers (FFDs that route messages and accept associations) and Zigbee end devices (which do not route). In distributed address assignment, the coordinator fixes the maximum children per parent Cm, the maximum router children Rm and the maximum depth Lm; a parent at depth d gives each router child a block of Cskip(d) addresses, where Cskip(d) = 1 + Cm(Lm - d - 1) if Rm = 1, and otherwise Cskip(d) = (1 + Cm - Rm - Cm × Rm^(Lm - d - 1)) / (1 - Rm). Router children get addresses A + 1, A + 1 + Cskip(d), ...; the n-th end device gets A + Rm × Cskip(d) + n. In tree (hierarchical) routing, a router at address A and depth d checks whether the destination D is a descendant, A < D < A + Cskip(d - 1); if so it forwards to the child whose block holds D, N = A + 1 + floor((D - (A + 1)) / Cskip(d)) × Cskip(d) (or to D itself if D is one of its end devices), and otherwise to its parent.

In mesh routing, a router finds a route by route discovery: it broadcasts a route request, each router adds the link cost from the previous hop to a path cost, and the destination unicasts a route reply back; the route with the lowest path cost wins. The link cost is min(7, round(1/p^4)), where p is the link's delivery probability, or a constant 7. A concentrator (a data sink) can set up routes to itself from every router at once with many-to-one route discovery. Zigbee secures its NWK and APS layers with AES in CCM* mode, using a network key shared by all devices (for broadcasts and network-layer frames) and link keys shared by pairs (for unicast application data), distributed by the Trust Center. IEEE 802.15.4e (2012) added time-slotted channel hopping (TSCH), merged into 802.15.4-2015: all nodes are synchronised, time is divided into timeslots grouped into repeating slotframes, each link has cells (slotOffset, channelOffset), and the frequency hops as F((ASN + channelOffset) mod nFreq), where the absolute slot number (ASN) counts every slot since the network began; hopping defeats multipath fading and interference on any one channel.

munotes.in479

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

Why MU files 802.15.4 under routing

The standard itself routes nothing. It defines how a frame crosses one hop and, as [IEEE 802.15.4: The Standard, Its Devices and Its Topologies] explained, leaves network formation beyond that and multi-hop routing to higher layers. What the MAC does provide is the material a router needs: addresses, beacons and association, which build a tree of parents and children, and per-packet link quality, which measures each hop ([The 802.15.4 Physical Layer]).

Zigbee turns that material into routes. Its device types are defined on 802.15.4's:

  • "ZigBee coordinator: an IEEE 802.15.4 PAN coordinator."
  • "ZigBee router: an IEEE 802.15.4 FFD participating in a ZigBee network, which is not the ZigBee coordinator but may act as an IEEE 802.15.4 coordinator within its personal operating space, that is capable of routing messages between devices and supporting associations."
  • "ZigBee end device: an IEEE 802.15.4 RFD or FFD participating in a ZigBee network, which is neither the ZigBee coordinator nor a ZigBee router."

Revision 22 defines the same roles on "IEEE 802.15.4-2015" devices, the edition described at the end of this chapter.

Addresses handed out in blocks

A Zigbee device needs a 16-bit network address. With the default setting, addresses are "assigned using a distributed addressing scheme that is designed to provide every potential parent with a finite sub-block of network addresses." No central server is needed: each parent hands out addresses from its own block.

The coordinator fixes three numbers: the maximum number of children a parent may have, nwkMaxChildren (Cm); the maximum number of those that may be routers, nwkMaxRouters (Rm); and the maximum depth, nwkMaxDepth (Lm). "The ZigBee coordinator itself has a depth of 0, while its children have a depth of 1." From these, the specification computes "the function, Cskip(d), essentially the size of the address sub-block being distributed by each parent at that depth to its router-capable child devices". Its formula, printed as an image, is:

  • if Rm = 1: Cskip(d) = 1 + Cm × (Lm - d - 1);
  • otherwise: Cskip(d) = (1 + Cm - Rm - Cm × Rm^(Lm - d - 1)) / (1 - Rm).

The addresses follow. "A parent assigns an address that is 1 greater than its own to its first router-capable child device. Subsequently assigned addresses to router-capable child devices are separated from each other by Cskip(d)." End devices come after the routers' blocks: the n-th end device of a parent at address A gets A + Cskip(d) × Rm + n. And "If a device has a Cskip(d) value of 0, then it shall not be capable of accepting children".

munotes.in480

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

Why the formula is right. A router's block must hold the router itself, the blocks of its Rm router children and its Cm - Rm end devices: Cskip(d - 1) = 1 + Rm × Cskip(d) + (Cm - Rm), with Cskip(Lm - 1) = 1. The closed formula solves that recurrence, and the program checks it both ways.

Worked, with the specification's example. The text sets Cm = 6, Rm = 4, Lm = 3. Then Cskip(0) = (1 + 6 - 4 - 6 × 16) / (1 - 4) = (3 - 96) / (-3) = 31; Cskip(1) = (3 - 24) / (-3) = 7; Cskip(2) = (3 - 6) / (-3) = 1; and Cskip(3) = 0, since a device at the maximum depth has no children. The coordinator's routers are 1, 32, 63 and 94, and its two end devices 125 and 126.

The specification's slip. Table 3-65 prints 31, 9, 1 and 0. The 9 cannot be right for Cm = 6: a depth-1 router's block would then need 1 + 4 × 9 + 2 = 39 addresses, more than the 31 its parent gave it. Figure 3-46 is drawn with a third set: its routers carry Cskip 41 and 9, the values for Cm = 8 (routers at 1, 42, 83 and 124), yet two end devices sit at 30 and 31, which fit neither. Trust the formula; the program uses Cm = 8 to match the figure's routers.

A tree of Zigbee network addresses for Cm = 8, Rm = 4, Lm = 3. The coordinator, address 0, with Cskip(0) = 41, has router children 1, 42, 83 and 124, and end devices 165 to 168. Router 42, at depth 1 with Cskip(1) = 9, has router children 43, 52, 61 and 70 and end devices 79 to 82. Router 124 has router children 125, 134, 143 and 152. Router 134, at depth 2 with Cskip(2) = 1, has children 135 to 138 and 139 to 142, which at depth 3 have no children. A bold line traces the tree route from end device 139 up through 134, 124 and the coordinator, then down through 42 to 70: five hops

Figure 67.1 Zigbee's distributed addresses, Cm = 8, Rm = 4, Lm = 3, and one tree route

The price of blocks. "Because an address sub-block cannot be shared between devices, it is possible that one parent exhausts its list of addresses while a second parent has addresses that go unused." A parent with no addresses left refuses new devices, which must find another parent or cannot join at all.

Routing by arithmetic: the tree

With addresses in blocks, a router can route without any table. "For hierarchical routing, if the destination is a descendant of the device, the device shall route the frame to the appropriate child. ... If the destination is not a descendant, the device shall route the frame to its parent."

The test and the next hop are both arithmetic:

  • Descendant test. For a Zigbee router with address A at depth d, a destination D is a descendant if A < D < A + Cskip(d - 1): D lies inside A's own block. "Every other device in the network is a descendant of the ZigBee coordinator and no device in the network is the descendant of any ZigBee end device."
  • Next hop. If D is one of A's end devices (D > A + Rm × Cskip(d)), the next hop is D itself. Otherwise it is the router child whose block holds D: N = A + 1 + floor((D - (A + 1)) / Cskip(d)) × Cskip(d).
munotes.in481

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

Worked. Take Cm = 8, Rm = 4, Lm = 3, and a frame from end device 139 (a child of router 134) to router 70. An end device always sends to its parent, 134. Router 134 is at depth 2, and its block runs from 134 to 142 (it was given Cskip(1), 9 addresses); 70 is outside it, so up to 124. Router 124's block runs to 164; 70 is outside, so up to the coordinator. The coordinator is everyone's ancestor. For it, (70 - 1) / 41 rounds down to 1, so the next hop is the router child whose block starts one block of 41 after address 1: router 42. At router 42 (depth 1), 70 lies between 42 and 42 + 41 = 83, so it is a descendant, and it is not above 42 + 4 × 9 = 78, so it is not an end device. Here (70 - 43) / 9 is exactly 3, so the next hop is 43 + 3 × 9 = 70, the destination. Five hops, drawn bold in the figure.

What it costs. Tree routing needs no route discovery and no routing table, which suits the smallest routers. But every route follows the tree, so two devices side by side in different branches talk through their common ancestor, perhaps the coordinator; and when a parent fails, its whole block is cut off. Mesh routing answers both.

Routing by discovery: the mesh

"Route discovery is the procedure whereby network devices cooperate to find and establish routes through the NWK." The steps work much as AODV's do ([AODV: Route Discovery and Route Maintenance]), with a path cost where AODV counts hops:

  1. A router with no route to a destination broadcasts a route request command.
  2. Each router that receives it computes "the link cost from the previous device that transmitted the frame" and adds it to the path cost carried in the request, then rebroadcasts.
  3. The destination (or the parent of an end device that is the destination) answers with a route reply, unicast hop by hop back toward the originator, carrying the path cost.
  4. Routers record routing table entries along the way; the route with the least path cost is kept.
munotes.in482

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

Path cost and link cost. "In order to compute this metric, a cost, known as the link cost, is associated with each link in the path and link cost values are summed to produce the cost for the path as a whole." The link cost for a link l with delivery probability p is either a constant 7 or min(7, round(1 / p^4)) (an image in the PDF). How to estimate p is left to implementers, but "the initial cost estimates shall be based on average LQI" ([The 802.15.4 Physical Layer]).

Many-to-one. In a sensor network most traffic flows to one sink. "Many-to-one route discovery is performed by a source device to establish routes to itself from all ZigBee routers and ZigBee coordinator, within a given radius. A source device that initiates a many-to-one route discovery is designated as a concentrator and referred to as such in this document." One discovery, from the sink, replaces a discovery from every source, much as the collection trees of [Routing Tables and What Happens When the Topology Changes] are built from the root.

Security in Zigbee

802.15.4's own security, its eight security levels, auxiliary header and frame counter, is taught in [Keys and Link Security: Key Predistribution, SPINS and 802.15.4]; the MAC secures a frame with whatever key it is given and does not say where keys come from. Zigbee adds that part.

Two layers. "The ZigBee security architecture includes security mechanisms at two layers of the protocol stack. The NWK and APS layers are responsible for the secure transport of their respective frames." Both use AES in CCM* mode.

Two kinds of key. "Security amongst a network of ZigBee devices is based on 'link' keys and a 'network' key. Unicast communication between APL peer entities is secured by means of a 128-bit link key shared by two devices, while broadcast communications and any network layer communications are secured by means of a 128-bit network key shared amongst all devices in the network."

The Trust Center. "The Trust Center is the device trusted by devices within a network to distribute keys for the purpose of network and potentially end-to-end application configuration management." In a centralised network there is exactly one; in a distributed one, all routers can hand out the network key.

Where it is weak. The specification is candid. "During initial key transport the keying material used for protection may be a well-known key, thus resulting in a brief moment of vulnerability where the key could be obtained by any device." And "due to the low-cost nature of ad hoc network devices, one cannot generally assume the availability of tamper-resistant hardware." A captured node gives up the network key, the weakness of any network-wide key.

munotes.in483

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

The later amendments: time-slotted channel hopping

802.15.4 kept changing after the 2006 edition this case study follows. RFC 7554 records the sequence: "IEEE 802.15.4e was published in 2012 as an amendment to the Medium Access Control (MAC) protocol defined by the IEEE 802.15.4 standard (of 2011)", to be "rolled into the next revision of IEEE 802.15.4, scheduled to be published in 2015." Its best-known part is TSCH.

"At its core is a medium access technique that uses time synchronization to achieve low-power operation and channel hopping to enable high reliability." It changes only the MAC: "TSCH does not amend the physical layer, i.e., it can operate on any hardware that is compliant with IEEE 802.15.4."

  • Timeslots. "All nodes in a TSCH network are synchronized. Time is sliced up into time slots. A time slot is long enough for a MAC frame of maximum size to be sent from node A to node B, and for node B to reply with an acknowledgment (ACK) frame indicating successful reception." At 2.4 GHz a 127-byte frame takes about 4 ms and a typical slot is 10 ms.
  • Slotframes. Slots are grouped into slotframes that repeat. "The shorter the slotframe, the more often a time slot repeats, resulting in more available bandwidth, but also in a higher power consumption."
  • Cells. The schedule tells each node, for every slot, to transmit, receive or sleep; a cell is a (slotOffset, channelOffset) pair reserved for one link, dedicated by default, or shared, with a backoff.
  • ASN. The absolute slot number counts every slot since the network began: ASN = k × S + t for slotframe cycle k, slotframe size S and slotOffset t.
  • Hopping. "frequency = F {(ASN + channelOffset) mod nFreq}", where F is a lookup table of the nFreq available channels (16 at 2.4 GHz). The ASN moves on each cycle, so the same cell lands on a different channel each time: "even with a static schedule, pairs of neighbors 'hop' between the different frequencies when communicating."

Why hop? Industrial sites are hard radio environments, where "vast deployment environments with large (metallic) equipment cause multi-path fading and interference to thwart any attempt of a single-channel solution to be reliable; the channel agility of TSCH is the key to its ultra-high reliability." Fading is taught in [Multipath, Fading and the Doppler Effect]. Who builds the schedule is not TSCH's business: "How the schedule is built, updated, and maintained, and by which entity, is outside of the scope of the IEEE 802.15.4e standard", which is where IETF's 6TiSCH work comes in.

munotes.in484

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

Compare it with the superframe of [The 802.15.4 Superframe and Guaranteed Time Slots]. Both give links reserved time; the superframe does it on one channel inside a coordinator's active period, TSCH across all sixteen channels for every link in a multi-hop network.

Zigbee and TSCH, computed

The program computes Cskip for the specification's parameters and for the figure's, checking the block recurrence; grows the tree for Cm = 8 and routes three frames along it; tabulates Zigbee's link cost and compares two paths; counts the channels a TSCH cell visits for several slotframe lengths; and delivers packets over a link beside Wi-Fi on one fixed channel, on another, and hopping.

# Zigbee's tree addressing and routing (specification r22, 3.6.1.6 and 3.6.3.3),
# its link cost (3.6.3.1), and the channel hopping of TSCH (RFC 7554, A.6, A.7).
import math
import random

def cskip(d, Cm, Rm, Lm):
    """The size of the address block a parent at depth d gives each router child."""
    if d >= Lm:
        return 0
    if Rm == 1:
        return 1 + Cm * (Lm - d - 1)
    return (1 + Cm - Rm - Cm * Rm ** (Lm - d - 1)) // (1 - Rm)

for Cm, Rm, Lm in ((6, 4, 3), (8, 4, 3)):
    values = [cskip(d, Cm, Rm, Lm) for d in range(Lm + 1)]
    sums = [1 + Rm * values[d + 1] + (Cm - Rm) for d in range(Lm - 1)]
    print("Cm = %d, Rm = %d, Lm = %d: Cskip(0..3) = %s; a router's block holds 1 + Rm x Cskip(d+1) + (Cm - Rm) = %s"
          % (Cm, Rm, Lm, values, sums))

# The tree for Cm = 8, Rm = 4, Lm = 3, the values Figure 3-46 labels its routers with.
Cm, Rm, Lm = 8, 4, 3
C = lambda d: cskip(d, Cm, Rm, Lm)
parent, depth, router = {0: None}, {0: 0}, {0}
def grow(a, d):
    if C(d) == 0:
        return
    for i in range(Rm):                                   # router children, Cskip(d) apart
        child = a + 1 + i * C(d)
        parent[child], depth[child] = a, d + 1
        router.add(child)
        grow(child, d + 1)
    for n in range(1, Cm - Rm + 1):                       # end devices after the routers' blocks
        child = a + Rm * C(d) + n
        parent[child], depth[child] = a, d + 1
grow(0, 0)
kids = lambda a: sorted(x for x, p in parent.items() if p == a)
print("\nThe coordinator (address 0) gives out", kids(0))
print("router 42 (depth 1) gives out", kids(42))
print("router 124 (depth 1) gives out", kids(124), "; router 134 (depth 2):", kids(134))
print("addresses used: %d, the highest %d" % (len(parent), max(parent)))

def next_hop(a, dst):
    """3.6.3.3: to the right child if dst is a descendant of router a, otherwise to
    the parent; an end device has no descendants and always sends to its parent."""
    d = depth[a]
    if a not in router:
        return parent[a]
    below = a == 0 or a < dst < a + C(d - 1)
    if not below:
        return parent[a]
    if dst > a + Rm * C(d):                                # one of a's end devices
        return dst
    return a + 1 + (dst - (a + 1)) // C(d) * C(d)

for src, dst in ((139, 70), (79, 82), (52, 61)):
    path = [src]
    while path[-1] != dst:
        path.append(next_hop(path[-1], dst))
    print("tree route %d to %d: %s, %d hop(s)" % (src, dst, " > ".join(map(str, path)), len(path) - 1))

# Zigbee's link cost: min(7, round(1 / p^4)) for a delivery probability p.
cost = lambda p: min(7, round(1 / p ** 4))
print("\ndelivery probability  0.99  0.95  0.90  0.80  0.70  0.60  0.50")
print("link cost          " + "".join("%6d" % cost(p) for p in (0.99, 0.95, 0.9, 0.8, 0.7, 0.6, 0.5)))
for name, links in (("2 hops at p = 0.7", [0.7, 0.7]), ("3 hops at p = 0.95", [0.95] * 3)):
    print("%-19s path cost %d; expected transmissions %.2f"
          % (name, sum(map(cost, links)), sum(1 / p for p in links)))

# TSCH: a cell (slotOffset t, channelOffset c) in a slotframe of S slots is used
# at ASN = k*S + t and on channel list[(ASN + c) mod 16]. How many channels does
# one cell visit?
print("\nslotframe length S:      16   17   20   24   31  101")
print("channels a cell visits: " + "".join("%5d" % len({(k * S + 3 + 5) % 16 for k in range(64)})
                                           for S in (16, 17, 20, 24, 31, 101)))

# Wi-Fi on 802.11b channel 6 (2426 to 2448 MHz) covers 802.15.4 channels 16 to 19.
# Assume (illustrative) 90% loss there and 2% on the others; one try per slotframe,
# up to 4 tries per packet.
channels = list(range(11, 27))
loss = {ch: 0.9 if 2426 < 2405 + 5 * (ch - 11) < 2448 else 0.02 for ch in channels}
rnd = random.Random(67)
def delivered(pick, packets=20000):
    good = 0
    for _ in range(packets):
        k0 = rnd.randrange(10 ** 6)
        good += any(rnd.random() > loss[pick(k0 + k)] for k in range(4))
    return 100 * good / packets
print("802.15.4 channels under Wi-Fi:", [ch for ch in channels if loss[ch] > 0.5])
print("fixed on channel 17: %.1f%% delivered; fixed on channel 25: %.1f%%; hopping (S = 17): %.1f%%"
      % (delivered(lambda k: 17), delivered(lambda k: 25),
         delivered(lambda k: channels[(k * 17 + 3 + 5) % 16])))
munotes.in485

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

Cm = 6, Rm = 4, Lm = 3: Cskip(0..3) = [31, 7, 1, 0]; a router's block holds 1 + Rm x Cskip(d+1) + (Cm - Rm) = [31, 7]
Cm = 8, Rm = 4, Lm = 3: Cskip(0..3) = [41, 9, 1, 0]; a router's block holds 1 + Rm x Cskip(d+1) + (Cm - Rm) = [41, 9]

The coordinator (address 0) gives out [1, 42, 83, 124, 165, 166, 167, 168]
router 42 (depth 1) gives out [43, 52, 61, 70, 79, 80, 81, 82]
router 124 (depth 1) gives out [125, 134, 143, 152, 161, 162, 163, 164] ; router 134 (depth 2): [135, 136, 137, 138, 139, 140, 141, 142]
addresses used: 169, the highest 168
tree route 139 to 70: 139 > 134 > 124 > 0 > 42 > 70, 5 hop(s)
tree route 79 to 82: 79 > 42 > 82, 2 hop(s)
tree route 52 to 61: 52 > 42 > 61, 2 hop(s)

delivery probability  0.99  0.95  0.90  0.80  0.70  0.60  0.50
link cost               1     1     2     2     4     7     7
2 hops at p = 0.7   path cost 8; expected transmissions 2.86
3 hops at p = 0.95  path cost 3; expected transmissions 3.16

slotframe length S:      16   17   20   24   31  101
channels a cell visits:     1   16    4    2   16   16
802.15.4 channels under Wi-Fi: [16, 17, 18, 19]
fixed on channel 17: 34.4% delivered; fixed on channel 25: 100.0%; hopping (S = 17): 95.8%
munotes.in486

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

Cskip. For the specification's Cm = 6 the values are 31, 7, 1, 0, and each router's block adds up: 1 + 4 × 7 + 2 = 31, and 1 + 4 × 1 + 2 = 7. For Cm = 8 they are 41, 9, 1, 0, with blocks of 41 and 9. Table 3-65's 9 belongs to the second set, its 31 to the first.

The tree. The coordinator hands out routers 1, 42, 83 and 124 and end devices 165 to 168; router 42 hands out 43, 52, 61 and 70 and end devices 79 to 82; router 134, at depth 2, hands out 135 to 142, whose Cskip of 0 makes them childless. The whole tree uses 169 addresses, 0 to 168, out of 65,536: small trees leave most of the address space idle, and deep or wide ones run out, since Cskip grows as a power of Rm.

Tree routes. 139 to 70 climbs to the coordinator and comes down, 5 hops. 79 to 82, two end devices of the same parent, takes 2 hops through router 42: an end device never routes, even to a sibling in range. 52 to 61 goes through their parent, 42.

Link costs. The cost stays at 1 down to p = 0.95, is 2 at 0.9 and 0.8, 4 at 0.7, and hits the cap of 7 at 0.6. The fourth power punishes a lossy link far more than counting expected transmissions would: two hops at p = 0.7 cost 8 against 3 for three hops at p = 0.95, although the expected transmissions (2.86 against 3.16) would slightly favour the two-hop path. Zigbee prefers more, better links.

munotes.in487

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

Why a prime slotframe. A cell in a slotframe of 16 slots always lands on the same channel, since the ASN advances by 16 each cycle. With 20 slots it visits 4 channels, with 24 only 2; with 17, 31 or 101, all 16. The rule underneath is that a cell visits 16 divided by the greatest common divisor of S and 16 channels, so any odd S works; a prime is the RFC's simple way to guarantee it.

Hopping beside Wi-Fi. An 802.11b network on its channel 6 covers 802.15.4 channels 16 to 19. With the illustrative loss rates, a link stuck on channel 17 delivers only 34.4 per cent of packets in four tries; a link fixed on channel 25, well away from the Wi-Fi, delivers all of them; a hopping link delivers 95.8 per cent without knowing which channels are bad. The program's table visits channels in order, so consecutive tries fall on neighbouring channels; a table in a scrambled order would spread them further apart.

Distinctions

Tree (hierarchical) routingMesh routing
How the next hop is foundArithmetic on addresses (Cskip)A routing table filled by route discovery
NeedsDistributed addresses, parent and childrenRoute request and reply, link costs
RouteAlong the tree, often through a common ancestorThe least-cost path found
StrengthNo tables, no discoveryShorter, better routes; survives a failed parent
WeaknessDetours, a single path, address blocks can run outDiscovery traffic, routing tables
Network keyLink key
Shared byAll devices in the networkTwo devices
ProtectsBroadcasts and all network-layer framesUnicast application (APS) frames
Obtained byKey transportKey transport or pre-installation
Zigbee coordinatorZigbee routerZigbee end device
802.15.4 devicePAN coordinatorFFDRFD or FFD
RoutesYesYesNo
Accepts childrenYesYesNo
802.15.4 superframe with GTSsTSCH (802.15.4e, 2012)
ChannelOneAll, hopping by ASN
Reserved timeUp to 7 GTSs, PAN coordinator's links onlyCells for any link in the schedule
Multi-hopSuperframes offset in a cluster treeDesigned for multi-hop schedules
Against fading and interferenceChange channel by handBuilt in

What it does not mean

Zigbee is not 802.15.4. Zigbee is a network and application layer standard from an industry alliance; 802.15.4 is the IEEE radio and MAC under it.

A Zigbee address is not an 802.15.4 extended address. Every device keeps its 64-bit extended address; the 16-bit network address is handed out by its parent and says where it sits in the tree.

munotes.in488

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

Tree routing is not shortest-path routing. It follows parent-child links, which can be far longer than the radio path between two neighbours.

A link cost is not a hop count. Mesh routing may choose three good hops over two poor ones.

A network key is not end-to-end security. Every device holds it; unicast application data needs a link key for that.

TSCH does not change the radio. It is a MAC redesign that runs on the same 802.15.4 PHY.

Quick revision

  • 802.15.4 has no routing; Zigbee's NWK layer adds it. Devices: coordinator (PAN coordinator), router (FFD; routes, accepts children), end device (RFD or FFD; does not route).
  • Distributed addresses: Cm, Rm, Lm; Cskip(d) = 1 + Cm(Lm - d - 1) if Rm = 1, else (1 + Cm - Rm - Cm × Rm^(Lm - d - 1)) / (1 - Rm). Routers at A + 1 + i × Cskip(d), end devices at A + Rm × Cskip(d) + n. Block recurrence: Cskip(d - 1) = 1 + Rm × Cskip(d) + (Cm - Rm).
  • Example Cm 6, Rm 4, Lm 3: 31, 7, 1, 0 (the specification's table prints 9 for 7; its figure uses Cm 8: 41, 9, 1, 0).
  • Tree routing: descendant if A < D < A + Cskip(d - 1); next hop D (end device, D > A + Rm × Cskip(d)) or A + 1 + floor((D - (A + 1)) / Cskip(d)) × Cskip(d); else the parent. Program: 139 to 70 in 5 hops via the coordinator.
  • Mesh routing: route request (broadcast, path cost accumulated), route reply (unicast); link cost = min(7, round(1/p^4)), initially from average LQI; many-to-one routes to a concentrator.
  • Security: NWK and APS layers, AES in CCM* mode; network key (all devices; broadcasts, NWK frames), link keys (pairs; unicast APS); the Trust Center distributes keys; weak moment: a well-known key at first transport; no tamper resistance.
  • 802.15.4e (2012), into 802.15.4-2015: TSCH: synchronised timeslots (about 10 ms), repeating slotframes, cells (slotOffset, channelOffset), ASN = k × S + t, frequency = F((ASN + channelOffset) mod nFreq); hopping beats fading and interference; prime (or odd) slotframe length to visit all 16 channels.

Test yourself

1. Why does MU study IEEE 802.15.4 under routing, when the standard defines no routing? IEEE 802.15.4 specifies only the physical layer and the MAC sublayer, and leaves network formation beyond one hop and multi-hop routing to higher layers. Those layers are built on what the MAC provides: addresses, beacons and association, which form parent-child trees, and link quality indication for each received packet, which measures links. Zigbee's network layer uses these to route, along the association tree or by mesh route discovery with LQI-based link costs, so the standard is the base on which WSN routing is built.

munotes.in489

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

2. Explain Zigbee's distributed address assignment and compute Cskip for Cm = 6, Rm = 4, Lm = 3. The coordinator fixes the maximum children per parent (Cm), the maximum router children (Rm) and the maximum depth (Lm). A parent at depth d gives each router child a block of Cskip(d) addresses: its first router child gets its own address plus 1 and the others follow Cskip(d) apart; end devices get the addresses after the routers' blocks, A + Rm × Cskip(d) + n. With Rm not 1, Cskip(d) = (1 + Cm - Rm - Cm × Rm^(Lm - d - 1)) / (1 - Rm). For Cm = 6, Rm = 4, Lm = 3: Cskip(0) = (3 - 96) / (-3) = 31, Cskip(1) = (3 - 24) / (-3) = 7, Cskip(2) = (3 - 6) / (-3) = 1, and Cskip(3) = 0.

3. How does a Zigbee router decide the next hop in tree routing? A router with address A at depth d checks whether the destination D is its descendant, that is, A < D < A + Cskip(d - 1). If D is not a descendant, it sends the frame to its parent. If D is a descendant and one of its end devices, D > A + Rm × Cskip(d), it sends directly to D; otherwise it sends to the router child whose block contains D, N = A + 1 + floor((D - (A + 1)) / Cskip(d)) × Cskip(d). End devices always send to their parent.

4. Describe Zigbee mesh route discovery and its routing cost. A router needing a route broadcasts a route request. Each router receiving it adds the cost of the link from the device it heard it from to the path cost in the request and rebroadcasts it. The destination replies with a route reply, unicast back along the reverse path, and routers along it record routing table entries, keeping the least-cost route. The path cost is the sum of link costs, each either the constant 7 or min(7, round(1/p^4)) for the link's delivery probability p, with the initial estimates based on average LQI. A sink acting as a concentrator can use many-to-one route discovery to build routes to itself from every router at once.

munotes.in490

Built on 802.15.4: Zigbee Routing, Security and the Later Amendments

5. Compare tree routing and mesh routing in Zigbee. Tree routing uses only address arithmetic, so it needs no routing tables and no discovery and suits routers with little memory; but it follows parent-child links, making detours through common ancestors, offers a single path, and loses a whole subtree when a parent fails. Mesh routing finds and stores routes by discovery, choosing the least-cost path, so routes are shorter and more reliable and survive failures, at the cost of discovery traffic and routing tables.

6. What keys does Zigbee use, and what does the Trust Center do? Zigbee uses a 128-bit network key shared by all devices, which protects broadcasts and all network-layer frames, and 128-bit link keys shared by two devices, which protect unicast application-layer frames. The Trust Center is the device trusted by the others to distribute these keys and to set and update the network's security policies; in a centralised network there is exactly one. Security is applied at the NWK and APS layers using AES in CCM* mode.

7. What is TSCH, and how does it compute the channel for a transmission? TSCH, time-slotted channel hopping, is a MAC mode introduced by the IEEE 802.15.4e amendment of 2012 and included in 802.15.4-2015. All nodes are synchronised; time is divided into timeslots long enough for a maximum-size frame and its acknowledgement, grouped into slotframes that repeat; a schedule assigns each link cells, each a slotOffset and a channelOffset. The absolute slot number counts slots since the network started, ASN = k × S + t, and the frequency used is F((ASN + channelOffset) mod nFreq), where F is a table of the available channels. Because the ASN changes from one slotframe to the next, a link hops across channels even with a fixed schedule, which combats multipath fading and interference.

8. In the program, why did a slotframe of 16 slots make a cell use only one channel, while 17 slots used all 16? The channel index is (ASN + channelOffset) mod 16, and between successive uses of a cell the ASN grows by the slotframe length S. With S = 16 it grows by a multiple of 16, so the index never changes. With S = 17 it grows by 17, which is 1 more than a multiple of 16, so the index advances by one channel each cycle and visits all 16. In general a cell visits 16 divided by the greatest common divisor of S and 16 channels, which is why 20 slots gave 4 channels and 24 gave 2.

Contents This chapter on its own page

munotes.in491

Chapter Sixty-Eight

Traditional Transport Control Protocols: TCP and UDP

Syllabus topic Module 1, "Transport Layer and Middleware in WSN: Traditional transport control protocols"

In one line

A transport protocol carries data between programs, not just between hosts: UDP adds only ports, a length and a checksum to IP and promises nothing more, while TCP sets up a connection with a three-way handshake and then delivers a reliable, in-order byte stream, numbering every byte, acknowledging and retransmitting, letting the receiver limit the flow with its window, and limiting itself with a congestion window that grows by slow start and congestion avoidance and is cut when loss signals congestion.

In the wording a student can write in an examination: the transport layer provides end-to-end communication between application processes, identified by port numbers, over the network layer. UDP (User Datagram Protocol, RFC 768) is connectionless and unreliable: it "provides a procedure for application programs to send messages to other programs with a minimum of protocol mechanism"; its 8-octet header holds source port, destination port, length and checksum, the checksum covering a pseudo header of IP addresses, protocol and length as well as the datagram. TCP (Transmission Control Protocol, RFC 9293) is connection-oriented and provides a reliable, in-order, byte-stream service: a connection is opened by the three-way handshake (SYN, SYN-ACK, ACK); every byte has a sequence number; the receiver returns cumulative acknowledgements; lost or damaged segments are detected (sequence numbers, checksums) and retransmitted; and flow control uses the receiver's advertised window.

Congestion control (RFC 5681) adds the congestion window (cwnd); the sender may have at most the minimum of cwnd and the receiver's window outstanding. In slow start (cwnd below the slow start threshold, ssthresh) cwnd grows by one segment per acknowledgement, doubling every round trip; in congestion avoidance it grows by about one segment per round trip. On a retransmission timeout, ssthresh becomes max(FlightSize / 2, 2 segments) and cwnd drops to 1 segment. On three duplicate ACKs, TCP performs fast retransmit and fast recovery: ssthresh is halved in the same way, cwnd is set to ssthresh + 3 segments while the lost segment is repaired, then deflated to ssthresh, so the window halves instead of collapsing. The result is TCP's sawtooth (additive increase, multiplicative decrease).

What a transport protocol is for

The network layer moves packets from one host to another. A host runs many programs, and they need more than that: a way to tell whose data is whose, and, for most of them, a way to get all the data, in order, without swamping the receiver or the network. Those are the transport layer's four jobs:

  • Multiplexing: port numbers that name the sending and receiving programs. "TCP uses port numbers to identify application services and to multiplex distinct flows between hosts."
  • Error detection: a checksum over each datagram or segment.
  • Reliability and order: detecting loss, retransmitting, and delivering in sequence.
  • Flow and congestion control: not sending faster than the receiver, or the path, can take.
munotes.in492

Traditional Transport Control Protocols: TCP and UDP

UDP does the first two. TCP does all four. The next chapters ask which of them a sensor network needs, and at what cost ([Why a Sensor Network Cannot Simply Run TCP]).

UDP

RFC 768 is three pages long, and its introduction says what the protocol is not. "This protocol provides a procedure for application programs to send messages to other programs with a minimum of protocol mechanism. The protocol is transaction oriented, and delivery and duplicate protection are not guaranteed. Applications requiring ordered reliable delivery of streams of data should use the Transmission Control Protocol (TCP)."

The header is four 16-bit fields, 8 octets in all:

  • Source Port, optional: "If not used, a value of zero is inserted."
  • Destination Port, which names the receiving program.
  • Length: "the length in octets of this user datagram including this header and the data. (This means the minimum value of the length is eight.)"
  • Checksum: "the 16-bit one's complement of the one's complement sum of a pseudo header of information from the IP header, the UDP header, and the data, padded with zero octets at the end (if necessary) to make a multiple of two octets."

The pseudo header is not sent. "The pseudo header conceptually prefixed to the UDP header contains the source address, the destination address, the protocol, and the UDP length. This information gives protection against misrouted datagrams." A datagram delivered to the wrong host fails its checksum there, although its own octets are intact.

There is no connection, no acknowledgement and no retransmission. A datagram arrives once, or twice, or not at all, and the application must cope. The program builds one and checks it.

TCP: the service

RFC 9293 states the service in a sentence: "TCP provides a reliable, in-order, byte-stream service to applications." The application writes a stream of bytes; TCP cuts it into segments, each sent in an IP datagram, and the receiving application reads the same bytes in the same order.

"TCP reliability consists of detecting packet losses (via sequence numbers) and errors (via per-segment checksums), as well as correction via retransmission." And "TCP is connection oriented": two endpoints set up shared state before data flows and tear it down after.

The header is 20 octets without options: source and destination ports; a 32-bit sequence number; a 32-bit acknowledgment number; the data offset (the header's length in 32-bit words, since options may follow); control bits (among them SYN, ACK, FIN and RST); the window; the checksum (over a pseudo header, as in UDP); and the urgent pointer.

munotes.in493

Traditional Transport Control Protocols: TCP and UDP

Opening a connection: the three-way handshake

"The 'three-way handshake' is the procedure used to establish a connection." RFC 9293's Figure 6 gives it with numbers:

  1. A sends SYN with SEQ = 100: "indicating that it will use sequence numbers starting with sequence number 100."
  2. B sends SYN and ACK with SEQ = 300 and ACK = 101: "Note that the acknowledgment field indicates TCP Peer B is now expecting to hear sequence 101, acknowledging the SYN that occupied sequence 100."
  3. A sends ACK with SEQ = 101 and ACK = 301. Both are now ESTABLISHED.

A's first data segment also carries SEQ = 101, because "the ACK does not occupy sequence number space (if it did, we would wind up ACKing ACKs!)." The SYN consumes one sequence number; a pure acknowledgement consumes none.

Why three messages? "The principal reason for the three-way handshake is to prevent old duplicate connection initiations from causing confusion." Each side must see its own starting number acknowledged before it trusts the connection, and that takes a question, an answer with a question, and an answer.

Sequence numbers, acknowledgements and the receive window

TCP numbers bytes, not segments. A segment carrying 100 octets from sequence number 101 covers 101 to 200, and the receiver answers ACK = 201: the next byte it expects. Acknowledgements are cumulative: ACK = 401 says everything up to 400 has arrived, so a lost acknowledgement is repaired by the next one.

When a segment is lost, the receiver keeps asking for the same byte: a duplicate acknowledgement for each later segment that arrives. The sender retransmits either after a retransmission timeout or, sooner, after three duplicate acknowledgements (fast retransmit, below).

Flow control. Every segment carries the receiver's window: "The number of data octets beginning with the one indicated in the acknowledgment field that the sender of this segment is willing to accept." A receiver with a full buffer advertises a small window, or zero, and the sender must wait. This protects the receiver; it does nothing for the network between them.

Congestion control: the congestion window

Congestion control protects the network. RFC 5681 adds a second limit, kept by the sender: "The congestion window (cwnd) is a sender-side limit on the amount of data the sender can transmit into the network before receiving an acknowledgment (ACK), while the receiver's advertised window (rwnd) is a receiver-side limit on the amount of outstanding data. The minimum of cwnd and rwnd governs data transmission."

A third variable chooses how cwnd grows: the slow start threshold, ssthresh. "The slow start algorithm is used when cwnd < ssthresh, while the congestion avoidance algorithm is used when cwnd > ssthresh."

munotes.in494

Traditional Transport Control Protocols: TCP and UDP

The initial window. A new connection starts with IW segments: 4 if the sender's maximum segment size (SMSS) is at most 1095 bytes, 3 if it is up to 2190, and 2 above that. ssthresh starts "arbitrarily high".

Slow start. "Beginning transmission into a network with unknown conditions requires TCP to slowly probe the network to determine the available capacity". During slow start, "a TCP increments cwnd by at most SMSS bytes for each ACK received that cumulatively acknowledges new data." A window of w segments brings back w acknowledgements in one round trip, so cwnd doubles every round trip: slow only at the start.

Congestion avoidance. Above ssthresh, cwnd grows by about one segment per round trip; one common way is cwnd += SMSS × SMSS / cwnd on each acknowledgement, "an acceptable approximation to the underlying principle of increasing cwnd by 1 full-sized segment per RTT."

A timeout. When the retransmission timer finds a loss, ssthresh is set to no more than max(FlightSize / 2, 2 × SMSS), "where, as discussed above, FlightSize is the amount of outstanding data in the network", and on a timeout "cwnd MUST be set to no more than the loss window, LW, which equals 1 full-sized segment (regardless of the value of IW)." TCP then slow-starts back up to the new ssthresh.

Fast retransmit and fast recovery. Three duplicate acknowledgements are a milder signal. "The fast retransmit algorithm uses the arrival of 3 duplicate ACKs ... as an indication that a segment has been lost." TCP retransmits at once, sets ssthresh as for a timeout, and sets cwnd to ssthresh plus 3 segments, which "artificially 'inflates' the congestion window by the number of segments (three) that have left the network and which the receiver has buffered." Each further duplicate adds one segment. When the retransmission is acknowledged, cwnd is set to ssthresh: "This is termed 'deflating' the window."

Why not slow-start after duplicate acknowledgements too? "since the receiver can only generate a duplicate ACK when a segment has arrived, that segment has left the network and is in the receiver's buffer, so we know it is no longer consuming network resources." Data is still flowing; only one segment is missing. So the window is halved, not reset.

UDP and TCP, computed

The program builds a UDP datagram of 9 data octets, computes its checksum over the pseudo header, and shows the receiver's check passing and failing; replays Figure 6's handshake and three data segments; and follows cwnd through forty round trips over a path that holds 24 segments, with one timeout placed at round trip 30.

munotes.in495

Traditional Transport Control Protocols: TCP and UDP

# UDP and TCP from their RFCs: a UDP datagram and its checksum (RFC 768), the
# three-way handshake of RFC 9293's Figure 6, and the congestion window of RFC
# 5681 round trip by round trip.
import struct

def ones_sum(data):
    """The 16-bit one's complement sum of RFC 768: carries wrap around."""
    if len(data) % 2:
        data += b"\0"
    total = 0
    for (word,) in struct.iter_unpack("!H", data):
        total += word
        total = (total & 0xFFFF) + (total >> 16)
    return total

src, dst, payload = bytes([10, 0, 0, 1]), bytes([10, 0, 0, 2]), b"temp=24.5"
length = 8 + len(payload)
pseudo = src + dst + bytes([0, 17]) + struct.pack("!H", length)   # addresses, protocol 17, length
checksum = 0xFFFF - ones_sum(pseudo + struct.pack("!HHHH", 5000, 6000, length, 0) + payload)
datagram = struct.pack("!HHHH", 5000, 6000, length, checksum) + payload
print("UDP datagram of %d octets: header %s, then %r" % (length, datagram[:8].hex(" "), payload))
print("the receiver's sum over pseudo header and datagram: 0x%04X" % ones_sum(pseudo + datagram))
damaged = datagram[:10] + b"5" + datagram[11:]
print("the same with one data octet changed:              0x%04X" % ones_sum(pseudo + damaged))

print("\nThe three-way handshake (RFC 9293, Figure 6), then three 100-octet segments:")
print("  A -> B  SYN      SEQ=100")
print("  B -> A  SYN,ACK  SEQ=300  ACK=101")
print("  A -> B  ACK      SEQ=101  ACK=301")
seq = 101
for _ in range(3):
    print("  A -> B  DATA     SEQ=%d, 100 octets;  B -> A  ACK=%d" % (seq, seq + 100))
    seq += 100

def cwnd_trace(rounds=40, capacity=24, iw=3, timeout_round=30):
    """RFC 5681, one entry per round trip, windows in full-sized segments. Slow
    start adds a segment per ACK, doubling cwnd each round trip; congestion
    avoidance adds one segment per round trip. A round trip that sends more
    than `capacity` segments overflows the path, and the loss is found by three
    duplicate ACKs (one loss event per such round trip, in this model): ssthresh
    becomes half the flight (equation 4) and, after fast recovery, cwnd deflates
    to ssthresh. At `timeout_round` the loss is found by the retransmission
    timer instead: cwnd falls to the loss window, 1, and slow start begins again."""
    cwnd, ssthresh, trace = iw, 10 ** 6, []
    for r in range(1, rounds + 1):
        event = "timeout" if r == timeout_round else "3 dup ACKs" if cwnd > capacity else ""
        trace.append((r, cwnd, ssthresh, event))
        if event:
            ssthresh = max(cwnd // 2, 2)
            cwnd = 1 if event == "timeout" else ssthresh
        elif cwnd < ssthresh:
            cwnd = min(2 * cwnd, ssthresh)
        else:
            cwnd += 1
    return trace

trace = cwnd_trace()
print("\nCongestion window by round trip (IW 3 segments, the path holds 24):")
for start in range(0, 40, 10):
    print("  rounds %2d to %2d: %s" % (start + 1, start + 10, " ".join("%3d" % t[1] for t in trace[start:start + 10])))
for r, cwnd, ssthresh, event in trace:
    if event:
        print("  round %2d: cwnd %2d, %-10s -> ssthresh %2d, cwnd %2d"
              % (r, cwnd, event, max(cwnd // 2, 2), 1 if event == "timeout" else max(cwnd // 2, 2)))
carried = sum(min(t[1], 24) for t in trace)
print("segments the path could carry: %d in 40 round trips, %.1f per round trip of a possible 24"
      % (carried, carried / 40))
munotes.in496

Traditional Transport Control Protocols: TCP and UDP

UDP datagram of 17 octets: header 13 88 17 70 00 11 38 9b, then b'temp=24.5'
the receiver's sum over pseudo header and datagram: 0xFFFF
the same with one data octet changed:              0xC7FF

The three-way handshake (RFC 9293, Figure 6), then three 100-octet segments:
  A -> B  SYN      SEQ=100
  B -> A  SYN,ACK  SEQ=300  ACK=101
  A -> B  ACK      SEQ=101  ACK=301
  A -> B  DATA     SEQ=101, 100 octets;  B -> A  ACK=201
  A -> B  DATA     SEQ=201, 100 octets;  B -> A  ACK=301
  A -> B  DATA     SEQ=301, 100 octets;  B -> A  ACK=401

Congestion window by round trip (IW 3 segments, the path holds 24):
  rounds  1 to 10:   3   6  12  24  48  24  25  12  13  14
  rounds 11 to 20:  15  16  17  18  19  20  21  22  23  24
  rounds 21 to 30:  25  12  13  14  15  16  17  18  19  20
  rounds 31 to 40:   1   2   4   8  10  11  12  13  14  15
  round  5: cwnd 48, 3 dup ACKs -> ssthresh 24, cwnd 24
  round  7: cwnd 25, 3 dup ACKs -> ssthresh 12, cwnd 12
  round 21: cwnd 25, 3 dup ACKs -> ssthresh 12, cwnd 12
  round 30: cwnd 20, timeout    -> ssthresh 10, cwnd  1
segments the path could carry: 609 in 40 round trips, 15.2 per round trip of a possible 24
A graph of the congestion window in segments against the round trip, from 1 to 40, with a dashed line at 24 segments marking what the path holds. The window starts at 3 and doubles to 48 by round trip 5, overshooting; three duplicate acknowledgements halve it to 24; it rises by one to 25 and is halved again to 12; it then climbs one segment per round trip to 25 and is halved again, a sawtooth; at round trip 30 a timeout drops it to 1, it doubles to 8, reaches the new threshold of 10 and climbs by one again

Figure 68.1 The congestion window of the program, round trip by round trip

The checksum. The datagram is 17 octets: ports 5000 (0x1388) and 6000 (0x1770), length 17 (0x0011) and checksum 0x389B, then the 9 data octets. The receiver adds up the pseudo header and the whole datagram, checksum included, and gets 0xFFFF, all ones, because the checksum was chosen as the one's complement of everything else. Change one data octet and the sum is 0xC7FF: the datagram is discarded. A one's complement sum catches any single changed octet, though some combinations of changes can cancel out.

The handshake. The numbers are Figure 6's: A's SYN takes sequence number 100, B's takes 300, and the third segment and the first data segment both carry 101. The three data segments cover bytes 101 to 400, and each acknowledgement names the next byte expected: 201, 301, 401.

munotes.in497

Traditional Transport Control Protocols: TCP and UDP

The sawtooth. From an initial window of 3, slow start doubles cwnd each round trip: 3, 6, 12, 24 and then 48, twice what the path holds; slow start overshoots because it doubles until something is lost. Three duplicate acknowledgements set ssthresh to 24 and cwnd to 24; the next increase overflows again and halves the window to 12. From there congestion avoidance adds one segment per round trip, 12 to 25 in thirteen round trips, and the loss at 25 halves it again: additive increase, multiplicative decrease. The timeout at round trip 30, with cwnd at 20, is harsher: ssthresh becomes 10 and cwnd 1, and slow start climbs 1, 2, 4, 8, to the threshold, 10, then one at a time.

What it costs. Over forty round trips the windows would fill the path for 609 segments, 15.2 per round trip of a possible 24. TCP keeps probing for capacity and backing off, and pays for its caution. Everything here assumes a loss means congestion, the assumption the next chapter tests on a wireless link.

Distinctions

UDPTCP
ConnectionNoneThree-way handshake, then state at both ends
UnitDatagram (message)Byte stream, cut into segments
ReliabilityNone: delivery and duplicate protection not guaranteedSequence numbers, acknowledgements, retransmission
OrderNot keptIn order
Flow controlNoneReceiver's window
Congestion controlNoneCongestion window (RFC 5681)
Header8 octets20 octets without options
ChecksumOver pseudo header, header and dataThe same, for segments
Flow controlCongestion control
ProtectsThe receiverThe network between them
Limitrwnd, advertised by the receivercwnd, kept by the sender
SignalThe window fieldLoss: a timeout or three duplicate ACKs
Slow startCongestion avoidance
Whencwnd below ssthresh (start, after a timeout)cwnd above ssthresh
GrowthOne segment per ACK: doubles each round tripAbout one segment per round trip
Program3, 6, 12, 24, 48; after the timeout 1, 2, 4, 8, 1012, 13, ..., 25
Retransmission timeoutThree duplicate ACKs
MeansNothing is getting throughOne segment lost, later ones arriving
ssthreshmax(FlightSize / 2, 2 segments)The same
cwnd1 segment, then slow startssthresh + 3 during fast recovery, then ssthresh

What it does not mean

Unreliable does not mean useless. UDP suits requests and answers that fit in one datagram, and data where a late copy is worthless; the application adds whatever reliability it needs.

TCP's reliability is not the network's. IP may still lose, duplicate or reorder packets; TCP hides that from the application at the ends.

Slow start is not slow. It doubles the window every round trip; it is only slow compared with sending the whole window at once.

munotes.in498

Traditional Transport Control Protocols: TCP and UDP

The receive window is not the congestion window. One protects the receiver, the other the network; the sender obeys the smaller.

A duplicate acknowledgement is not a retransmission request for everything. It repeats the next byte expected, because a later segment arrived first.

Loss is not always congestion. TCP treats it as congestion; on a wired path that is usually right, and on a radio link often wrong.

Quick revision

  • Transport layer: end-to-end between programs, ports; jobs: multiplexing, error detection, reliability and order, flow and congestion control.
  • UDP (RFC 768): connectionless, no guarantees; header 8 octets: source port (optional, 0 if unused), destination port, length (at least 8), checksum (one's complement over pseudo header + header + data).
  • TCP (RFC 9293): reliable, in-order, byte stream, connection oriented; header 20 octets + options: ports, sequence, acknowledgment, data offset, control bits, window, checksum, urgent pointer.
  • Three-way handshake: SYN (SEQ 100) / SYN-ACK (SEQ 300, ACK 101) / ACK (SEQ 101, ACK 301); prevents old duplicate connections; the SYN uses a sequence number, a pure ACK does not.
  • Bytes numbered; cumulative ACK = next byte expected; duplicate ACKs after a loss; flow control by the receiver's window.
  • Congestion control (RFC 5681): send at most min(cwnd, rwnd); IW 2 to 4 segments; slow start below ssthresh (+1 segment per ACK, doubles per RTT); congestion avoidance above (+1 segment per RTT); timeout: ssthresh = max(FlightSize / 2, 2 SMSS), cwnd = 1; 3 duplicate ACKs: fast retransmit, fast recovery (cwnd = ssthresh + 3, then deflate to ssthresh).
  • Program: checksum sum 0xFFFF when intact, 0xC7FF when one octet changed; cwnd 3, 6, 12, 24, 48, then a sawtooth 12 to 25; timeout at 20: back to 1, threshold 10; 15.2 of 24 segments per round trip.

Test yourself

1. What services does the transport layer provide, and which does each of UDP and TCP provide? The transport layer carries data end to end between application processes, identified by port numbers, and may add error detection, reliable in-order delivery, flow control and congestion control. UDP provides multiplexing by ports and a checksum, with no connection, no reliability, no ordering and no flow or congestion control. TCP provides all of them: ports, checksums, a connection, reliable in-order byte-stream delivery by sequence numbers, acknowledgements and retransmission, flow control by the receiver's window, and congestion control by the congestion window.

2. Draw the UDP header and explain each field. The UDP header has four 16-bit fields: source port, the sending process's port, optional and zero if unused; destination port, the receiving process's port; length, the length in octets of the header and data, at least 8; and checksum, the 16-bit one's complement of the one's complement sum of a pseudo header (source and destination IP addresses, protocol and UDP length), the UDP header and the data, padded to a multiple of two octets. The pseudo header protects against misrouted datagrams.

munotes.in499

Traditional Transport Control Protocols: TCP and UDP

3. Explain TCP's three-way handshake. The initiating peer sends a SYN segment carrying its initial sequence number, for example 100. The other peer replies with a segment that carries its own SYN, for example sequence number 300, and acknowledges the first SYN with acknowledgment number 101, the next sequence number it expects. The initiator then acknowledges that SYN with acknowledgment number 301, using sequence number 101, and both sides are established; data can then flow, starting at sequence number 101, since a pure acknowledgement uses no sequence number. The main reason for three messages is to prevent old duplicate connection requests from being mistaken for new ones.

4. How does TCP achieve reliable, in-order delivery? TCP numbers every byte of the stream, and each segment carries the sequence number of its first byte. The receiver returns cumulative acknowledgements naming the next byte it expects, uses the sequence numbers to put segments in order and discard duplicates, and uses the checksum to discard damaged segments. The sender retransmits a segment when its acknowledgement does not arrive before the retransmission timer expires, or when three duplicate acknowledgements show that it was lost while later segments arrived.

5. Distinguish flow control from congestion control in TCP. Flow control prevents the sender from overrunning the receiver: the receiver advertises in every segment a window, the number of octets beyond the acknowledged one that it can accept. Congestion control prevents the sender from overloading the network: the sender keeps a congestion window, grown by slow start and congestion avoidance and reduced when loss is detected. The sender may have outstanding at most the minimum of the two windows.

6. Explain slow start and congestion avoidance. In slow start, used at the start of a connection and after a retransmission timeout, while the congestion window is below the slow start threshold, the sender increases the window by up to one segment for each acknowledgement of new data, so the window doubles every round trip. Once the window exceeds the threshold, congestion avoidance increases it by about one segment per round trip, for example by SMSS × SMSS / cwnd per acknowledgement. The threshold starts high and is reduced whenever loss shows congestion.

7. How does TCP react to a timeout, and how to three duplicate acknowledgements? On a retransmission timeout, the sender sets the slow start threshold to at most max(FlightSize / 2, 2 × SMSS), sets the congestion window to one segment, retransmits and slow-starts back up. On the third duplicate acknowledgement, it performs fast retransmit, resending the missing segment at once, and fast recovery: it sets the threshold the same way, sets the window to the threshold plus three segments, adds one segment for each further duplicate, and when new data is acknowledged sets the window to the threshold. Duplicate acknowledgements show that later segments are still arriving, so the window is halved rather than reset.

munotes.in500

Traditional Transport Control Protocols: TCP and UDP

8. In the program, why did the window reach 48 segments on a path that holds 24? In slow start the window doubles every round trip, and the threshold starts arbitrarily high, so nothing stops the doubling until a loss is detected. The window went 3, 6, 12, 24; 24 fitted the path, so the next round trip doubled it to 48, and only then did the overflow cause a loss and three duplicate acknowledgements, which set the threshold to 24. Slow start always overshoots in this way, and the losses of that round trip are the price of probing a path of unknown capacity.

Contents This chapter on its own page

munotes.in501

Chapter Sixty-Nine

Why a Sensor Network Cannot Simply Run TCP

Syllabus topic Module 1, "Transport Layer and Middleware in WSN: Traditional transport control protocols, Transport protocol design issues in WSNs"

In one line

TCP was built for wired paths where loss means congestion, so on a multi-hop radio path it halves its window at every bit error, recovers end to end although errors multiply hop by hop, spends a handshake and 20-octet headers on readings of a few octets, and delivers every byte of one flow to one address, while a sensor network wants hop-by-hop recovery, event-level reliability, little overhead and data named by what it is.

In the wording a student can write in an examination: TCP is unsuitable for wireless sensor networks for several reasons. (1) Loss is taken for congestion. TCP assumes packet loss is caused by congestion and responds by reducing its congestion window, but in sensor networks most losses come from radio errors, collisions and fading, so TCP needlessly throttles itself. (2) End-to-end recovery over many hops. With a per-hop error rate p, a packet crosses n hops with probability (1 - p)^n, which falls exponentially, so end-to-end retransmission costs many transmissions over the whole path, while hop-by-hop recovery, with a cache at each node, repairs a loss where it happened. (3) Overhead. The three-way handshake, the connection close and 20-octet headers (with 40 octets of IPv6) are large compared with small, frequent sensor readings and a frame of at most 127 octets, and every extra transmission costs energy. (4) A different model. Sensor traffic is mostly many-to-one, toward a sink, driven by events, and data-centric rather than addressed to hosts; the sink needs event reliability (enough reports about an event), not every packet from every node; and TCP's per-connection state and acknowledgement traffic burden small nodes and a shared channel. (5) Fairness and energy across many sources are not TCP's concern. Sensor networks therefore end TCP at the sink (TCP splitting) and use lightweight, hop-by-hop, energy-aware transport protocols inside.

Loss is not always congestion

The last chapter's congestion control rests on one assumption. RMST's authors state it plainly: "Traditional transport layers, like TCP, assume that the primary cause of packet loss is congestion. As such, their focus is on congestion control. In sensor nets the primary problem is packet loss due to interference or low power."

On a wired path the assumption is sound: links rarely corrupt bits, and a lost packet was almost always dropped by a full router queue. On a radio link, frames are lost to collisions ([Hidden and Exposed Terminals, and RTS and CTS]), to interference and to fading ([Multipath, Fading and the Doppler Effect]) whatever the load. TCP cannot tell the difference. Every loss, detected by a timeout or three duplicate acknowledgements, halves its window or resets it to one segment ([Traditional Transport Control Protocols: TCP and UDP]), so a lossy but idle path ends up carried at a crawl. The program measures how fast the window falls.

munotes.in502

Why a Sensor Network Cannot Simply Run TCP

Errors multiply over many hops

The second problem is geometry. A sensor network is multi-hop, and PSFQ's authors show what that does to recovery done only at the ends. "Despite the various differences in the communication and service model, the biggest problem with end-to-end recovery has to do with the physical characteristic of the transport medium: sensor networks usually operate in harsh radio environments, and rely on multi-hop forwarding techniques to exchange messages. Error accumulates exponentially over multihops."

With a per-hop error rate p, one hop succeeds with probability 1 - p, and n hops with (1 - p)^n. PSFQ's conclusion: "for larger network it is almost impossible to deliver a single message using an end-to-end approach in a lossy link environment when the error rate is larger than 20%." Wired networks and wireless LANs, with error rates under 1 per cent, manage well above 90 per cent; low-power sensor radios "cannot rely on using high power to boost the link reliability when operating under harsh radio conditions."

The remedy is to recover at each hop. "We propose hop-by-hop error recovery in which intermediate nodes also take responsibility for loss detection and recovery so reliable data exchange is done on a hop-by-hop manner rather than an end-to-end one." A loss is then repaired across the one hop where it happened, and the chance of delivery no longer falls with the length of the path.

RMST separates two places where hops can retry. At the MAC, a hop that retries up to R times succeeds with 1 - (1 - s)^R, where s is the chance of one try, and the whole path with that raised to the number of hops. At the transport layer, nodes can cache fragments and repair them hop by hop. RMST's finding joins them: "The loss rate presented to the transport layer by the MAC layer needs to get below one percent for the advantages of caching and hop-by-hop repair to be marginalized." The program computes both.

A handshake and headers for a few octets

A sensor reading is often a few octets: a temperature, a count, an alarm. TCP wraps it in a connection. RFC 9293's handshake takes three segments before the first byte of data (Figure 6), and its normal close four more (Figure 12), each with a 20-octet TCP header inside an IP header.

The frame has little room to give. RFC 4919, which set the goals for IPv6 over 802.15.4, did the sum: "Given that in the worst case the maximum size available for transmitting IP packets over an IEEE 802.15.4 frame is 81 octets, and that the IPv6 header is 40 octets long, (without optional headers), this leaves only 41 octets for upper-layer protocols, like UDP and TCP. UDP uses 8 octets in the header and TCP uses 20 octets. This leaves 33 octets for data over UDP and 21 octets for data over TCP." Every octet on the air is energy spent by a battery ([MAC Protocols for Sensor Networks: The Job and Where the Energy Goes]).

munotes.in503

Why a Sensor Network Cannot Simply Run TCP

The Internet's own protocol for constrained devices chose UDP. CoAP is "bound to unreliable transports such as UDP" and adds "a lightweight reliability mechanism, without trying to re-create the full feature set of a transport like TCP" ([Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP]).

A different kind of traffic

Even with a perfect channel, TCP answers a different question from the one a sensor network asks.

  • What must be reliable. ESRT's authors: "Reliable event detection at the sink is based on collective information provided by source nodes and not on any individual report. Hence, conventional end-to-end reliability definitions and solutions are inapplicable in the WSN regime and would only lead to a waste of scarce sensor resources." Twenty nodes reporting one fire need not each deliver every packet; the sink needs enough of them.
  • Who talks to whom. TCP joins two addressed endpoints. In a sensor network data flows mostly from many sources to one sink, in bursts when events happen, and is often named by its attributes, not by host addresses. The 2002 survey: "Unlike protocols such as TCP, the end-to-end communication schemes in sensor networks are not based on global addressing." ([Design Principles: Data Centricity, Location, Activity and Heterogeneity])
  • Which way reliability is needed. Readings flowing up can often tolerate loss; code and commands flowing down to every node cannot lose a byte, and must reach many receivers at once, which one TCP connection cannot do.
  • State and acknowledgements. Each TCP connection keeps buffers, timers and windows at both ends, costly in a few kilobytes of memory, and its acknowledgements share the same channel as the data, adding to the contention they are meant to measure.
  • Fairness and energy. Many sources must share the path to the sink fairly, and the protocol should spend as little energy as possible; neither is part of TCP's design.

What sensor networks do instead

The survey's answer for the boundary with the Internet is to split TCP at the sink: "An approach such as TCP splitting may be needed to make sensor networks interact with other networks such as the Internet. In this approach, TCP connections are ended at sink nodes, and a special transport layer protocol can handle the communications between the sink node and sensor nodes". Users reach the sink by TCP or UDP over the Internet; inside, the sink and nodes use protocols built for the job. [Transport Protocol Design Issues in WSNs] sets out what those must decide, and [Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST] shows four that answer it.

munotes.in504

Why a Sensor Network Cannot Simply Run TCP

The misfit, computed

The program tabulates PSFQ's end-to-end success for five error rates and five path lengths; reproduces RMST's Figure 1 (40 hops at a 0.10 error rate, with R tries per hop) and Table 1 (link transmissions for M fragments over H hops, 0.9 per link), adding the cost when a lost fragment stops where it was lost; counts the octets UDP and TCP spend on one 10-octet reading with uncompressed IPv6 headers; and runs TCP's congestion avoidance on a five-hop path with no congestion at all, only radio losses.

# Why TCP fits a multi-hop sensor network badly, computed. PSFQ's end-to-end
# success (1 - p)^n, RMST's retries per hop and its Table 1, the octets TCP
# spends on one reading, and TCP's window under losses that are not congestion.
import random

print("End-to-end success over n hops, per-hop error rate p (PSFQ, Figure 1):")
print("   p      1 hop   5 hops  10 hops  20 hops  40 hops")
for p in (0.01, 0.05, 0.10, 0.20, 0.30):
    print("  %.2f " % p + "".join("%8.3f" % (1 - p) ** n for n in (1, 5, 10, 20, 40)))

# RMST, equations 2 and 3: R tries per hop succeed with 1 - (1 - s)^R, and all
# H hops with that to the power H. Its Figure 1: 40 hops, error rate 0.10.
s = 0.9
print("\n40 hops at error 0.10, R tries per hop (RMST, Figure 1):",
      ", ".join("R = %d: %.3f" % (R, (1 - (1 - s) ** R) ** 40) for R in range(1, 6)))

# RMST's Table 1: link transmissions to move M fragments over H hops at 0.9 per
# link. With a cache at every hop each fragment costs H / s. End to end, an
# attempt that fails must start again from the source; charging each attempt the
# whole path gives H / s^H per fragment (Table 1's figures), and stopping a lost
# fragment where it was lost gives (1 - s^H) / ((1 - s) s^H).
print("\nLink transmissions for M fragments at 0.9 per link:")
print("   M  hops   hop by hop   end to end (Table 1)   end to end (lost stops)")
for M in (5, 10, 20):
    for H in (5, 10):
        print("  %2d  %3d  %11.2f  %21.2f  %22.2f"
              % (M, H, M * H / s, M * H / s ** H, M * (1 - s ** H) / ((1 - s) * s ** H)))

# One 10-octet reading, IPv6 headers uncompressed (40 octets): UDP sends one
# datagram; TCP opens, sends, is acknowledged and closes (RFC 9293 Figures 6, 12).
udp = 40 + 8 + 10
tcp_segments = ["SYN", "SYN,ACK", "ACK", "DATA", "ACK", "FIN,ACK", "ACK", "FIN,ACK", "ACK"]
tcp = sum(40 + 20 + (10 if kind == "DATA" else 0) for kind in tcp_segments)
print("\nOne 10-octet reading: UDP %d octets in 1 datagram; TCP %d octets in %d segments (%.1f times)"
      % (udp, tcp, len(tcp_segments), tcp / udp))

# TCP in congestion avoidance on a path with no congestion at all: each segment
# is lost with probability q per hop over 5 hops, and every loss halves cwnd as
# if it were congestion (fast recovery); otherwise cwnd grows by 1 per round trip.
rnd = random.Random(69)
print("\nAverage cwnd (segments per round trip) with only radio losses, 5 hops, 2000 round trips:")
for q in (0.0, 0.001, 0.01, 0.02, 0.05, 0.10):
    lost_end_to_end = 1 - (1 - q) ** 5
    cwnd, total = 1.0, 0.0
    for _ in range(2000):
        total += cwnd
        if any(rnd.random() < lost_end_to_end for _ in range(int(cwnd))):
            cwnd = max(cwnd / 2, 1.0)
        else:
            cwnd = min(cwnd + 1, 64)
    print("  per-hop loss %.3f (end to end %.3f): %5.1f" % (q, lost_end_to_end, total / 2000))
munotes.in505

Why a Sensor Network Cannot Simply Run TCP

End-to-end success over n hops, per-hop error rate p (PSFQ, Figure 1):
   p      1 hop   5 hops  10 hops  20 hops  40 hops
  0.01    0.990   0.951   0.904   0.818   0.669
  0.05    0.950   0.774   0.599   0.358   0.129
  0.10    0.900   0.590   0.349   0.122   0.015
  0.20    0.800   0.328   0.107   0.012   0.000
  0.30    0.700   0.168   0.028   0.001   0.000

40 hops at error 0.10, R tries per hop (RMST, Figure 1): R = 1: 0.015, R = 2: 0.669, R = 3: 0.961, R = 4: 0.996, R = 5: 1.000

Link transmissions for M fragments at 0.9 per link:
   M  hops   hop by hop   end to end (Table 1)   end to end (lost stops)
   5    5        27.78                  42.34                   34.68
   5   10        55.56                 143.40                   93.40
  10    5        55.56                  84.68                   69.35
  10   10       111.11                 286.80                  186.80
  20    5       111.11                 169.35                  138.70
  20   10       222.22                 573.59                  373.59

One 10-octet reading: UDP 58 octets in 1 datagram; TCP 550 octets in 9 segments (9.5 times)

Average cwnd (segments per round trip) with only radio losses, 5 hops, 2000 round trips:
  per-hop loss 0.000 (end to end 0.000):  63.0
  per-hop loss 0.001 (end to end 0.005):  19.4
  per-hop loss 0.010 (end to end 0.049):   5.9
  per-hop loss 0.020 (end to end 0.096):   4.0
  per-hop loss 0.050 (end to end 0.226):   2.6
  per-hop loss 0.100 (end to end 0.410):   1.8
munotes.in506

Why a Sensor Network Cannot Simply Run TCP

End to end, hop after hop. At 1 per cent per hop, even 40 hops deliver 66.9 per cent; at 10 per cent, 10 hops deliver 34.9 per cent and 40 hops 1.5 per cent; at 20 per cent, 20 hops deliver 1.2 per cent, which is PSFQ's "almost impossible". The exponent is the path length, so every extra hop multiplies the loss.

Retries at each hop. Over 40 hops at a 0.10 error rate, one try per hop delivers 1.5 per cent; two tries, 66.9 per cent; three, 96.1 per cent; five, practically all. RMST's conclusion that "the use of at least 3 retries is vital to reliable data delivery in this scenario" is the table's third column.

What end-to-end recovery costs. With a cache at each hop, moving a fragment costs H / 0.9 link transmissions: 27.78 for 5 fragments over 5 hops, doubling with the hops or the fragments. End to end, a failed attempt starts again from the source. Charging each attempt the whole path gives RMST's printed figures (42.33, 143.39, 84.67, 286.79, 169.35, 573.59, which are the program's values truncated): over 10 hops, 2.6 times the hop-by-hop cost, because 0.9^10 is only 0.349. Even letting a lost fragment stop where it was lost, end to end costs 93.40 against 55.56 over 10 hops. The gap widens with every hop.

One reading by TCP. Sending 10 octets costs UDP one 58-octet datagram. TCP opens (three segments), sends the data, has it acknowledged and closes (four segments): nine segments and 550 octets, 9.5 times as much, before counting any retransmission. Header compression ([Sensor Networks in the Internet of Things: 6LoWPAN, RPL and CoAP]) shrinks the headers but not the seven segments that carry no data.

Losses read as congestion. With no loss the window grows to its limit of 64 segments per round trip. A per-hop loss of just 0.1 per cent (0.5 per cent over five hops) cuts the average window to 19.4 segments; 1 per cent per hop to 5.9; 10 per cent per hop to 1.8. The path was never congested: every halving was a response to a radio error that more caution could not prevent. Recovering losses at the hop where they happen, and signalling congestion explicitly instead of inferring it from loss, are the two answers the next chapters build on.

Distinctions

End-to-end recoveryHop-by-hop recovery
Who detects and repairs lossOnly the two end pointsEvery node along the path
Delivery over n hops at error pFalls as (1 - p)^nEach hop repaired, so it no longer falls with n
Cost of one lossA resend over the whole pathA resend over one hop
NeedsState at the endsA cache at each node
ExampleTCPPSFQ, RMST's caching mode
munotes.in507

Why a Sensor Network Cannot Simply Run TCP

What TCP assumesWhat a sensor network has
Why packets are lostCongestionMostly radio errors, collisions, fading
PathFew, reliable hopsMany lossy hops
TrafficA byte stream between two addressed hostsShort readings, many sources to one sink, event bursts
What must be reliableEvery byteOften the event, not each report; code and commands, every byte
ResourcesPlenty of memory, energy is freeKilobytes of memory, energy is everything

What it does not mean

TCP is not useless near sensor networks. It works well from the sink outward, over the Internet; the survey's TCP splitting keeps it there.

Hop-by-hop recovery is not free. Every node must cache what it forwards and track what is missing; RMST found the benefit shrinks once the MAC's own retries bring losses below 1 per cent.

MAC retries do not replace transport recovery. They make each hop better, but a node can still fail or drop a packet from a full buffer, so end-to-end or hop-by-hop recovery above the MAC remains.

A sensor network does not always need less reliability. Readings can tolerate loss; a program image sent to every node cannot lose a byte, which is PSFQ's case.

UDP is not the answer on its own. It removes TCP's costs and its guarantees together; sensor transport protocols add back only what their data needs.

Quick revision

  • TCP assumes loss = congestion and cuts its window; in WSNs loss is mostly interference, low power, collisions, fading (RMST).
  • End to end over n hops: (1 - p)^n, exponential decay; PSFQ: almost impossible above 20 per cent error on long paths. Answer: hop-by-hop recovery with caches.
  • RMST: MAC tries per hop, 1 - (1 - s)^R per hop; 40 hops at 0.10: 0.015 / 0.669 / 0.961 for R = 1 / 2 / 3; transport caching matters until MAC losses are below 1 per cent.
  • Cost at 0.9 per link: hop by hop M × H / 0.9; end to end M × H / 0.9^H (RMST Table 1): 10 hops, 2.6 times as many transmissions.
  • Overhead: 3-segment handshake, 4-segment close, 20-octet TCP header + 40-octet IPv6; RFC 4919: 21 octets of data over TCP against 33 over UDP in the worst case; program: one reading 550 octets TCP against 58 UDP.
  • Model misfit: many-to-one, event-driven, data-centric; event reliability (ESRT) not per-packet; downstream code needs full reliability to many nodes; per-connection state, ACK traffic, fairness, energy.
  • Instead: TCP splitting at the sink; lightweight, hop-by-hop, energy-aware transport inside (PSFQ, ESRT, CODA, RMST); CoAP over UDP.
  • Program: radio losses alone cut TCP's average window from 63 to 5.9 segments at 1 per cent per hop.
munotes.in508

Why a Sensor Network Cannot Simply Run TCP

Test yourself

1. Why is TCP not suitable for wireless sensor networks? Give five reasons. First, TCP assumes that packet loss signals congestion and cuts its window, while in sensor networks losses come mostly from radio errors, collisions and fading, so TCP throttles itself needlessly. Second, it recovers losses end to end, but over many lossy hops the chance of an end-to-end delivery falls exponentially and each retransmission crosses the whole path, wasting energy; hop-by-hop recovery works better. Third, its connection setup and teardown and its 20-octet header, with an IP header, are large compared with small sensor readings and short frames. Fourth, sensor traffic is many-to-one, event-driven and data-centric, and often needs reliability of the event rather than of every packet, which TCP's model of a reliable byte stream between two addressed hosts does not express. Fifth, TCP keeps per-connection state and sends acknowledgements that burden small nodes and the shared channel, and it does not address fairness among many sources or energy.

2. Compare end-to-end and hop-by-hop error recovery for a multi-hop sensor network. In end-to-end recovery only the destination detects loss and the source retransmits, so with a per-hop error rate p a packet crosses n hops with probability (1 - p)^n, which falls exponentially with the path length, and every retransmission costs transmissions on every hop. In hop-by-hop recovery each intermediate node caches the data it forwards and detects and repairs losses on its own hop, so the chance of detecting and repairing a loss does not depend on the path length and each repair costs one hop. Hop-by-hop recovery scales better in lossy multi-hop networks, at the cost of buffers and processing at every node.

3. With a per-hop error rate of 10 per cent, what fraction of packets crosses 10 hops end to end? What changes with three MAC tries per hop? With one try per hop the fraction is 0.9^10, about 0.349, so about 35 per cent. With up to three tries per hop, each hop succeeds with 1 - 0.1^3 = 0.999, and 10 hops with 0.999^10, about 0.990, so about 99 per cent.

4. What does RMST's Table 1 show about hop-by-hop and end-to-end repair? It counts the link transmissions needed to move M fragments across H hops when each link succeeds with probability 0.9. With caching and repair at every hop, each fragment costs H / 0.9 transmissions, 27.77 for 5 fragments over 5 hops and twice that for 10 hops. Without caching, a lost fragment must be resent from the source, and the costs are 42.33 over 5 hops and 143.39 over 10, about 2.6 times the hop-by-hop cost at 10 hops, and the gap grows with the number of hops and fragments.

munotes.in509

Why a Sensor Network Cannot Simply Run TCP

5. Why does TCP's congestion control perform badly on a lossy wireless path? TCP reduces its congestion window whenever it detects a loss, halving it on three duplicate acknowledgements and resetting it to one segment on a timeout, because on wired paths loss almost always means a congested router. On a wireless path losses happen even when there is no congestion, so TCP keeps cutting its window and rarely lets it grow; its throughput falls far below what the path could carry, and the reduction does nothing to prevent the next radio error.

6. What is TCP splitting, and why is it used with sensor networks? In TCP splitting, TCP connections from users on the Internet end at the sink, and a transport protocol designed for the sensor network carries data between the sink and the sensor nodes. The user and the sink communicate over TCP or UDP across the Internet, where TCP works well, while inside the sensor network the lighter, energy-aware, often hop-by-hop protocol suits the lossy multi-hop radio paths and the small memory of the nodes.

Contents This chapter on its own page

munotes.in510

Chapter Seventy

Transport Protocol Design Issues in WSNs

Syllabus topic Module 1, "Transport Layer and Middleware in WSN: Transport protocol design issues in WSNs"

In one line

Designing transport for a sensor network is a sequence of choices the traditional protocols made once, for a different network: which direction the data flows, whether reliability is owed to every packet or to the event, whether loss is found by acknowledgements or by gaps, whether it is repaired end to end or at each hop, how congestion is noticed, told and relieved, and how little energy, memory and unfairness the whole thing can be built with.

In the wording a student can write in an examination: the design issues for a WSN transport protocol are:

  1. Direction of traffic. Upstream (many sources to one sink) traffic is high-volume, often loss-tolerant; downstream (sink to many nodes: code, queries, commands) is one-to-many and usually needs complete reliability. One protocol rarely serves both.
  2. What reliability means. Packet reliability (every packet delivered) against event reliability (enough reports for the sink to detect the event, ESRT's "observed event reliability"); also block or message reliability, where a whole object must arrive.
  3. Loss detection. ACK (positive), NACK (negative, only gaps are reported), implicit acknowledgement (hearing the next hop forward the packet), or sequence numbers with timeouts; each costs different control traffic and energy.
  4. Loss recovery. End-to-end (only the ends repair) against hop-by-hop (every node caches and repairs), and whether to cache in the network at all.
  5. Congestion detection. By queue length or buffer occupancy, by channel loading, or by rate mismatch between arrival and service; each has a cost and a delay.
  6. Congestion notification. Explicit messages or a bit piggybacked on data; hop-by-hop backpressure or end-to-end regulation from the sink.
  7. Rate adjustment. Throttling, packet dropping, AIMD or rate assignment, and whether the sink sets one rate for all sources.
  8. Energy. Every retransmission, acknowledgement and listening period costs battery; the protocol must be judged in transmissions, not only in delivery.
  9. Memory and state. Caches, per-flow state and timers must fit in a few kilobytes.
  10. Fairness and priority. Sources far from the sink cost more per packet than near ones; equal rates and maximum throughput pull in opposite directions, and some events matter more than others.
  11. Quality of service and timeliness. Some data is worthless late; delay bounds, not just delivery, may be the requirement.

1. Which way does the data flow?

The two directions of a sensor network are not symmetrical.

Upstream, from many sources to a sink, carries readings and event reports. It is the heavy direction, it converges (so the last hop before the sink is the busiest), and it is often tolerant: one lost temperature reading among a hundred changes nothing.

Downstream, from the sink to many or all nodes, carries queries, commands, configuration and new program images. It is light in volume but demanding: a program image with one missing fragment is useless, as RMST notes, "A single missing fragment from a large binary object (such as executable code) may render the data entity useless; therefore, transport layer facilities are required." It is also one-to-many, which no single TCP connection can serve.

munotes.in511

Transport Protocol Design Issues in WSNs

PSFQ was built for the downstream direction, ESRT and CODA for the upstream. A designer must say which problem is being solved before anything else is decided.

2. What does reliability mean here?

The traditional answer is per-packet: deliver every byte. A sensor network can often afford less, and sometimes needs something different.

Event reliability. ESRT's definitions: "The observed event reliability, ri, is the number of received data packets in decision interval i at the sink" and "The desired event reliability, R, is the number of data packets required for reliable event detection. This is determined by the application". The transport problem is then not to save every packet but "to configure the reporting rate, f, of source nodes so as to achieve the required event detection reliability, R, at the sink with minimum resource utilization."

Block reliability. For code and other objects, every fragment matters, and the unit of success is the whole block.

Nothing at all. Where readings are frequent and correlated, an application may want no reliability mechanism, only speed and low cost.

The choice decides everything downstream: with event reliability the protocol can control the rate instead of retransmitting, which is what ESRT does, and which can save energy rather than spend it.

3. How is a loss detected?

Four mechanisms, with different costs:

  • Positive acknowledgement (ACK). The receiver confirms each packet; the sender resends when no confirmation comes. Simple and prompt, but one control packet per data packet, and a lost ACK causes a needless resend.
  • Negative acknowledgement (NACK). The receiver notices a gap in the sequence numbers and asks for the missing packet. Cheap when losses are rare, but a loss at the end of a burst leaves no later packet to reveal the gap, so a timeout is still needed. PSFQ's note: "For a negative acknowledgement system, at least one message has to be received correctly at the destination after a loss has happened in order to detect the loss."
  • Implicit acknowledgement. On a broadcast medium, a sender that hears its neighbour forward the packet knows it arrived. No control packet at all, but the radio must stay on to hear it, and the last hop has no onward transmission to overhear.
  • Sequence numbers and timeouts at the receiver, used with any of the above.

Where the loss detection lives matters as much as its kind: a NACK to the source crosses the whole path, a NACK to the previous hop crosses one.

munotes.in512

Transport Protocol Design Issues in WSNs

4. Where is a loss repaired?

This is the issue [Why a Sensor Network Cannot Simply Run TCP] sized. RMST puts it as the central decision: "The design decisions examined by this paper for the transport layer are primarily concerned with the balance of hop-by-hop vs. end-to-end functionality. Repair requests could be initiated by sinks (receiver end-points), or by in-network nodes on an established path."

  • End to end. Only the source keeps the data; only the destination notices loss. Simple nodes, no caches, but each repair crosses every hop, and the chance of a repair itself getting through falls as (1 - p)^n.
  • Hop by hop. Each node caches what it forwards and repairs its own hop, which PSFQ chose because it "essentially segments multihop forwarding operations into a series of single hop transmission processes that eliminate error accumulation."
  • How much to cache. RMST makes caching optional: in caching mode every node keeps recent fragments; in non-caching mode "only the sources and sinks maintain a cache". Memory is the price.
  • What the MAC already does. MAC retries repair most losses before transport sees them, and RMST found the transport's caching stops paying once the MAC brings losses below 1 per cent.

5. How is congestion noticed?

Congestion in a sensor network is not only a full queue. CODA's setting: "Sensor networks typically operate under light load and then suddenly become active in response to a detected or monitored event", and "The transport of event impulses is likely to lead to varying degrees of congestion in the network depending on the sensing application." The load arrives as a burst, exactly when the data matters most.

Three signals, each considered by CODA:

  • Buffer occupancy or queue length, the wired network's choice. CODA found it weak here: "the buffer occupancy does not provide an accurate indication of congestion even when the link ARQ is enabled ... except in the extreme case when the queue is empty or about to overflow", a bimodal signal, "not responsive enough and too coarse to provide smooth and efficient congestion control."
  • Channel loading. "In CSMA networks, it is straightforward for sensors to listen to the channel, trace the channel busy time and calculate the local channel loading conditions." It measures what a queue cannot: neighbours competing for the same air. But listening costs energy, so CODA samples rather than listening all the time.
  • Rate mismatch, comparing what arrives with what can be forwarded, or, at the sink, comparing the reports received with the reports expected. This is inherently slow: CODA notes that detection based on the reporting rate "is inherently slow and end-to-end in nature".
munotes.in513

Transport Protocol Design Issues in WSNs

Note what makes this different from a wired network: congestion is shared, since a node's neighbours use the same channel, so a node can be congested without a single packet in its own queue.

6. How is congestion told, and to whom?

  • Explicit against implicit. A backpressure message says so directly; a bit set in a forwarded data packet costs nothing extra but travels only where data travels.
  • Hop by hop. CODA's "open-loop hop-by-hop backpressure": "a node broadcasts backpressure messages as long as it detects congestion. Backpressure signals are propagated upstream toward the source. ... Nodes that receive backpressure signals can throttle their sending rates or drop packets based on the local congestion policy". It acts in one hop time, and each node decides whether to pass it on.
  • From the sink. CODA's "closed-loop, multi-source regulation" "operates over a slower time scale and is capable of asserting congestion control over multiple sources from a single sink in the event of persistent congestion." Slow, but it can see and settle the whole event.

Both are needed because the hotspots differ: dense sources make "persistent hotspots" near the sources, sparse ones make "transient hotspots potentially anywhere in the sensor field but likely farther from the sources, toward the sink."

7. How is the rate adjusted?

Once congestion is known, something must give: throttle the sending rate, drop packets by a local policy, or have the sink assign each source a new reporting rate. TCP's additive-increase, multiplicative-decrease is one policy among several, and in a sensor network the rate is often the application's reporting rate, so transport and application decide together. That is ESRT's approach: control f, the reporting rate, to reach R with the least resource use, which also saves energy when the network is delivering more than the sink needs.

8 to 11. The constraints over all of it

  • Energy. A protocol is judged in transmissions, not only in delivery ratio. CODA measures an "energy tax" against a "fidelity penalty"; ESRT uses congestion control to cut energy as well as loss.
  • Memory and state. Caches, sequence-number maps and timers must fit in kilobytes; per-flow state at every node does not scale, which is why sensor protocols are usually stateless or keep only per-hop state.
  • Fairness. Sources far from the sink spend more of the network's transmissions per delivered packet than near ones. Equal rates are fair but reduce total throughput; favouring near sources raises throughput and silences the far ones.
  • Timeliness and priority. A fire alarm that arrives late is worthless, and a protocol may need to carry some traffic ahead of the rest. PSFQ's own goal, for its direction, is "to provide loose delay bounds for data delivery to all the intended receivers."
munotes.in514

Transport Protocol Design Issues in WSNs

The issues, computed

The program computes event reliability from per-packet reliability; counts the transmissions ACK, NACK and implicit acknowledgement each cost over one lossy hop; counts the packets a hotspot drops before hop-by-hop backpressure and before end-to-end regulation reach it; and prices fairness in a field where four of six sources are five hops away.

# The design issues of a sensor transport protocol, put in numbers: what
# reliability means, what loss detection costs, how fast congestion control can
# react, and what fairness costs. All models are ours; the issues are the
# sources'.
from math import comb

# 1. Packet reliability against event reliability (ESRT's notion). n sources
#    each send one report per decision interval; each arrives with probability
#    p; the sink needs R reports to call the event detected.
def at_least(n, p, R):
    return sum(comb(n, k) * p ** k * (1 - p) ** (n - k) for k in range(R, n + 1))

print("Event detected (R of n reports arrive) against per-packet delivery p:")
print("  n   R      p=0.30   p=0.50   p=0.70   p=0.90")
for n, R in ((5, 3), (10, 3), (20, 3), (20, 8), (50, 8)):
    print("  %2d  %2d   " % (n, R) + "".join("%9.4f" % at_least(n, p, R) for p in (0.3, 0.5, 0.7, 0.9)))

# 2. What loss detection costs, in transmissions, to get 100 packets across one
#    hop that loses a fraction q. Every scheme resends the lost data; they
#    differ in the control traffic. ACK: one acknowledgement per arrival, and an
#    ACK may itself be lost. NACK: one per loss noticed, but a loss at the end
#    of a run is found only by a timeout. Implicit ACK: the forwarder's own
#    transmission onward is the acknowledgement, so no control packet at all,
#    but the sender must stay awake to hear it.
print("\nTransmissions to deliver 100 packets over one hop (data + control):")
print("  loss q       ACK: data + ACKs      NACK: data + NACKs   implicit: data")
for q in (0.0, 0.01, 0.05, 0.10, 0.20):
    # ACK: a packet is confirmed only if the data and its ACK both get through,
    # so data costs 100 / (1 - q)^2 and each arrival costs one ACK.
    ack_data, acks = 100 / (1 - q) ** 2, 100 / (1 - q)
    nack_data, nacks = 100 / (1 - q), 100 * q / (1 - q)
    print("  %6.2f  %7.0f + %-5.0f = %5.0f  %6.0f + %-4.0f = %5.0f  %8.0f"
          % (q, ack_data, acks, ack_data + acks, nack_data, nacks, nack_data + nacks, nack_data))

# 3. How fast congestion control can react. A hotspot u hops upstream of the
#    sink receives 10 packets per hop time and can forward 6, so it drops 4 a
#    hop time until the flow slows. Hop-by-hop backpressure throttles the
#    neighbour one hop time later. End-to-end regulation waits for the sink to
#    notice (u hop times for the thinned flow to reach it, plus a decision
#    interval of 5) and for its message to reach the sources (u + 3 hops).
print("\nPackets dropped at a hotspot before the flow slows (10 arrive, 6 forwarded):")
print("  hops from the sink   hop-by-hop backpressure   end-to-end regulation")
for u in (2, 5, 10, 20):
    print("  %18d   %22d   %21d" % (u, 1 * 4, (u + 5 + u + 3) * 4))

# 4. Fairness. Six sources, two 1 hop from the sink and four 5 hops away, share
#    a sink that can take 12 packets per second. Each packet costs one
#    transmission per hop, and the whole network can make 30 transmissions a
#    second. Equal rates against rates proportional to 1 / hops.
def jain(rates):
    return sum(rates) ** 2 / (len(rates) * sum(r * r for r in rates))

hops = [1, 1, 5, 5, 5, 5]
budget = 30.0
equal = [budget / sum(hops)] * 6
weighted = [budget / (len(hops) * h) for h in hops]
for name, rates in (("equal rate to every source", equal), ("rate proportional to 1 / hops", weighted)):
    print("\n%s:" % name)
    print("  rates (packets/s): " + ", ".join("%.2f" % r for r in rates))
    print("  transmissions/s %.1f of %.0f; packets reaching the sink %.2f/s; Jain's fairness %.3f"
          % (sum(r * h for r, h in zip(rates, hops)), budget, sum(rates), jain(rates)))
munotes.in515

Transport Protocol Design Issues in WSNs

Event detected (R of n reports arrive) against per-packet delivery p:
  n   R      p=0.30   p=0.50   p=0.70   p=0.90
   5   3      0.1631   0.5000   0.8369   0.9914
  10   3      0.6172   0.9453   0.9984   1.0000
  20   3      0.9645   0.9998   1.0000   1.0000
  20   8      0.2277   0.8684   0.9987   1.0000
  50   8      0.9927   1.0000   1.0000   1.0000

Transmissions to deliver 100 packets over one hop (data + control):
  loss q       ACK: data + ACKs      NACK: data + NACKs   implicit: data
    0.00      100 + 100   =   200     100 + 0    =   100       100
    0.01      102 + 101   =   203     101 + 1    =   102       101
    0.05      111 + 105   =   216     105 + 5    =   111       105
    0.10      123 + 111   =   235     111 + 11   =   122       111
    0.20      156 + 125   =   281     125 + 25   =   150       125

Packets dropped at a hotspot before the flow slows (10 arrive, 6 forwarded):
  hops from the sink   hop-by-hop backpressure   end-to-end regulation
                   2                        4                      48
                   5                        4                      72
                  10                        4                     112
                  20                        4                     192

equal rate to every source:
  rates (packets/s): 1.36, 1.36, 1.36, 1.36, 1.36, 1.36
  transmissions/s 30.0 of 30; packets reaching the sink 8.18/s; Jain's fairness 1.000

rate proportional to 1 / hops:
  rates (packets/s): 5.00, 5.00, 1.00, 1.00, 1.00, 1.00
  transmissions/s 30.0 of 30; packets reaching the sink 14.00/s; Jain's fairness 0.605
munotes.in516

Transport Protocol Design Issues in WSNs

Event reliability is cheap. Five sources needing three reports fail half the time when packets arrive with probability 0.5. Twenty sources needing three succeed 99.98 per cent of the time at p = 0.5, and 96.45 per cent even at p = 0.3. Redundant sources do what retransmission would have done, at no cost in protocol: this is why a sensor network can often ask for rate control rather than reliability. But the demand matters: twenty sources needing eight reports succeed only 22.77 per cent of the time at p = 0.3, and there the protocol must act.

What detection costs. Over a hop losing 10 per cent, delivering 100 packets costs 111 data transmissions. NACK adds 11 control packets; a positive acknowledgement, which must itself get through, brings the total to 235 transmissions, more than twice the data. Implicit acknowledgement adds nothing on the air, but the sender's radio must be listening, which is where [Duty Cycling: Preamble Sampling, B-MAC and X-MAC] counts its cost. At 20 per cent loss, ACK costs 281 transmissions against NACK's 150.

What a slow signal costs. A hotspot receiving 10 packets a hop time and forwarding 6 drops 4 a hop time. Hop-by-hop backpressure stops the inflow one hop time later: 4 packets lost, wherever the hotspot is. End-to-end regulation waits for the sink to notice and for its message to travel out: 48 packets at 2 hops, 192 at 20. The first is fast and local, the second sees the whole event: CODA runs both.

What fairness costs. With a budget of 30 transmissions a second, giving all six sources the same rate delivers 8.18 packets a second to the sink, with Jain's fairness index at 1. Giving each source a rate proportional to 1 divided by its hop count delivers 14.00 packets a second, 71 per cent more, with fairness at 0.605: the two near sources send five times as often as the four far ones. A far source is expensive, and every sensor transport protocol must choose how much throughput to give up for it.

Distinctions

Upstream (sources to sink)Downstream (sink to nodes)
CarriesReadings, event reportsQueries, commands, code images
VolumeHigh, bursty at eventsLow
PatternMany to one, convergingOne to many
Reliability neededOften the event, not each packetUsually every fragment
ExamplesESRT, CODA, RMSTPSFQ
Packet reliabilityEvent reliability
AsksEvery packet deliveredEnough reports for the sink to decide
Measured byDelivery ratio per flowri against R in a decision interval
Answered byRetransmissionAdjusting the reporting rate
Programp per packet20 sources, 3 needed: 0.9998 at p = 0.5
munotes.in517

Transport Protocol Design Issues in WSNs

ACKNACKImplicit ACK
Sent whenEvery packet arrivesA gap is noticedNever: the onward transmission serves
Cost at 10 per cent loss235 transmissions per 100 packets122111, and the radio must listen
WeaknessControl traffic, lost ACKsA loss at the end of a burst needs a timeoutNo overhearing at the last hop
Queue lengthChannel loadingRate mismatch
SeesThis node's backlogNeighbours competing for the airToo little arriving at the sink
CostFreeListening, so CODA samplesFree
WeaknessBimodal, coarseLocal onlySlow, end to end
Hop-by-hop backpressureSink regulation
SpeedOne hop timeDetection plus the path, both ways
SeesOne hop's troubleThe whole event
Program (hotspot 10 hops out)4 packets dropped112

What it does not mean

Less reliability is not lower quality. Event reliability asks for what the application actually needs; delivering every duplicate report of one fire is the waste.

Hop-by-hop recovery is not always right. It costs a cache at every node, and RMST found it stops paying once the MAC keeps losses below 1 per cent.

A NACK scheme is not free. It is cheap when losses are rare, and it still needs timeouts for a loss with nothing behind it.

An empty queue does not mean no congestion. The channel is shared: a node's neighbours can be saturating the air while its own buffer is empty.

Fairness is not automatic. Equal rates cost throughput; maximum throughput silences the far sources. A protocol must choose, and say which.

These issues are not independent. Choosing event reliability makes rate control the natural remedy; choosing hop-by-hop recovery demands caches and memory; saving energy limits how often the channel can be sampled.

Quick revision

  • Direction: upstream (many to one, bursty, often loss-tolerant) against downstream (one to many, code and commands, must be complete).
  • Reliability: packet, block (every fragment: code), or event (ESRT: observed ri against desired R in a decision interval; control the reporting rate f).
  • Loss detection: ACK, NACK (needs a later packet or a timeout), implicit ACK (overhear the forward), sequence numbers and timers.
  • Loss recovery: end to end against hop by hop; caching at every node or only at the ends; the MAC's retries first.
  • Congestion detection: queue length (bimodal, coarse), channel loading (accurate, costs listening, so sampled), rate mismatch (slow, end to end). Congestion is shared, not only local.
  • Notification: explicit message or piggybacked bit; hop-by-hop backpressure (fast, local) and sink regulation (slow, whole-event). CODA runs both.
  • Rate adjustment: throttle, drop, AIMD, or the sink assigns a reporting rate.
  • Constraints: energy (count transmissions; CODA's energy tax and fidelity penalty), memory and state, fairness (far sources cost more), timeliness and priority.
  • Program: 20 sources, 3 needed: 0.9998 at p = 0.5; ACK 235 transmissions against NACK 122 at 10 per cent loss; hotspot drops 4 with backpressure against 112 with sink regulation at 10 hops; equal rates 8.18 packets/s (fairness 1.000) against 14.00 (fairness 0.605).
munotes.in518

Transport Protocol Design Issues in WSNs

Test yourself

1. List the design issues for a transport protocol in a wireless sensor network. The direction of the traffic (upstream from sources to sink, or downstream from sink to nodes); what reliability is required (every packet, a whole block, or the event); how loss is detected (positive acknowledgements, negative acknowledgements, implicit acknowledgements, sequence numbers and timeouts); where loss is repaired (end to end or hop by hop, with or without caching); how congestion is detected (queue length, channel loading, rate mismatch); how it is notified (explicit or piggybacked, hop-by-hop backpressure or regulation from the sink); how the rate is adjusted; and the constraints of energy, memory and state, fairness among near and far sources, and timeliness or priority.

2. Distinguish packet reliability from event reliability, and say why the difference matters. Packet reliability requires every packet to be delivered, and is achieved by retransmission. Event reliability requires only that the sink receive enough reports about an event to detect it: ESRT defines the observed event reliability as the number of data packets received in a decision interval and the desired reliability as the number needed by the application. It matters because many sensors report the same event, so redundancy replaces retransmission: the protocol can control the sources' reporting rate instead of repairing losses, which saves energy and avoids the congestion that extra traffic would cause.

3. Compare ACK, NACK and implicit acknowledgement as loss-detection mechanisms. A positive acknowledgement is sent for every packet received, so loss is detected quickly but one control packet is spent per data packet and a lost acknowledgement causes an unnecessary retransmission. A negative acknowledgement is sent only when the receiver notices a gap, which is cheap when losses are rare, but a loss at the end of a burst is not revealed by any later packet and needs a timeout. An implicit acknowledgement costs nothing on the air, since the sender overhears its neighbour forwarding the packet, but the sender's radio must be listening and the last hop cannot be confirmed this way.

4. Why is buffer occupancy a poor indicator of congestion in a sensor network, and what does CODA use instead? The medium is shared, so a node can be unable to send because its neighbours are using the channel while its own queue is nearly empty; CODA's simulations found the buffer occupancy bimodal, empty or about to overflow, too coarse and unresponsive for smooth control. CODA therefore combines present and past channel loading, measured by listening to the channel and tracing its busy time, with the current buffer occupancy, and samples the channel rather than listening continuously in order to save energy.

munotes.in519

Transport Protocol Design Issues in WSNs

5. Compare hop-by-hop backpressure with regulation from the sink. Hop-by-hop backpressure is open-loop and fast: a congested node broadcasts backpressure to its upstream neighbours, which throttle or drop, and each decides whether to propagate it further, so the inflow falls after about one hop time. Sink regulation is closed-loop and slower: the sink notices that reliability or the reporting rate has fallen and sends new rates to all sources, which takes a detection interval plus the travel time of the message. The first resolves local, transient hotspots; the second settles persistent congestion across many sources; CODA uses both.

6. Why is fairness a problem in sensor transport, and what does it cost? A source many hops from the sink consumes one transmission on every hop for each packet delivered, while a source one hop away consumes one, so a fixed budget of transmissions buys far fewer packets from distant sources. Giving every source the same rate is fair but wastes capacity; giving near sources higher rates raises the total delivered but silences the far ones, and the sink then hears mostly about its own neighbourhood. In the chapter's example, equal rates delivered 8.18 packets per second with Jain's fairness index 1.000, while rates proportional to the inverse of the hop count delivered 14.00 packets per second with fairness 0.605.

Contents This chapter on its own page

munotes.in520

Chapter Seventy-One

Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST

Syllabus topic Module 1, "Transport Layer and Middleware in WSN: Transport protocol design issues in WSNs"

In one line

Four protocols split the problem the last chapter set out: PSFQ pumps code downstream slowly and repairs every gap at the next hop quickly, ESRT gives up per-packet reliability upstream and instead has the sink dictate one reporting rate that reaches the required number of reports with the least energy, CODA watches the channel and pushes back hop by hop and then from the sink when congestion persists, and RMST adds NACK-based recovery, with or without caches, to a directed diffusion path.

In the wording a student can write in an examination: PSFQ (Pump Slowly, Fetch Quickly) is a downstream, hop-by-hop, NACK-based protocol for distributing code or commands to every node: the user node pumps each fragment slowly, one every Tmin, relays cache it and forward it in sequence after a random delay between Tmin and Tmax, and a relay that sees a gap goes into fetch mode and asks its neighbour aggressively (one fetch may name several losses: loss aggregation); a report operation feeds delivery status back to the user, aggregated hop by hop.

ESRT (Event-to-Sink Reliable Transport) is an upstream protocol built on event reliability: the sink compares the observed reliability with the desired one, normalises it as eta, decides which of five regions the network is in, from (NC, LR) to (C, HR), and broadcasts an updated reporting frequency f, increasing it aggressively when reliability is low and decreasing it when congestion appears, until it holds in the optimal operating region (OOR).

CODA (Congestion Detection and Avoidance) is a congestion control scheme with three parts: receiver-based congestion detection from channel loading and buffer occupancy, sampled to save energy; open-loop hop-by-hop backpressure broadcast upstream, which neighbours answer by throttling or dropping; and closed-loop multi-source regulation, in which the sink acknowledges at a set rate and sources need those acknowledgements to keep sending.

RMST (Reliable Multi-Segment Transport) runs over directed diffusion as a filter, fragments and reassembles data entities, and recovers loss with NACKs sent along the reverse reinforced path, either with caches at every node (hop-by-hop repair) or without (only source and sink cache, leaning on MAC retries).

PSFQ: pump slowly, fetch quickly

PSFQ carries data down to the nodes: a script, a binary image, a command. Losing a fragment is not acceptable, and the receivers are many.

Its reasoning starts from the observation that loss here is not congestion. "Since most sensor network applications generate light traffic most of the time, message loss in the sensor networks usually occurs because of transmission errors due to poor quality wireless links and not because of traffic congestion." Slow injection keeps congestion away; the work is repairing errors.

"PSFQ comprises three functions: message relaying (pump operation), relay-initiated error recovery (fetch operation) and selective status reporting (report operation)."

munotes.in521

Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST

Pump. "A user node broadcasts a packet to its neighbors every Tmin until all the data fragments has been sent out." A neighbour checks its cache, drops duplicates, decrements the TTL and then, only "if the TTL value is not zero and there is no gap in the sequence number", schedules the fragment to be forwarded after "a random period between Tmin and Tmax". The random delay matters on a broadcast medium, "to avoid collisions because RTS/CTS dialogues are inappropriate in broadcasting operations when the timing of rebroadcasts among interfering nodes can be highly correlated."

In-sequence forwarding is the second mechanism. A node with a gap forwards nothing past it. PSFQ thereby localises the damage: "By localizing loss events and not relaying any higher sequence number messages until recovery has taken place, this mechanism operates in a similar fashion to a store-and-forward" scheme. The gap is repaired where it happened, not carried onward.

Fetch. "A node goes into fetch mode once a sequence number gap in a file fragments is detected. A fetch operation is the proactive act of requesting a retransmission from neighboring nodes once loss is detected at a receiving node." It is aggressive, on a timer much shorter than the pump's, which is the protocol's name: pump slowly, fetch quickly. To save messages, "PSFQ uses the concept of 'loss aggregation' whenever loss is detected; that is, it attempts to batch up all message losses in a single fetch operation whenever possible."

Report. A NACK protocol leaves the sender blind, so PSFQ adds feedback on demand: the user sets a report bit, and "The report message is designed to travel from the furthest target node back to the user on a hop-by-hop basis. Each node en route toward the user is capable of piggybacking their report message in an aggregated manner." One message gathers the status of a whole path instead of one message per node.

Timers. Tmin paces the pump, Tmax bounds the forwarding delay, and together they give the loose delay bound the protocol promises: D(n) = Tmax × n × (number of hops), for a file of n fragments.

ESRT: reliability of the event, not of the packet

ESRT carries reports up, and begins by refusing the usual goal. Its notion, from [Transport Protocol Design Issues in WSNs], is event reliability: the observed reliability ri is the number of data packets received in a decision interval, the desired reliability R is what the application needs, and the problem is "to configure the reporting rate, f, of source nodes so as to achieve the required event detection reliability, R, at the sink with minimum resource utilization."

munotes.in522

Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST

Write eta for ri divided by R. As f rises, eta rises, until the network cannot carry the load and congestion pulls it down. That curve divides into five regions, which the paper defines with a tolerance eps and the frequency fmax at which congestion begins:

  • (NC, LR): no congestion, low reliability, f below fmax and eta below 1 - eps.
  • (NC, HR): no congestion, high reliability, f at most fmax and eta above 1 + eps.
  • (C, HR): congestion, high reliability, f above fmax and eta above 1.
  • (C, LR): congestion, low reliability, f above fmax and eta at most 1.
  • OOR, the optimal operating region: f below fmax and eta within eps of 1.
A graph of normalised reliability eta against the reporting frequency f. The curve rises from the origin, flattens and peaks at f = fmax where eta is 1, then collapses beyond fmax. A horizontal band marks eta within eps of 1 and a vertical line marks fmax, dividing the plane into five labelled regions: no congestion and low reliability on the left below the band, no congestion and high reliability above the band left of fmax, the optimal operating region where the band meets the curve, and congestion with high or low reliability to the right of fmax. Marked points trace the program's run from f = 4 up to the optimal region, and from f = 90 collapsing back to f = 1.17 and climbing again

Figure 71.1 ESRT's five regions, after its Fig. 4, with the program's two trajectories

The sink runs one rule at the end of each decision interval and broadcasts the new f. In the paper's own pseudo-code (its Figure 6), with k a counter of consecutive congested intervals:

  • (C, LR): decrease aggressively, f becomes f raised to the power eta / k, and k increases. Reliability is low and the network is congested, so the rate must come down hard.
  • (C, HR): relieve the congestion without giving up reliability, f becomes f / eta, and k returns to 1.
  • (NC, LR): increase aggressively, f becomes f / eta.
  • (NC, HR): decrease cautiously, f becomes (f / 2)(1 + 1 / eta), which is half way between f and f / eta: more reliability than needed is energy wasted, but the cut is gentle.
  • OOR: hold, f unchanged.

Two things are worth noticing. The sink does the deciding, so the nodes stay simple: they listen for the broadcast and set their rate. And congestion is reported by the nodes cheaply: each watches its own buffer and, if the level would overflow in the next interval, sets a congestion notification bit in the header of the packets it forwards.

CODA: detect congestion, push back, then regulate

CODA does not deliver data; it keeps the network from drowning in it. Its setting is the event impulse: "Sensor networks typically operate under light load and then suddenly become active in response to a detected or monitored event", and it is exactly then that "the information it delivers is of greatest importance."

Its three mechanisms:

  • Receiver-based congestion detection. "CODA uses a combination of the present and past channel loading conditions, and the current buffer occupancy, to infer accurate detection of congestion at each receiver with low cost." Listening costs energy, so "CODA uses a sampling scheme that activates local channel monitoring at the appropriate time to minimize cost while forming an accurate estimate."
  • Open-loop hop-by-hop backpressure. "Once congestion is detected, the receiver will broadcast a suppression message to its neighbors and at the same time make local adjustments to prevent propagating the congestion downstream." Upstream nodes "could throttle their sending rates or drop packets based on some local congestion policy (e.g., packet drop, AIMD, etc.)", and each decides whether to pass the signal further. CODA names the distance the signal travels the depth of congestion: "the number of hops that the backpressure message has traversed before a non-congested node is encountered."
  • Closed-loop multi-source regulation. For congestion that will not go away, the sink takes charge. A source whose rate passes a fraction of the channel's theoretical maximum sets a regulate bit; the sink then sends acknowledgements at a set rate (the paper's example is one per hundred events), and "The reception of ACKs at sources would serve as a self-clocking mechanism allowing the sources to maintain the current event rate". A source that stops hearing them must slow down.
munotes.in523

Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST

CODA is judged by two measures of its own: the energy tax it charges and the fidelity penalty it imposes on the application.

RMST: NACKs along a diffusion path

RMST is a reliability layer for directed diffusion ([Directed Diffusion and Rumour Routing]), implemented "as a filter that could be attached to any diffusion node on an as needed basis". It exists for the data that cannot lose a fragment: "A single missing fragment from a large binary object (such as executable code) may render the data entity useless; therefore, transport layer facilities are required."

  • Fragments and entities. A data entity is split into fragments, each with a fragment number, and the total is known, so a receiver can see exactly what is missing. "Reliability in RMST refers to the eventual delivery to all subscribing sinks of any and all fragments related to a unique RMST entity."
  • Loss detection. A watchdog timer. "In non-caching mode, only sinks set timers to detect loss. In caching mode, each caching node on the reinforced path from source to sink detects loss."
  • Repair. A single control message, the NACK, travels back along the reverse reinforced path; a node with the fragment cached answers, otherwise the NACK travels on toward the source.
  • Two modes. "In caching mode, the caching of fragments along reinforced paths is used to limit power loss due to end-to-end retransmission. In non-caching mode, the underlying MAC layer is exploited to limit the transport layer overhead."
  • Node failure is diffusion's business: a new reinforced path appears, and the RMST filter follows it.

RMST's wider finding is where reliability belongs: "We conclude that reliability is important at the MAC layer and the transport layer. MAC-level reliability is important not just to provide hop-by-hop error recovery for the transport layer, but also because it is needed for route discovery and maintenance."

munotes.in524

Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST

The protocols, run

The program delivers a 40-fragment file along a chain of ten relays, each hop losing 15 per cent, with PSFQ's pump, cache and in-sequence forwarding, once with fetch and once without; then it runs ESRT's Figure 6 rule on a model network whose reports peak at f = 40 and collapse beyond it, with 20 reports required, starting below, just above and far above the peak.

# Two of the four protocols, run. PSFQ's pump with in-sequence forwarding and
# its fetch, against the same pump without fetch; then ESRT's own algorithm
# (its Figure 6) driven to the optimal operating region.
import math
import random

# ---- PSFQ.
# A user node injects a 40-fragment file along a chain of 10 relays.
# The pump sends each fragment once; a node caches what arrives and forwards
# only in sequence, so a gap stops it. With fetch, a node that sees a gap asks
# the node upstream, which resends from its cache, and the request is repeated
# until the hole is filled (loss aggregation is not modelled: one fetch per
# gap). Without fetch, a lost fragment is lost and the node stalls there.
def psfq(hops=10, frags=40, q=0.15, fetch=True, seed=71):
    rnd = random.Random(seed)
    have = [set(range(1, frags + 1))] + [set() for _ in range(hops)]
    sends, fetches = 0, 0
    for n in range(hops):
        for frag in range(1, frags + 1):
            if frag not in have[n]:
                break                                   # nothing upstream to pump
            if len(have[n + 1]) + 1 != frag:
                break                                   # a gap: in-sequence forwarding stops
            sends += 1
            if rnd.random() > q:
                have[n + 1].add(frag)
                continue
            if not fetch:
                break                                   # lost, and nobody will ask
            while True:                                 # fetch until the hole is filled
                fetches, sends = fetches + 1, sends + 1
                if rnd.random() > q:
                    have[n + 1].add(frag)
                    break
    return [len(h) for h in have], sends, fetches

for fetch in (True, False):
    got, sends, fetches = psfq(fetch=fetch)
    print("PSFQ %-7s fetch: fragments at each node %s" % ("with" if fetch else "without", got))
    print("                     %d transmissions, %d of them fetch requests and their answers"
          % (sends, 2 * fetches))

# ---- ESRT, exactly the algorithm of its Figure 6. eta is the normalised
# reliability, reports received over reports required. The sink runs this at
# the end of every decision interval and broadcasts the new frequency f.
def esrt_step(f, eta, congestion, k, eps=0.05):
    if congestion:
        if eta < 1:
            return f ** (eta / k), k + 1                # (C, LR): decrease aggressively
        if eta > 1:
            return f / eta, 1                           # (C, HR): relieve the congestion
        return f, k
    if eta < 1 - eps:
        return f / eta, 1                               # (NC, LR): increase aggressively
    if eta > 1 + eps:
        return f / 2 * (1 + 1 / eta), 1                 # (NC, HR): decrease cautiously
    return f, 1                                         # OOR: hold

# A model network, after the shape of the paper's Fig. 4: reports rise with f
# but contention eats into them, they peak at fmax, and beyond it congestion
# collapses them. R reports are required in a decision interval.
fmax, R, eps = 40.0, 20.0, 0.05
def eta_of(f):
    got = f * (1 - f / (2 * fmax)) if f <= fmax else fmax / 2 * math.exp(-(f - fmax) / 15)
    return got / R

def state(f, eta):
    if f > fmax:
        return "(C, LR)" if eta < 1 else "(C, HR)"
    if eta < 1 - eps:
        return "(NC, LR)"
    if eta > 1 + eps:
        return "(NC, HR)"
    return "OOR"

print("\nESRT (Figure 6): reports peak at f = %.0f, %.0f are required, eps = %.2f" % (fmax, R, eps))
for start in (4.0, 55.0, 90.0):
    f, k, row = start, 1, []
    for _ in range(9):
        eta = eta_of(f)
        row.append("%6.2f %5.3f %-8s" % (f, eta, state(f, eta)))
        f, k = esrt_step(f, eta, f > fmax, k, eps)
    print("  from f = %5.1f" % start)
    for line in row:
        print("     f =%s" % line)
munotes.in525

Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST

PSFQ with    fetch: fragments at each node [40, 40, 40, 40, 40, 40, 40, 40, 40, 40, 40]
                     468 transmissions, 136 of them fetch requests and their answers
PSFQ without fetch: fragments at each node [40, 2, 2, 2, 1, 1, 0, 0, 0, 0, 0]
                     11 transmissions, 0 of them fetch requests and their answers

ESRT (Figure 6): reports peak at f = 40, 20 are required, eps = 0.05
  from f =   4.0
     f =  4.00 0.190 (NC, LR)
     f = 21.05 0.776 (NC, LR)
     f = 27.14 0.897 (NC, LR)
     f = 30.27 0.941 (NC, LR)
     f = 32.17 0.962 OOR
     f = 32.17 0.962 OOR
     f = 32.17 0.962 OOR
     f = 32.17 0.962 OOR
     f = 32.17 0.962 OOR
  from f =  55.0
     f = 55.00 0.368 (C, LR)
     f =  4.37 0.206 (NC, LR)
     f = 21.15 0.778 (NC, LR)
     f = 27.19 0.897 (NC, LR)
     f = 30.30 0.941 (NC, LR)
     f = 32.19 0.962 OOR
     f = 32.19 0.962 OOR
     f = 32.19 0.962 OOR
     f = 32.19 0.962 OOR
  from f =  90.0
     f = 90.00 0.036 (C, LR)
     f =  1.17 0.058 (NC, LR)
     f = 20.30 0.757 (NC, LR)
     f = 26.80 0.891 (NC, LR)
     f = 30.08 0.938 (NC, LR)
     f = 32.05 0.960 OOR
     f = 32.05 0.960 OOR
     f = 32.05 0.960 OOR
     f = 32.05 0.960 OOR
munotes.in526

Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST

Fetch is the protocol. With fetch, every one of the ten relays ends with all 40 fragments, at 468 transmissions, of which 136 belong to fetch requests and their answers: 29 per cent overhead for complete delivery over a chain losing 15 per cent per hop. Without fetch, the first relay receives two fragments before one is lost, and in-sequence forwarding stops it there; the file never reaches the rest, and only 11 transmissions are made. In-sequence forwarding without fetch is a dam; with fetch it is what makes each hop repair its own loss.

ESRT finds its region. From f = 4, eta is 0.19, so (NC, LR) increases the rate aggressively to 21.05, then 27.14, 30.27, 32.17, where eta is 0.962, within eps of 1: OOR, and the rate holds. From f = 55 the network is congested and eta only 0.368, so the (C, LR) rule, f raised to the power eta, cuts the rate to 4.37 in one interval, and the climb begins again. From f = 90, eta is 0.036 and the rate collapses to 1.17. The aggressive cut is deliberate: while the network is congested and unreliable, every packet sent is wasted, so ESRT would rather start again from almost nothing than creep down. Notice also where it stops: eta = 0.96, not 1. The optimal region is a band, and holding inside it is the goal, because chasing eta = 1 exactly would cost more energy than the accuracy is worth.

Distinctions

PSFQESRTCODARMST
DirectionDownstream (sink to nodes)Upstream (sources to sink)UpstreamUpstream (source to sinks)
GoalEvery fragment to every nodeEnough reports per event, least energyRelieve congestionEvery fragment of an entity
ReliabilityPer fragment, hop by hopEvent, by rate controlNone (congestion only)Per fragment
Loss detectionGap in sequence numbers at a relayThe sink counts reportsNot its jobWatchdog timer, NACK
RecoveryFetch from a neighbour, cachedNone: raise the rate insteadNoneNACK along the reverse path, cached or not
CongestionAvoided by pumping slowlyRate cut when nodes report itBackpressure, then sink regulationLeft to the MAC and diffusion
Runs onAny MAC with broadcastAnyCSMA MACDirected diffusion
PumpFetch
TimerEvery Tmin, forwarded after Tmin to TmaxMuch shorter, aggressive
What it doesInjects and relays fragments in sequenceAsks a neighbour for the gap
ScopeHop by hop outward, TTL limitedOne hop, with loss aggregation
munotes.in527

Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST

ESRT stateMeaningThe sink's action
(NC, LR)No congestion, too few reportsIncrease aggressively: f / eta
(NC, HR)No congestion, more than neededDecrease cautiously: (f / 2)(1 + 1 / eta)
(C, HR)Congestion, enough reportsDecrease to relieve it: f / eta
(C, LR)Congestion, too few reportsDecrease hard: f to the power eta / k
OORWithin eps of what is neededHold

What it does not mean

PSFQ is not a routing protocol. "Recall that PSFQ is not a routing solution but a transport scheme"; it can run over diffusion, DSDV or plain broadcast.

ESRT does not make packets reliable. It never retransmits: it changes how often sources report, so enough arrive.

A high eta is not good news. Above 1 + eps the network is spending energy on reports the application does not need, which is why ESRT cuts the rate there too.

CODA does not deliver anything. It is congestion control alone, to be run beside a delivery scheme such as diffusion.

RMST's caches are not always worth it. Its own conclusion is that once the MAC's retries bring losses below about 1 per cent, caching and hop-by-hop repair stop paying.

None of the four is a TCP replacement. Each answers part of the problem, for one direction and one kind of reliability.

Quick revision

  • PSFQ: downstream, hop-by-hop, NACK. Pump every Tmin, relay after a random Tmin to Tmax, cache, forward only in sequence, TTL. Fetch aggressively on a gap, with loss aggregation. Report on demand, aggregated hop by hop. Delay bound D(n) = Tmax × n × hops.
  • ESRT: upstream, event reliability; eta = observed / desired; five regions (NC/C, LR/HR, OOR); rules f / eta (NC, LR), (f / 2)(1 + 1 / eta) (NC, HR), f / eta (C, HR), f^(eta / k) (C, LR), hold in OOR; nodes set a congestion notification bit from their buffer level.
  • CODA: congestion detection (channel loading + buffer, sampled), open-loop hop-by-hop backpressure (depth of congestion), closed-loop sink regulation (regulate bit, ACKs as self-clocking). Metrics: energy tax, fidelity penalty.
  • RMST: over directed diffusion, a filter; entities in fragments; watchdog timers, NACK back along the reinforced path; caching or non-caching mode; reliability belongs in the MAC and the transport layer.
  • Program: PSFQ with fetch, all 40 fragments at all 10 relays, 468 transmissions (136 fetch); without fetch the file stops at the first relay. ESRT reaches OOR (eta 0.962) from f = 4, and cuts 55 to 4.37 and 90 to 1.17 when congested.
munotes.in528

Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST

Test yourself

1. Explain PSFQ's three operations. The pump operation injects fragments: the user node broadcasts one every Tmin, and each relay caches what it receives, discards duplicates, decrements the TTL and forwards the fragment after a random delay between Tmin and Tmax, but only if there is no gap in the sequence numbers. The fetch operation is error recovery: a node that detects a gap goes into fetch mode and aggressively requests the missing fragments from its neighbours, batching several losses into one request by loss aggregation. The report operation returns delivery status to the user on demand: the furthest node starts a report message that travels back hop by hop, and each node on the way appends its own status, so one message carries the status of a whole path.

2. Why does PSFQ pump slowly and fetch quickly? Pumping slowly keeps the injected traffic light, so congestion, which is not the main cause of loss in sensor networks, does not arise, and neighbouring relays do not collide with one another. Fetching quickly repairs a gap in far less time than the interval between pumped fragments, so the loss is made good before the next fragment is due and the node can keep forwarding in sequence; the loss is contained at the hop where it happened rather than propagating downstream.

3. Define ESRT's reliability measure and its five regions. The observed event reliability is the number of data packets received at the sink in a decision interval, the desired reliability is the number the application needs, and eta is their ratio. With fmax the reporting frequency at which congestion begins and eps a tolerance, the regions are: (NC, LR), no congestion and eta below 1 - eps; (NC, HR), no congestion and eta above 1 + eps; (C, HR), congestion with eta above 1; (C, LR), congestion with eta at most 1; and the optimal operating region, no congestion and eta within eps of 1.

4. Give ESRT's rule for updating the reporting frequency in each region. In (NC, LR) the frequency is increased aggressively to f / eta. In (NC, HR) it is decreased cautiously to (f / 2)(1 + 1 / eta). In (C, HR) it is decreased to f / eta, relieving the congestion without giving up reliability. In (C, LR) it is decreased aggressively to f raised to the power eta / k, where k counts consecutive congested intervals. In the optimal operating region it is left unchanged.

5. Describe CODA's three mechanisms. Congestion detection at the receiver, which combines the present and past channel loading, measured by sampling the channel to save energy, with the current buffer occupancy. Open-loop hop-by-hop backpressure: a congested node broadcasts backpressure messages upstream, and each node that receives one throttles its sending rate or drops packets by its local policy and decides whether to propagate the signal further; the number of hops the signal travels is the depth of congestion. Closed-loop multi-source regulation: when a source's rate exceeds a fraction of the channel's theoretical maximum it sets a regulate bit, and the sink then sends acknowledgements at a set rate which the sources need in order to keep sending, so the sink controls all the sources of an event.

munotes.in529

Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST

6. What is RMST, and what are its two modes? RMST, reliable multi-segment transport, is a transport layer for sensor networks implemented as a filter over directed diffusion. It splits a data entity into numbered fragments, guarantees the eventual delivery of every fragment to all subscribing sinks, detects loss with watchdog timers and repairs it with NACKs sent back along the reverse reinforced path. In caching mode every node on the path caches fragments and can answer a NACK, so repairs are hop by hop; in non-caching mode only the source and the sinks cache and only sinks set timers, and the protocol relies on the MAC's own retries, trading memory against transmissions.

7. Which of the four protocols would you choose to send a new program image to every node, and why? PSFQ, because the traffic goes downstream from the user node to many receivers, every fragment must arrive since a single missing fragment makes an executable useless, and the load is light enough to pump slowly. Its in-sequence forwarding with hop-by-hop fetch repairs each loss at the hop where it happens, which is what a long lossy path needs, and its report operation tells the user which nodes have the complete image before the new task is started.

Contents This chapter on its own page

munotes.in530

Chapter Seventy-Two

WSN Middleware: Why It Is Needed, and Its Architecture

Syllabus topic Module 1, "Transport Layer and Middleware in WSN: WSN middleware architecture"

In one line

Middleware sits between the operating system on one node and an application that wants an answer from a whole field: it lets a user state a high-level sensing task once, splits it across the nodes, fuses what they return, and reports the result, while respecting what sensor networks demand of any software on them, that it be data-centric, localized, lightweight, and willing to trade the quality of an answer for the energy it costs.

In the wording a student can write in an examination: middleware sits between the operating system and the application.

Why it is needed. It is needed because a sensor network application must otherwise be written directly against the hardware, the operating system and the network protocols of hundreds of nodes, node by node, and because several applications may have to share one network. "The main purpose of middleware for sensor networks is to support the development, maintenance, deployment, and execution of sensing-based applications." Its functions are to provide standardised, portable abstractions hiding the hardware, operating system and network; mechanisms for formulating a high-level sensing task, communicating it to the network, coordinating the nodes to split and distribute it, fusing their readings into one result and reporting it back; to handle the heterogeneity of nodes; to support automatic configuration and error handling for unattended nodes; and to manage time and location, which matter because the data is physical.

Its principles. They are data-centric communication, using application knowledge, localized algorithms, being lightweight, and trading the quality of service of applications against each other and against resources (adaptive fidelity). A common architecture is layered: over the operating system and network, a layer that organises the nodes (in Yu's design, a cluster layer) and a layer that allocates and adapts resources to the application's requirements, with the application above stating what it wants, not how to get it.

Where middleware sits, and why it is needed

"Middleware sits between the operating system and the application." That one sentence gives its place; the reason it is needed is what lies on either side.

Below it, an operating system such as TinyOS ([Why a Sensor Node Needs an Operating System]) manages one node: its timers, its radio, its sensors, its few kilobytes of memory. It knows nothing of the field.

Above it, the application wants something the field knows: the average temperature in a wing of a building, the place where a vibration began, the track of a vehicle. No single node holds that answer.

Without middleware, the gap is filled by the application programmer, node by node. Yu, Krishnamachari and Prasanna describe the practice: "the design of the network protocols and applications are usually closely-coupled, or even combined as a monolithic procedure. However, such procedures are sometimes ad hoc and impose direct interaction with the underlying embedded operating system, or even the hardware components, of sensor nodes."

munotes.in531

WSN Middleware: Why It Is Needed, and Its Architecture

Two further pressures make that unsustainable. The first is portability: a program written against one mote's hardware must be written again for the next. The second is sharing: "multiple applications will be required to be concurrently executed over a single WSN. For instance, a building monitoring system may need to simultaneously monitor the temperature and luminance, check cracks on the wall, track traversing persons, and even communicate with systems in nearby buildings."

So a middleware is asked to sit between and provide "(1) standardized system services to diverse applications, (2) a runtime environment that can support and coordinate multiple applications, and (3) mechanisms to achieve adaptive and efficient utilization of system resources."

What it must do

Roemer, Kasten and Mattern set out the scope. "The main purpose of middleware for sensor networks is to support the development, maintenance, deployment, and execution of sensing-based applications. This includes mechanisms for formulating complex high-level sensing tasks, communicating this task to the WSN, coordination of sensor nodes to split the task and distribute it to the individual sensor nodes, data fusion for merging the sensor readings of the individual sensor nodes into a high-level result, and reporting the result back to the task issuer."

Read as a sequence, that is the life of one query:

  1. Formulate. The user says what is wanted, at a level such as "report size, speed and direction of vehicles over 40 tons", or, in TinyDB, a line of SQL.
  2. Communicate. The task reaches the nodes, usually by flooding or a dissemination protocol.
  3. Coordinate and split. The nodes decide who samples, who forwards, who aggregates.
  4. Fuse. Readings are merged in the network, not at the sink, because a merged reading costs one transmission instead of many.
  5. Report. One result returns to whoever asked.

Two more duties follow from the setting. Nodes are unattended: "In traditional systems, each computing device belongs to someone who is responsible for configuration, maintenance, and error handling. In contrast, WSN nodes must operate unattended, which means that middleware for WSN has to provide new levels of support for automatic configuration and error handling." And the data is physical: "Since WSN process real-world data, the concepts of physical time and location play a much more important role than in traditional computing systems. Time and location of sensed real-world events are key elements for fusing individual sensor readings in order to obtain a high-level sensing result." That is why [Time Synchronisation and Localisation] belongs under middleware as much as under the network.

munotes.in532

WSN Middleware: Why It Is Needed, and Its Architecture

Middleware also reaches outside the field: "The scope of middleware for WSN is not restricted to the sensor network alone, but also covers devices and networks connected to the WSN", since resource-heavy work has to happen on the gateway or beyond.

The design principles

A middleware that ignored the nature of sensor networks would be worse than none: it would hide exactly the costs that matter. Both papers therefore begin from principles, Roemer's from the software design principles already proposed for sensor networks, Yu's as a list of five.

  • Data-centric communication.

"Data-centric communication introduces a new style of node addressing by focusing on the data produced by nodes, since applications are unlikely to request the current sensor reading such as temperature at a specific node, but instead ask for locations where temperature exceeds a certain value."

It gains robustness "by decoupling data from the sensor that produced it." ([Design Principles: Data Centricity, Location, Activity and Heterogeneity])

  • Application knowledge in the nodes. "application knowledge in nodes can significantly improve the resource and energy efficiency, for example by application-specific data caching and aggregation in intermediate nodes." Yu adds the tension it creates: "due to the mission to support and optimize for a broad class of applications, tradeoffs need to be explored between the degree of application-specific and the generality of the middleware."
  • Localized algorithms. "Localized algorithms are distributed algorithms that achieve a global goal by communicating with nodes in some neighborhood only. Such algorithms scale well with increasing network size and are robust to network partitions and node failures."
  • Lightweight. Yu's fourth: "Since the available resources of sensor nodes are low, the middleware itself should be lightweight in terms of the computation and communication requirements. The lightweight requirement necessitates simple and efficient heuristics to be used for suboptimal solutions." A middleware that needs more memory than the application leaves is not a middleware.
  • Trading quality against resources. Roemer's "Adaptive fidelity algorithms allow to trade the quality of the result against resource usage and are thus a key element for resource efficiency." Yu's fifth principle adds the other trade, between applications: "it is very likely that the performance requirements of all the running applications cannot be simultaneously satisfied. Therefore, it's necessary for the middleware to smartly trade the QoS of various applications against each other."

Roemer ties the principles back to the design: "All mechanisms provided by a middleware system should respect the design principles sketched above and the special characteristics of WSN, which mostly boils down to energy efficiency, robustness, and scalability."

A layered architecture

The usual picture has the application on top, the middleware in the middle and the operating system, network protocols and hardware below, with the middleware split into layers of its own: one that organises the nodes into a working group, and one that allocates the group's resources to what the application asked for.

munotes.in533

WSN Middleware: Why It Is Needed, and Its Architecture

Yu's design is of that shape. Its unit is the cluster: "a set of spatially adjacent sensor nodes that reside around the target phenomena and are capable of detecting and/or processing the data of interest. Clusters are dynamically formed during the lifetime of the system, triggered by the changing conditions of the environment, data source, and sensor nodes." One node in each is elected cluster head, "responsible for the control and coordination of sensor nodes within the cluster", as in LEACH ([LEACH: Clusters That Take Turns]) but formed when and where a phenomenon appears rather than fixed in advance.

Above the network, two layers:

  • The cluster layer forms and maintains clusters. "Typically, the data accessibility, node capability (including remaining energy and CCS capabilities), and network connectivity are the criteria for determining the membership of a sensor node", and "the gathering and exchanging of such information should be performed in a distributed way." It also distributes control decisions from the cluster head.
  • The resource management layer: "As the key component of the middleware, the resource management layer commands the allocation and adaptation of resources, such that the QoS requirements specified by the applications can be met." Its two halves are resource allocation, "generating an initial solution when the cluster is formed", and resource adaptation, which "controls the runtime behavior of the cluster".

Between the application and the middleware passes an application specification: what is wanted, the quality of service required, and the policy for adapting when it cannot all be had. That is the interface that lets the application say what and leave how to the middleware.

Yu's paper calls the whole abstraction a virtual machine "because of its similarity to the virtual machine concept in traditional distributed systems in terms of providing application semantic transparency from the physical infrastructure", which is a different sense from the byte-code virtual machine of [Middleware Approaches, and TinyDB].

Why this is harder than middleware on a desktop

Roemer's paper lists what the middleware must survive, and every item is a demand on its design:

  • Limited size and energy and therefore "restricted resources (CPU performance, memory, wireless communication bandwidth and range)".
  • Dynamics. "Node mobility, node failures, and environmental obstructions cause a high degree of dynamics in WSN. This includes frequent network topology changes and network partitions."
  • Communication failures, the everyday state of a low-power radio.
  • Heterogeneity. "WSN may consist of a large number of rather different nodes in terms of sensors, computing power, and memory."
  • Scale, with redundancy. "The large number raises scalability issues on the one hand, but provides a high level of redundancy on the other hand."
  • Unattended operation. "nodes have to operate unattended, since it is impossible to service a large number of nodes in remote, possibly inaccessible locations."
munotes.in534

WSN Middleware: Why It Is Needed, and Its Architecture

And one honest note on the state of the field in 2002: "For sensor nodes, however, the identification and implementation of appropriate operating system primitives is still a research issue." And "In many current projects, applications are executing on the bare hardware without a separate operating system component." And so, "at this early stage of WSN technology it is not clear on which basis future middleware for WSN can typically be built." Middleware was being designed before the floor under it was settled.

The principles, priced

The program prices three principles. It counts the transmissions a global algorithm and a localized one cost as a network grows; it asks what is lost by waking only k of 400 nodes to estimate a mean; and it splits one network's budget of reports between three applications, equally and then by usefulness.

# What the middleware design principles buy, in numbers: localized against
# global algorithms, adaptive fidelity, and the sharing of one network between
# concurrent applications. The models are ours; the principles are Yu,
# Krishnamachari and Prasanna's, and Romer, Kasten and Mattern's.
import math

# 1. Localized algorithms (principle 3). A global algorithm has every node's
#    reading reach every other node; a localized one talks only to the nodes
#    within r hops. On a grid of n nodes with an average path of about sqrt(n)
#    hops, count the transmissions one round costs.
print("Transmissions in one round, global against localized:")
print("   nodes   global (all to all)   localized r = 1   localized r = 2")
for n in (25, 100, 400, 1600, 6400):
    hops = math.sqrt(n)                       # average path across a square grid
    neighbours = lambda r: min(n - 1, 2 * r * (r + 1))     # a grid's r-hop neighbourhood
    print("  %6d %21.0f %17d %17d"
          % (n, n * (n - 1) * hops, n * neighbours(1), n * neighbours(2) * 2))

# 2. Adaptive fidelity (principle 5, and Romer's "trade the quality of the
#    result against resource usage"). The sink wants the mean of a field of 400
#    nodes whose readings vary with standard deviation 2.0 degrees. Waking k of
#    them gives a standard error of 2.0 / sqrt(k) and costs k reports.
n, sigma = 400, 2.0
print("\nAsking only k of %d nodes for the mean (readings vary by %.1f degrees):" % (n, sigma))
print("     k   reports   standard error   error per report saved")
prev = None
for k in (400, 100, 25, 16, 9, 4):
    err = sigma / math.sqrt(k) * math.sqrt((n - k) / (n - 1))    # finite field correction
    extra = "" if prev is None else "%.3f degrees for %d fewer" % (err - prev[0], prev[1] - k)
    print("  %4d %9d %16.3f   %s" % (k, k, err, extra))
    prev = (err, k)

# 3. Concurrent applications (principle 5 again). Three applications share one
#    network with 100 reports a minute to give out. Each has a utility that
#    grows with its share but flattens: 1 - exp(-rate / scale). Equal shares
#    against shares that maximise the total utility, found by giving each
#    report to whichever application gains most from it.
apps = [("fire alarm", 8.0), ("temperature log", 40.0), ("crack survey", 120.0)]
budget = 100
utility = lambda rate, scale: 1 - math.exp(-rate / scale)
equal = {name: budget / len(apps) for name, _ in apps}
greedy = {name: 0.0 for name, _ in apps}
for _ in range(budget):
    name = max(apps, key=lambda a: utility(greedy[a[0]] + 1, a[1]) - utility(greedy[a[0]], a[1]))[0]
    greedy[name] += 1
print("\nThree applications sharing %d reports a minute:" % budget)
for label, share in (("equal shares", equal), ("most useful first", greedy)):
    total = sum(utility(share[name], scale) for name, scale in apps)
    print("  %-18s %s  total usefulness %.3f"
          % (label, ", ".join("%s %3.0f" % (name, share[name]) for name, _ in apps), total))
munotes.in535

WSN Middleware: Why It Is Needed, and Its Architecture

Transmissions in one round, global against localized:
   nodes   global (all to all)   localized r = 1   localized r = 2
      25                  3000               100               600
     100                 99000               400              2400
     400               3192000              1600              9600
    1600             102336000              6400             38400
    6400            3276288000             25600            153600

Asking only k of 400 nodes for the mean (readings vary by 2.0 degrees):
     k   reports   standard error   error per report saved
   400       400            0.000
   100       100            0.173   0.173 degrees for 300 fewer
    25        25            0.388   0.214 degrees for 75 fewer
    16        16            0.491   0.103 degrees for 9 fewer
     9         9            0.660   0.169 degrees for 7 fewer
     4         4            0.996   0.336 degrees for 5 fewer

Three applications sharing 100 reports a minute:
  equal shares       fire alarm  33, temperature log  33, crack survey  33  total usefulness 1.792
  most useful first  fire alarm  23, temperature log  52, crack survey  25  total usefulness 1.859

Localized algorithms. At 25 nodes, all-to-all exchange costs 3,000 transmissions against 100 for a one-hop neighbourhood: bad, but survivable. At 6,400 nodes it costs 3,276,288,000 against 25,600. The global cost grows as n squared times the path length, the localized cost only as n. That is the whole argument for the third principle: not that global algorithms are inelegant, but that they stop being possible.

Adaptive fidelity. Asking all 400 nodes gives the exact mean. Asking 100 costs a quarter of the energy and leaves a standard error of 0.173 degrees; asking 25 costs a sixteenth and leaves 0.388. The returns are lopsided: dropping from 400 reports to 100 costs 0.173 degrees, while dropping from 9 to 4 costs another 0.336 for only five reports saved. An application that can say how much error it will accept lets the middleware stop at the right point on that curve; one that cannot must be given everything.

munotes.in536

WSN Middleware: Why It Is Needed, and Its Architecture

Sharing between applications. Three applications with different appetites share 100 reports a minute. Equal thirds give a total usefulness of 1.792. Giving each report to whichever application gains most from it gives the fire alarm 23, the temperature log 52 and the crack survey 25, for 1.859. The fire alarm needs few reports to be satisfied, so more of them go to the applications still learning something from each one. The gain is small here and the split is uneven, which is Yu's point exactly: the middleware must "smartly trade the QoS of various applications against each other", and the policy for doing so belongs in the application specification, not buried in the code.

Distinctions

Operating systemMiddlewareApplication
ScopeOne nodeThe network as a wholeThe user's question
ManagesTimers, radio, sensors, memory, tasksTasks, coordination, fusion, resources, qualityWhat is wanted
ExampleTinyOS, ContikiTinyDB, Mate, AgillaHabitat monitoring
Knows aboutHardwareNodes and their stateNeither, ideally
PrincipleWhat it asksWhat it saves
Data-centric communicationAddress data, not nodesRobustness when nodes fail
Application knowledgePut what the application knows into the nodesCaching and aggregation in the network
Localized algorithmsTalk to a neighbourhoodProgram: 25,600 against 3,276,288,000 transmissions at 6,400 nodes
LightweightSimple heuristics, small footprintRoom for the application
Trade quality against resourcesAccept a worse answer for less energyProgram: a quarter of the reports for 0.173 degrees of error
Cluster layerResource management layer
JobForm and maintain clusters around a phenomenonAllocate and adapt resources to the application's requirements
Decides byData access, node capability, remaining energy, connectivityThe application specification and its quality requirements
WhenOn the changing conditions of the fieldAt formation (allocation) and while running (adaptation)

What it does not mean

Middleware is not an operating system. The operating system runs one node; the middleware makes many nodes look like one service.

It is not a library the application calls once. It runs while the application runs, coordinating nodes and reallocating resources.

It does not remove the physics. Energy, loss and latency remain; the middleware decides how to spend them and tells the application what it bought.

General is not automatically better. More application knowledge means more efficiency and less generality, and the papers treat the balance as a design decision.

munotes.in537

WSN Middleware: Why It Is Needed, and Its Architecture

A middleware "virtual machine" is not always a byte-code interpreter. Yu uses the phrase for the transparency the abstraction provides; Mate means an interpreter on the mote.

Quick revision

  • Middleware sits between the operating system and the application. Needed because applications otherwise bind to the hardware, the operating system and the protocols of every node, and because several applications share one network.
  • Provides: standardised, portable abstractions; a runtime for concurrent applications; adaptive, efficient use of resources.
  • Functions: formulate a high-level task, communicate it, coordinate and split it, fuse the readings, report the result; plus heterogeneity, automatic configuration and error handling, and time and location.
  • Principles: data-centric communication, application knowledge in nodes, localized algorithms, lightweight, trade quality against resources (adaptive fidelity, and between applications).
  • Architecture: application (with its specification: what, what quality, what to do when it cannot be met) over the middleware (cluster layer, then resource management layer: allocation and adaptation) over the operating system, protocols and hardware.
  • Harder than on a desktop: restricted resources, dynamics and partitions, communication failures, heterogeneity, scale with redundancy, unattended operation.
  • Program: localized 25,600 against global 3,276,288,000 transmissions at 6,400 nodes; 100 of 400 reports leave a standard error of 0.173 degrees; sharing by usefulness 1.859 against 1.792 for equal shares.

Test yourself

1. What is middleware in a wireless sensor network, and why is it needed? Middleware is software that sits between the operating system and the application. It is needed because the operating system manages only a single node, while applications need results from the whole network, so without middleware the programmer must write against the hardware, the operating system and the network protocols of every node, producing ad hoc, monolithic and unportable programs; and because several applications may have to run concurrently on one network, sharing its scarce resources. Middleware provides standardised, portable abstractions, a runtime environment that coordinates concurrent applications, and adaptive management of the network's resources.

2. List the functions of WSN middleware. Supporting the development, maintenance, deployment and execution of sensing applications: formulating complex high-level sensing tasks, communicating the task to the network, coordinating the nodes so the task is split and distributed among them, fusing the individual readings into a high-level result, and reporting that result to whoever issued the task. It must also provide abstractions for the heterogeneity of nodes, new levels of automatic configuration and error handling because nodes are unattended, and support for physical time and location, which are needed to fuse readings; and its scope extends to the devices and networks connected to the sensor network.

munotes.in538

WSN Middleware: Why It Is Needed, and Its Architecture

3. State the design principles a WSN middleware should respect. Data-centric communication, addressing the data rather than the node that produced it; putting application knowledge into the nodes, so that caching and aggregation can be specific to the task, while balancing this against the generality of the middleware; localized algorithms, which achieve a global goal by communicating only within a neighbourhood and so scale and tolerate failures; being lightweight in computation and communication, using simple heuristics rather than optimal solutions; and trading quality of service against resource usage (adaptive fidelity) and between concurrent applications when not all can be satisfied.

4. Describe a layered middleware architecture for a sensor network. The application sits on top and passes down an application specification: what it wants, the quality of service it requires and the policy for adapting when the requirements cannot all be met. Below it the middleware has two layers: a cluster layer, which forms and maintains clusters of spatially adjacent nodes around the phenomenon of interest, choosing members by data access ability, node capability including remaining energy, and network connectivity, and electing a cluster head to coordinate them; and a resource management layer, the core of the middleware, which allocates resources when a cluster is formed and adapts them while it runs, to meet the application's requirements. Below the middleware are the operating system, the network protocols and the hardware.

5. Why do localized algorithms matter so much in sensor network middleware? Because the cost of a global algorithm grows far faster than the network. In the chapter's model, all-to-all exchange over 25 nodes costs 3,000 transmissions but over 6,400 nodes costs more than three thousand million, while a one-hop localized algorithm costs 100 and 25,600 respectively, growing only in proportion to the number of nodes. Localized algorithms also survive partitions and node failures, since a node depends only on its neighbourhood.

6. What is adaptive fidelity, and what does it buy? It is the principle that a middleware may trade the quality of a result against the resources used to obtain it. In the chapter's example, estimating the mean of a field of 400 nodes exactly needs 400 reports, but sampling 100 of them costs a quarter of the energy for a standard error of only 0.173 degrees, and 25 nodes cost a sixteenth for 0.388 degrees. Because the returns are uneven, an application that states how much error it can accept lets the middleware choose a point on that curve instead of always paying for the exact answer.

Contents This chapter on its own page

munotes.in539

Chapter Seventy-Three

Middleware Approaches, and TinyDB

Syllabus topic Module 1, "Transport Layer and Middleware in WSN: WSN middleware architecture"

In one line

Middleware differs by the abstraction it offers the programmer: a database, where the network is one table and the application writes a query; a virtual machine, where a few bytes of byte-code are installed and spread through the field; mobile agents, which carry their own state from node to node; a tuple space, where nodes leave data for one another to read; messages, published to whoever subscribed; or an application-driven stack the application itself tunes, and TinyDB is the first of these carried furthest.

In the wording a student can write in an examination: the main middleware approaches for WSNs are:

  1. Database approach. The network is treated as a distributed database; the user issues a declarative query and the middleware decides how to answer it. TinyDB and Cougar are the examples: Cougar "adopts a database approach where sensor readings are treated like 'virtual' relational database tables. An SQL-like query language is used to issue tasks to the WSN."
  2. Virtual machine approach. A small interpreter runs on every node, and programs are short byte-code capsules that can be injected and can spread. Mate is the example: "a tiny communication-centric virtual machine designed for sensor networks", whose "high-level interface allows complex programs to be very short (under 100 bytes), reducing the energy cost of transmitting new programs."
  3. Mobile agent approach. The application is a set of agents that carry their code and state from node to node. Agilla is the example: agents "can explicitly migrate or clone from node to node while maintaining their state", coordinating through tuple spaces.
  4. Tuple space (shared data) approach. Nodes leave tuples in a shared space, read and removed by pattern; agents or tasks are decoupled in space and time. Agilla's tuple spaces, and TinyLIME, are of this kind.
  5. Message-oriented (publish and subscribe) approach. Producers publish readings on topics and consumers subscribe, which suits event reporting.
  6. Application-driven approach. The application states its requirements and the middleware tunes the stack to them, as in MiLAN.

TinyDB is an acquisitional query processor: it decides not only how to process data but when and whether to acquire it. Its data model is one table, sensors, with "one row per node per instant in time, with one column per attribute", partitioned across the nodes. Queries are SELECT-FROM-WHERE-GROUP BY with a SAMPLE PERIOD, the interval between samples; the period between the start of each sample period is an epoch. A LIFETIME clause instead names how long the network must last and lets TinyDB compute the sample rate. Aggregates are computed inside the network by three functions, a merging function f, an initializer i and an evaluator e, passing partial state records up the routing tree so that each node sends one record rather than forwarding every reading.

munotes.in540

Middleware Approaches, and TinyDB

Why there are approaches and not one answer

Middleware offers the programmer an abstraction, and the abstraction decides what is easy and what is impossible. A database makes the average temperature where the light is high a single line, and a moving target hard; a virtual machine makes any algorithm expressible and pushes the work back to the programmer; agents make an application that walks toward an event natural and make reasoning about the whole network hard.

The approaches are therefore not competitors so much as different bets about what applications will ask for. Roemer's 2002 survey of the field lists several at once: Cougar's database, the Smart Messages Project "based on agent-like messages containing code and data, which migrate through the sensor network", NEST's microcells, and SCADDS built on directed diffusion.

The database approach

The idea. A sensor network holds data; a database language asks for data without saying how to get it. The user writes a query, the middleware plans it, distributes it, and the data comes back.

Cougar took readings as virtual relational tables with an SQL-like language. TinyDB went further: it is "acquisitional", because in a sensor network the data does not exist until somebody spends energy to sample it. Its own statement of the difference: "we advocate acquisitional query processing (ACQP), where we focus not" on data already stored but on "where, when, and how often data is physically acquired".

What it buys. The query is short, portable and can be optimised: the middleware can order predicates so the cheapest sensor is read first, drop nodes that cannot qualify, and aggregate on the way to the sink.

What it costs. Only what the language can express can be asked. TinyDB is detailed below.

The virtual machine approach

The idea. Put a small interpreter on every node and send programs, not images.

Mate is "a bytecode interpreter that runs on TinyOS. It is a single TinyOS component that sits on top of several system components, including sensors, the network stack, and nonvolatile storage (the 'logger')." Its unit is the capsule: "Code is broken in capsules of 24 instructions, each of which is a single byte long; larger programs can be composed of multiple capsules. In addition to bytecodes, capsules contain identifying and version information."

It is a stack machine with three execution contexts, which "correspond to three events: clock timers, message receptions and message send requests", each with an operand stack (depth 16) and a return address stack (depth 8).

Code spreads by itself. "A capsule can be transmitted to other motes using the forw instruction, which broadcasts the issuing capsule for network neighbors to install. These motes will then issue forw when they execute the capsule, forwarding the capsule to their local neighbors." Version numbers stop the spread from looping: a node installs a capsule only if it is newer than the one it holds.

munotes.in541

Middleware Approaches, and TinyDB

Why it matters. Reprogramming is the problem: a deployed network is "reprogrammable although physically unreachable, and this reprogramming can be a significant energy cost." A capsule is 24 bytes; an image is tens of kilobytes, and the program below prices that difference.

What it costs. Interpretation is slower than native code, and the instruction set fixes what can be said concisely.

The mobile agent approach

The idea. Move the computation to the data instead of the data to the computation.

Agilla "structures an application in terms of one or more mobile agents, which are special processes that can explicitly migrate or clone from node to node while maintaining their state." An application therefore follows the phenomenon: "By facilitating agent migration, an application can restrict itself to reside only on relevant nodes. As the environment changes, the application can self-adapt by migrating its agents to positions that best fulfill its goals."

Agents coordinate through tuple spaces, one per node: "To ensure that agents remain autonomous while enabling interagent coordination, Agilla provides localized tuple spaces that are remotely accessible."

What it costs. An agent that migrates carries its state over the air, and a field full of agents is hard to reason about. Agilla was "the first mobile agent system to operate in resource-constrained wireless sensor platforms", built on TinyOS.

The tuple space, message-oriented and application-driven approaches

Tuple space. A shared space of tuples, written with out, read with rd and taken with in, matched by pattern rather than by address: the data-centric principle made into a programming model. Agilla's design shows both the appeal and the limit. The appeal: "since tuple spaces facilitate spatiotemporal decoupling between agents, each agent can be replaced without affecting the other agents in the network." The limit is the cost of a shared space across a changing network, so Agilla keeps each space on one node: "A tuple space in Agilla does not span multiple nodes to avoid the overhead of keeping it consistent in a dynamic environment and to ensure scalability." Its blocking operations, in and rd, wait for a matching tuple; inp and rdp probe without blocking.

Message-oriented. Producers publish readings under a topic and consumers subscribe; the middleware delivers. It fits event reporting, where a source does not know or care who wants the event, and it decouples the two in time.

Application-driven. The application declares what quality it needs and the middleware tunes the network to it, reaching down into the protocol stack. MiLAN, "middleware linking applications and networks", is the named example, and the architecture of [WSN Middleware: Why It Is Needed, and Its Architecture], with its application specification of requirements and adaptation policy, is of this kind.

munotes.in542

Middleware Approaches, and TinyDB

TinyDB in detail

One table. "In TinyDB, sensor tuples belong to a table sensors which, logically, has one row per node per instant in time, with one column per attribute (e.g., light, temperature, etc.) that the device can produce." The table is virtual and acquisitional: "records in this table are materialized (i.e., acquired) only as needed to satisfy the query". Physically it is spread out: "the sensors table is partitioned across all of the devices in the network, with each device producing and storing its own readings", so comparing two nodes' readings requires bringing them to a common node.

A node without a sensor is not excluded: it inserts NULLs, and the reading is filtered out only if the query says so.

Queries and epochs. "Queries in TinyDB, as in SQL, consist of a SELECT-FROM-WHERE-GROUPBY clause supporting selection, join, projection, and aggregation." What SQL does not have is time:

SELECT nodeid, light, temp
FROM sensors
SAMPLE PERIOD 1s FOR 10s

"This query specifies that each device should report its own id, light, and temperature readings (contained in the virtual table sensors) once per second for 10 seconds." The interval is the epoch: "The period of time between the start of each sample period is known as an epoch. Epochs provide a convenient mechanism for structuring computation to minimize power consumption." Nodes agree on when an epoch begins: "Nodes in TinyDB run a simple time synchronization protocol to agree on a global time base that allows them to start and end each epoch at the same time."

A query gets an identifier and can be stopped with STOP QUERY, limited with FOR, or ended by an event.

Lifetime instead of a rate. A user who cares about the battery rather than the sample rate can say so:

SELECT nodeid, accel
FROM sensors
LIFETIME 30 days

"This query specifies that the network should run for at least 30 days, sampling light and acceleration sensors at a rate that is as quick as possible and still satisfies this goal." TinyDB then performs lifetime estimation: "The goal of lifetime estimation is to compute a sampling and transmission rate given a number of Joules of energy remaining." The reasoning behind the clause is worth keeping: "Specifying lifetime is a much more intuitive way for users to reason about power consumption. Especially in environmental monitoring scenarios, scientific users are not particularly concerned with small adjustments to the sample rate ... Such users, however, are very concerned with the lifetime of the network executing the queries."

munotes.in543

Middleware Approaches, and TinyDB

Aggregation inside the network. This is where the database approach earns its place. "Aggregation has the attractive property that it reduces the quantity of data that must be transmitted through the" network. An aggregate is written as three functions: a merging function f, an initializer i and an evaluator e. For AVERAGE, each partial state record is a pair, a sum and a count; the merging function adds them pairwise; the initializer turns one reading x into the pair (x, 1); and "the evaluator e(< S, C >) simply returns S/C." The only requirement is "that the merging function be commutative and associative".

The records travel up the routing tree in step with the epoch: each node waits for its children, combines "any child values it heard with its own local readings", sends one record up, and sleeps. "Notice that motes are idle for a significant portion of each epoch so they can enter a low power sleeping state." The result is one record per node per epoch rather than one reading per node per hop, which the program measures.

Temporal aggregates cover windows: WINAVG(volume, 30s, 5s) "will report the average volume over the last 30 seconds once every 5 seconds, sampling once per second".

Materialisation points are stored tables produced by logging queries, which later queries may read, giving TinyDB sub-queries and windows over stored data.

TinyDB and Mate, computed

The program builds a 60-node field with a routing tree, answers the query SELECT AVG(light) FROM sensors WHERE light > 400 with and without aggregation in the network, prices the difference over 30 days of 30-second epochs, and compares flooding a Mate capsule with flooding a whole program image.

# Two middleware approaches, put to work: the database approach (a TinyDB query
# answered over a routing tree, with and without aggregation in the network)
# and the virtual machine approach (what a Mate capsule costs to send against a
# whole program image).
import random

# ---- The network: 60 nodes in a field, the sink at one corner, a tree built by
# hop count, as a collection protocol would build it.
rnd = random.Random(73)
nodes = [(0.0, 0.0)] + [(rnd.uniform(0, 50), rnd.uniform(0, 50)) for _ in range(59)]
RANGE = 14.0
dist = lambda a, b: ((a[0] - b[0]) ** 2 + (a[1] - b[1]) ** 2) ** 0.5
depth, parent = {0: 0}, {0: None}
frontier = [0]
while frontier:
    nxt = []
    for a in frontier:
        for i, p in enumerate(nodes):
            if i not in depth and dist(nodes[a], p) <= RANGE:
                depth[i], parent[i] = depth[a] + 1, a
                nxt.append(i)
    frontier = nxt
joined = sorted(depth)
children = {i: [c for c, p in parent.items() if p == i] for i in joined}
print("%d of %d nodes joined the tree, deepest %d hops" % (len(joined), len(nodes), max(depth.values())))

# A reading of the light sensor, and the query
#   SELECT AVG(light) FROM sensors WHERE light > 400 SAMPLE PERIOD 30s
light = {i: rnd.randrange(100, 900) for i in joined}
qualify = [i for i in joined if light[i] > 400]
print("nodes whose light is over 400: %d of %d" % (len(qualify), len(joined)))

def subtree(i):
    return [i] + [d for c in children[i] for d in subtree(c)]

# Without aggregation every qualifying reading is carried to the sink, one
# transmission per hop. With it, each node sends one partial state record (a
# sum and a count) up, whatever its subtree holds, as TinyDB's merging function
# allows; a node with nothing to report and no reporting descendants is silent.
plain = sum(depth[i] for i in qualify)
reports = [i for i in joined if i and any(light[j] > 400 for j in subtree(i))]
print("\nSELECT AVG(light) FROM sensors WHERE light > 400")
print("  readings carried to the sink one by one: %d transmissions" % plain)
print("  partial state records, aggregated in the network: %d transmissions" % len(reports))
print("  the sink's answer is the same either way: %.1f from %d readings"
      % (sum(light[i] for i in qualify) / len(qualify), len(qualify)))

# The same query for a month of epochs, and what the saving is worth: at 30 s a
# sample, an epoch every 30 s, and 20 mA for the 4.3 ms a 26-octet frame takes
# on a CC2420-class radio at 250 kb/s.
epochs = 30 * 24 * 60 * 60 // 30
frame_mAs = 20e-3 * (26 * 8 / 250000)
print("  over 30 days (%d epochs): %.1f mA-s against %.1f mA-s of transmit current"
      % (epochs, epochs * plain * frame_mAs, epochs * len(reports) * frame_mAs))

# ---- The virtual machine approach. Mate's capsules are 24 instructions of one
# byte; a mote's program image is tens of kilobytes. What does it cost to
# retask 60 nodes by flooding, one way and the other?
PAYLOAD = 100                                   # octets a frame can carry
for name, size in (("a Mate capsule", 24), ("a whole program image", 16 * 1024)):
    frames = -(-size // PAYLOAD)
    print("\n%s: %d octets, %d frame(s); flooded to %d nodes: %d transmissions"
          % (name, size, frames, len(joined), frames * len(joined)))
munotes.in544

Middleware Approaches, and TinyDB

60 of 60 nodes joined the tree, deepest 7 hops
nodes whose light is over 400: 42 of 60

SELECT AVG(light) FROM sensors WHERE light > 400
  readings carried to the sink one by one: 144 transmissions
  partial state records, aggregated in the network: 47 transmissions
  the sink's answer is the same either way: 667.8 from 42 readings
  over 30 days (86400 epochs): 207.0 mA-s against 67.6 mA-s of transmit current

a Mate capsule: 24 octets, 1 frame(s); flooded to 60 nodes: 60 transmissions

a whole program image: 16384 octets, 164 frame(s); flooded to 60 nodes: 9840 transmissions
munotes.in545

Middleware Approaches, and TinyDB

Aggregation. All 60 nodes joined a tree 7 hops deep, and 42 of them had light over 400. Carrying each qualifying reading to the sink costs one transmission per hop: 144 transmissions an epoch. Aggregating in the network costs one partial state record from each node with something to report, its own reading or a descendant's: 47. The answer is identical, 667.8 over 42 readings, because a sum and a count merge exactly. Over 30 days of 30-second epochs, that is 67.6 mA-s of transmit current instead of 207.0: a third. The saving is not in the radio or the routing but in the middleware knowing what the query means.

Retasking. A Mate capsule is 24 bytes, one frame, so retasking all 60 nodes costs 60 transmissions. A 16 KB program image is 164 frames, so the same field costs 9,840 transmissions, 164 times as much, before any retry. That is the argument for a virtual machine on a mote: not that interpretation is fast, but that the network, not the processor, is where the energy goes.

Distinctions

ApproachAbstractionExampleStrengthWeakness
DatabaseThe network as a table; a declarative queryTinyDB, CougarShort queries, in-network optimisation and aggregationOnly what the language can express
Virtual machineByte-code capsules on every nodeMateTiny programs, cheap retasking, safe executionInterpretation is slow; the instruction set limits
Mobile agentProcesses that migrate with their stateAgillaFollows the phenomenon; self-adaptiveState travels; hard to reason about
Tuple spaceA shared space written and read by patternAgilla's spaces, TinyLIMEDecoupled in space and timeConsistency across nodes is costly
Message-orientedPublish and subscribe on topicsMiresSuits events; producers need no addressesWeak for continuous queries
Application-drivenThe application states its requirementsMiLANTunes the stack to the needThe application must know what it needs
SQL on a serverTinyDB
The tableStored rowssensors, acquired when a query needs them
Where it livesOne machinePartitioned across every node
TimeNot in the languageSAMPLE PERIOD, epoch, FOR, LIFETIME
AggregationIn the serverIn the network, by merging functions and partial state records
Cost modelDisk and CPUEnergy, and the lifetime the user asked for
AVERAGE in TinyDBWhat it is
Partial state recordA pair: the sum and the count
Initializer i(x)Turns one reading into (x, 1)
Merging function fAdds two records pairwise; must be commutative and associative
Evaluator e(S, C)Returns S / C
munotes.in546

Middleware Approaches, and TinyDB

What it does not mean

The database approach is not a database on a mote. The table is virtual; rows are acquired only when a query needs them, and most are never stored.

A Mate capsule is not a whole program. It is 24 instructions; larger programs are composed of several capsules, and the interpreter provides the rest.

Mobile agents are not mobile nodes. The agent moves; the hardware stays.

In-network aggregation is not approximation. A sum and a count merge exactly; the answer equals the one the sink would have computed from every reading.

LIFETIME is not a promise about the hardware. It is a rate computed from the energy left and the query's cost, and it moves if the estimate does.

No approach is the right one. Each fixes what is easy to say, and the application decides which of those is the right thing to make easy.

Quick revision

  • Approaches: database (TinyDB, Cougar), virtual machine (Mate), mobile agent (Agilla), tuple space (Agilla's spaces, TinyLIME), message-oriented (publish and subscribe), application-driven (MiLAN).
  • Mate: byte-code interpreter on TinyOS; capsules of 24 one-byte instructions with a version number; three contexts (clock, receive, send), operand stack 16, call stack 8; forw spreads a capsule; programs "under 100 bytes".
  • Agilla: agents migrate and clone with their state; localized tuple spaces, remotely accessible; out, in, rd, and the probing inp, rdp; a space does not span nodes.
  • TinyDB: acquisitional; one virtual table sensors (a row per node per instant), partitioned across nodes, NULLs for missing sensors; SELECT-FROM-WHERE-GROUP BY plus SAMPLE PERIOD, epoch, FOR, STOP QUERY; LIFETIME with lifetime estimation; materialisation points; temporal aggregates (WINAVG).
  • In-network aggregation: merging function f, initializer i, evaluator e; partial state records; AVERAGE = (sum, count), merge by adding, evaluate S / C; f must be commutative and associative; nodes send in depth order and sleep.
  • Program: 60 nodes, 7 hops, 42 qualify: 144 transmissions an epoch one by one against 47 aggregated; over 30 days 207.0 against 67.6 mA-s. A Mate capsule floods 60 nodes in 60 transmissions; a 16 KB image takes 9,840.

Test yourself

1. Describe the main middleware approaches for wireless sensor networks. The database approach treats the network as a distributed table queried declaratively, as in TinyDB and Cougar. The virtual machine approach runs a small interpreter on each node and distributes short byte-code programs, as in Mate. The mobile agent approach lets processes carry their code and state from node to node, as in Agilla. The tuple space approach has nodes leave tuples that others read or take by pattern, decoupling them in space and time. The message-oriented approach publishes readings on topics to subscribers. The application-driven approach lets the application state its requirements and has the middleware tune the network and the protocol stack to them, as in MiLAN.

munotes.in547

Middleware Approaches, and TinyDB

2. What does it mean to call TinyDB acquisitional, and why does it matter? TinyDB does not assume that the data already exists: it decides where, when and how often readings are physically acquired, as well as how they are processed, and records of the sensors table are materialised only as needed to satisfy a query. It matters because in a sensor network sampling and transmitting cost energy, so choosing not to acquire a reading is often the largest optimisation available, and an ordinary query processor, which assumes stored data, cannot make that choice.

3. Explain TinyDB's data model and the role of an epoch. Logically there is one table, sensors, with one row per node per instant in time and one column per attribute the device can produce; devices lacking a sensor insert NULLs. Physically the table is partitioned across the nodes, each producing and storing its own readings, so comparing readings from different nodes requires collecting them at a common node, usually the root. A query names a SAMPLE PERIOD, and the period between the starts of successive sample periods is an epoch; nodes run a time synchronisation protocol so they begin and end each epoch together, which lets them sample, communicate and then sleep for the rest of the epoch.

4. How does TinyDB compute an aggregate inside the network? Illustrate with AVERAGE. An aggregate is defined by three functions: an initializer, which turns a single reading into a partial state record; a merging function, which combines two partial state records; and an evaluator, which turns a partial state record into the final value. For AVERAGE, the partial state record is a pair of a sum and a count, the initializer turns a reading x into (x, 1), the merging function adds two records pairwise, and the evaluator returns the sum divided by the count. Each node waits for its children in the routing tree, merges their records with its own reading, sends one record to its parent and sleeps; the root evaluates the final record. The merging function must be commutative and associative.

5. Why is in-network aggregation so much cheaper than sending readings to the sink? Without it, each qualifying reading must be forwarded on every hop of its path, so the cost is the sum of all the path lengths. With it, each node sends exactly one partial state record per epoch whatever its subtree contains, so the cost is at most the number of nodes with something to report. In the chapter's 60-node, 7-hop network, 42 qualifying readings cost 144 transmissions per epoch one by one and 47 when aggregated, and over 30 days of 30-second epochs the transmit current fell from 207.0 to 67.6 mA-s, with an identical answer.

munotes.in548

Middleware Approaches, and TinyDB

6. What is Mate, and how does a program reach the nodes? Mate is a tiny byte-code virtual machine for sensor networks, a single TinyOS component over the sensors, the network stack and the logger. Programs are capsules of 24 one-byte instructions, with identifying and version information; larger programs are several capsules. There are three execution contexts, for clock timers, message receptions and message send requests, each with an operand stack and a return address stack. A capsule spreads itself: the forw instruction broadcasts it to the neighbours, which install it if its version is newer than the one they hold and then forward it in turn, so a new program propagates through the network.

7. Why is a virtual machine attractive for reprogramming a deployed network? Because the expensive part of retasking is the radio, not the processor. A deployed network is physically unreachable but must be reprogrammable, and a full binary image is tens of kilobytes: in the chapter's field of 60 nodes, a 16 KB image takes 164 frames per node, 9,840 transmissions, while a 24-byte Mate capsule takes one frame per node, 60 transmissions. The interpreter's slower execution costs far less energy than the transmissions it saves, and the virtual machine also provides a safe execution environment on hardware with no protection mechanisms.

Contents This chapter on its own page

munotes.in549

Module II

Wireless transmission: frequencies, signals, antennas, propagation, multiplexing, modulation and spread spectrum, then cellular systems and GSM, DECT, TETRA and UMTS, and satellite and broadcast systems

munotes.in

Chapter Seventy-Four

Frequencies for Radio Transmission

Syllabus topic Module 2, "Wireless Transmission: Frequency for radio transmission"

In one line

A radio transmission is an electromagnetic wave whose frequency fixes its wavelength, and with it the antenna, the reach and the room for data; the ITU names the bands from ULF to EHF in decades, divides the world into three Regions and allocates each band to services in its Radio Regulations; India's WPC Wing turns that into the National Frequency Allocation Plan and issues the licences; and a few small bands, the ISM bands, are designated for industrial, scientific and medical use, which is where unlicensed sensor radios live.

In the wording a student can write in an examination: a radio wave travels at the speed of light c, so its wavelength is c divided by its frequency: 1 MHz is 300 m, 2.45 GHz is about 12 cm. The ITU names bands by decades of frequency (Recommendation ITU-R V.431): band number N extends from 0.3 × 10^N to 3 × 10^N Hz, giving VLF (3 to 30 kHz), LF (30 to 300 kHz), MF (300 kHz to 3 MHz), HF (3 to 30 MHz), VHF (30 to 300 MHz), UHF (300 MHz to 3 GHz), SHF (3 to 30 GHz) and EHF (30 to 300 GHz), with metric names from myriametric to millimetric waves. Lower frequencies travel further, bend around obstacles and penetrate better, but have little bandwidth and need large antennas; higher frequencies carry more data with small antennas but are blocked more easily.

Spectrum is a shared, finite resource, so its use is regulated. The ITU's Radio Regulations, "an international treaty signed by India and other Member States of the International Telecommunication Union (ITU)", govern spectrum and satellite orbits worldwide, and divide the world into three Regions; India is within Region 3. Each band is allocated to one or more of the radiocommunication services. India's National Frequency Allocation Plan (NFAP-2022), published by the Wireless Planning and Coordination (WPC) Wing of the Department of Telecommunications, applies this nationally and "covers the frequency range up to 3000 GHz"; it does not itself grant the right to transmit, for which "a licence is required to be obtained from the Wireless Planning and Coordination Wing (WPC Wing), Ministry of Communications, unless such a requirement is exempted by the WPC Wing". Certain bands, including 2 400 to 2 500 MHz, are "designated for industrial, scientific and medical (ISM) applications", and there "Radiocommunication services operating within these bands must accept harmful interference which may be caused by these applications": that is the bargain a Wi-Fi, Bluetooth or 802.15.4 radio accepts.

Frequency and wavelength

A radio wave is an electromagnetic wave. All such waves travel in free space at the same speed, c, about 300,000 km/s, so a wave's frequency f and its wavelength (lambda) are tied together:

munotes.in550

Frequencies for Radio Transmission

lambda = c / f

Every property a designer cares about follows from that one relation, because the wavelength is the size of the thing the wave interacts with:

  • The antenna. An efficient antenna is a fraction of a wavelength, typically a half or a quarter ([Antennas: Radiators, Dipoles and Radiation Patterns]). At 1 MHz a quarter wave is 75 m; at 2.45 GHz it is 3 cm, which fits on a mote.
  • How it travels. Long waves follow the ground and bend around hills; short waves travel in straight lines and are stopped by walls and foliage ([Signal Propagation: Ranges, Path Loss and How a Signal Travels]).
  • How much data. The room for data goes with the width of the band in hertz, and the higher the band, the more hertz there are to allocate.

The ITU's band names

Recommendation ITU-R V.431 fixes the vocabulary. Its rule is simple: "Band number N extends from 0.3 × 10^N to 3 × 10^N Hz", so each band is a decade and each has a metric name taken from its wavelengths, from myriametric to millimetric waves. The Recommendation asks that "administrations should always use the nomenclature of the frequency and wavelength bands given in Table 1".

The program prints the whole table with the wavelengths worked out. In brief, and with what each band is used for:

BandFrequenciesWavelengthsTypically used for
VLF3 to 30 kHz100 to 10 kmSubmarine communication, navigation
LF30 to 300 kHz10 to 1 kmTime signals, long-wave broadcasting
MF300 kHz to 3 MHz1 km to 100 mAM broadcasting
HF3 to 30 MHz100 to 10 mShort-wave broadcasting, amateur radio, over-the-horizon links
VHF30 to 300 MHz10 to 1 mFM broadcasting, television, aviation
UHF300 MHz to 3 GHz1 m to 10 cmCellular, Wi-Fi, Bluetooth, 802.15.4, television
SHF3 to 30 GHz10 to 1 cmSatellite links, radar, 5 GHz Wi-Fi
EHF30 to 300 GHz10 to 1 mmMillimetre-wave links, automotive radar

Above EHF the Recommendation continues with decimillimetric, centimillimetric, micrometric and decimicrometric waves, up to 3,000 THz, which is light.

The sensor radio of [The 802.15.4 Physical Layer] uses three slices of UHF: 868 to 868.6 MHz, 902 to 928 MHz and 2 400 to 2 483.5 MHz.

Who decides who may transmit

Spectrum cannot be owned by use: two transmitters on one frequency in one place interfere, and a wave does not stop at a border. So the use of spectrum is decided by treaty and then by each country.

The ITU. "The Radio Regulations, an international treaty signed by India and other Member States of the International Telecommunication Union (ITU), governs the use of radio-frequency spectrum and satellite-orbits (geostationary and non-geostationary) at the global level." They are revised at World Radiocommunication Conferences, whose year is printed against each footnote in the tables (WRC-07, WRC-12, and so on).

munotes.in551

Frequencies for Radio Transmission

Three Regions. "For the purpose of frequency allocation, the world has been divided into three Regions. They are referred to as Region 1, Region 2 and Region 3 in the Radio Regulations. ... India is within Region 3." Allocations may differ between Regions, which is why 902 to 928 MHz is an ISM band in Region 2 (the Americas) and not everywhere, and why an 802.15.4 radio for that band is sold in some countries and not others.

Allocation. "the spectrum is divided into frequency bands and each band is allocated to one or more radiocommunication services. The principle of designating a band for the use by specified radiocommunication services is referred to as frequency allocation." The NFAP counts "forty one in total (the 41st being special service)" such services.

India. The NFAP-2022 is the national plan: "The National Frequency Allocation Plan-2022 of India provides a broad regulatory framework, identifying which frequency bands are available for cellular mobile service, Wi-fi, sound and television broadcasting, radionavigation for aircrafts and ships, defence and security communications, disaster relief and emergency communications, satellite communications and satellite-broadcasting, and amateur service, to name just a few." It is built on the treaty: "the Radio Regulations (Edition of 2020) is the foundational text used for drawing up the National Frequency Allocation Plan 2022 (NFAP-2022)", and it "covers the frequency range up to 3000 GHz".

A plan is not a permit. "NFAP-2022, though governing the use of spectrum in India, does not by itself provide the right to use the spectrum. Before any part of the spectrum is put to use in India, a licence is required to be obtained from the Wireless Planning and Coordination Wing (WPC Wing), Ministry of Communications, unless such a requirement is exempted by the WPC Wing." That last clause is the one a sensor network depends on.

The ISM bands

Some bands are set aside for equipment that uses radio energy for something other than communication: industrial heaters, microwave ovens, medical diathermy. Radio Regulations footnote 5.150, as the NFAP prints it, names them:

"40.66-40.70 MHz (centre frequency 40.68 MHz), 902-928 MHz in Region 2 (centre frequency 915 MHz), 2 400-2 500 MHz (centre frequency 2 450 MHz), 5 725-5 875 MHz (centre frequency 5 800 MHz), and 24-24.25 GHz (centre frequency 24.125 GHz) are also designated for industrial, scientific and medical (ISM) applications."

(The same footnote lists three more below these: 13 553 to 13 567 kHz, 26 957 to 27 283 kHz and, through footnote 5.138, 6 765 to 6 795 kHz.)

munotes.in552

Frequencies for Radio Transmission

The condition attached is the important part: "Radiocommunication services operating within these bands must accept harmful interference which may be caused by these applications." A communication service in an ISM band has no protection. The higher bands in footnote 5.138, "61-61.5 GHz (centre frequency 61.25 GHz), 122-123 GHz (centre frequency 122.5 GHz), and 244-246 GHz (centre frequency 245 GHz)", carry a stronger condition: "The use of these frequency bands for ISM applications shall be subject to special authorization by the administration concerned, in agreement with other administrations whose radiocommunication services might be affected."

Why this matters for sensor networks. Because an ISM band needs no licence for low-power equipment in most countries, the whole family of short-range radios lives there: Wi-Fi, Bluetooth, Zigbee and 802.15.4 all share 2 400 to 2 500 MHz. Nobody there is protected from anybody, which is why 802.15.4 spreads its signal over 32 chips per symbol, gives sixteen channels and prefers those clear of Wi-Fi ([The 802.15.4 Physical Layer]), and why [Built on 802.15.4: Zigbee Routing, Security and the Later Amendments] hops channels: interference is not a fault in the band, it is the band's defining condition.

The spectrum, computed

The program prints the ITU's table with the wavelengths of each band's edges, computes the quarter-wave antenna a radio would need in each band a sensor network might use, and measures the ISM bands.

# The radio spectrum: the ITU's bands with their wavelengths, the ISM bands
# every unlicensed radio shares, and what a band's width is worth.
C = 299792458.0                                  # the speed of light, m/s

def wavelength(hz):
    return C / hz

BANDS = [(3, "ULF", 300, 3e3, "hectokilometric"), (4, "VLF", 3e3, 30e3, "myriametric"),
         (5, "LF", 30e3, 300e3, "kilometric"), (6, "MF", 300e3, 3e6, "hectometric"),
         (7, "HF", 3e6, 30e6, "decametric"), (8, "VHF", 30e6, 300e6, "metric"),
         (9, "UHF", 300e6, 3e9, "decimetric"), (10, "SHF", 3e9, 30e9, "centimetric"),
         (11, "EHF", 30e9, 300e9, "millimetric")]

def show(m):
    for limit, unit, scale in ((1000, "km", 1000), (1, "m", 1), (0, "mm", 0.001)):
        if m >= limit:
            v = m / scale
            return ("%.0f %s" if v >= 100 else "%.1f %s") % (v, unit)

def hz(f):
    for limit, unit, scale in ((1e9, "GHz", 1e9), (1e6, "MHz", 1e6), (1e3, "kHz", 1e3)):
        if f >= limit:
            v = f / scale
            return ("%.0f %s" if v == int(v) else "%.2f %s") % (v, unit)
    return "%.0f Hz" % f

print("ITU band nomenclature (Recommendation ITU-R V.431-8): band number N runs from")
print("0.3 x 10^N to 3 x 10^N Hz, and the wavelength is the speed of light divided by it.")
print("   N  symbol   frequency range         wavelengths             metric subdivision")
for n, sym, lo, hi, metric in BANDS:
    print("  %2d  %-6s %9s to %-9s %9s down to %-9s %s waves"
          % (n, sym, hz(lo), hz(hi), show(wavelength(lo)), show(wavelength(hi)), metric))

# A quarter-wave antenna is a quarter of the wavelength: why low bands are not
# used by small devices.
print("\nA quarter-wave antenna at the bands a sensor network might use:")
for name, hz in (("AM radio, 1 MHz", 1e6), ("VHF, 100 MHz", 100e6), ("868 MHz", 868.3e6),
                 ("915 MHz", 915e6), ("2450 MHz", 2450e6), ("5.8 GHz", 5800e6), ("24 GHz", 24.125e9)):
    lam = wavelength(hz)
    print("  %-16s wavelength %-9s quarter-wave %s" % (name, show(lam), show(lam / 4)))

# The ISM bands of Radio Regulations 5.138 and 5.150, as the NFAP prints them,
# and how wide each is in relative terms.
ISM = [("6 765-6 795 kHz", 6.765e6, 6.795e6, "5.138"), ("13 553-13 567 kHz", 13.553e6, 13.567e6, "5.150"),
       ("26 957-27 283 kHz", 26.957e6, 27.283e6, "5.150"), ("40.66-40.70 MHz", 40.66e6, 40.70e6, "5.150"),
       ("902-928 MHz (Region 2)", 902e6, 928e6, "5.150"), ("2 400-2 500 MHz", 2400e6, 2500e6, "5.150"),
       ("5 725-5 875 MHz", 5725e6, 5875e6, "5.150"), ("24-24.25 GHz", 24e9, 24.25e9, "5.150"),
       ("61-61.5 GHz", 61e9, 61.5e9, "5.138"), ("122-123 GHz", 122e9, 123e9, "5.138"),
       ("244-246 GHz", 244e9, 246e9, "5.138")]
print("\nISM bands (Radio Regulations 5.138 and 5.150, as the NFAP prints them):")
print("  band                      width      width as a share   16 channels of 2 MHz fit?")
for name, lo, hi, rr in ISM:
    width = hi - lo
    print("  %-24s %7.2f MHz %14.2f%%   %s"
          % (name, width / 1e6, 100 * width / ((lo + hi) / 2), "yes" if width >= 16 * 2e6 else "no"))
munotes.in553

Frequencies for Radio Transmission

ITU band nomenclature (Recommendation ITU-R V.431-8): band number N runs from
0.3 x 10^N to 3 x 10^N Hz, and the wavelength is the speed of light divided by it.
   N  symbol   frequency range         wavelengths             metric subdivision
   3  ULF       300 Hz to 3 kHz        999 km down to 99.9 km   hectokilometric waves
   4  VLF        3 kHz to 30 kHz      99.9 km down to 10.0 km   myriametric waves
   5  LF        30 kHz to 300 kHz     10.0 km down to 999 m     kilometric waves
   6  MF       300 kHz to 3 MHz         999 m down to 99.9 m    hectometric waves
   7  HF         3 MHz to 30 MHz       99.9 m down to 10.0 m    decametric waves
   8  VHF       30 MHz to 300 MHz      10.0 m down to 999 mm    metric waves
   9  UHF      300 MHz to 3 GHz        999 mm down to 99.9 mm   decimetric waves
  10  SHF        3 GHz to 30 GHz      99.9 mm down to 10.0 mm   centimetric waves
  11  EHF       30 GHz to 300 GHz     10.0 mm down to 1.0 mm    millimetric waves

A quarter-wave antenna at the bands a sensor network might use:
  AM radio, 1 MHz  wavelength 300 m     quarter-wave 74.9 m
  VHF, 100 MHz     wavelength 3.0 m     quarter-wave 749 mm
  868 MHz          wavelength 345 mm    quarter-wave 86.3 mm
  915 MHz          wavelength 328 mm    quarter-wave 81.9 mm
  2450 MHz         wavelength 122 mm    quarter-wave 30.6 mm
  5.8 GHz          wavelength 51.7 mm   quarter-wave 12.9 mm
  24 GHz           wavelength 12.4 mm   quarter-wave 3.1 mm

ISM bands (Radio Regulations 5.138 and 5.150, as the NFAP prints them):
  band                      width      width as a share   16 channels of 2 MHz fit?
  6 765-6 795 kHz             0.03 MHz           0.44%   no
  13 553-13 567 kHz           0.01 MHz           0.10%   no
  26 957-27 283 kHz           0.33 MHz           1.20%   no
  40.66-40.70 MHz             0.04 MHz           0.10%   no
  902-928 MHz (Region 2)     26.00 MHz           2.84%   no
  2 400-2 500 MHz           100.00 MHz           4.08%   yes
  5 725-5 875 MHz           150.00 MHz           2.59%   yes
  24-24.25 GHz              250.00 MHz           1.04%   yes
  61-61.5 GHz               500.00 MHz           0.82%   yes
  122-123 GHz              1000.00 MHz           0.82%   yes
  244-246 GHz              2000.00 MHz           0.82%   yes
munotes.in554

Frequencies for Radio Transmission

The bands. Each decade of frequency is a decade of wavelength: ULF's waves are hundreds of kilometres long, EHF's are millimetres. The metric names follow the wavelength exactly, which is what makes them worth learning: decimetric waves are UHF because a decimetre is 10 cm, the wavelength at 3 GHz.

Why sensor radios are in UHF. A quarter-wave antenna at 1 MHz is 74.9 m, at 100 MHz 749 mm, at 868 MHz 86.3 mm and at 2 450 MHz 30.6 mm. A mote is a few centimetres across, so the wavelength has to be a few centimetres too. At 24 GHz the antenna would be 3.1 mm, small enough, but the wave would be stopped by a leaf.

How much room there is. The 2 400 to 2 500 MHz band is 100 MHz wide, which is why 802.15.4 can put sixteen 2 MHz channels 5 MHz apart in it and Wi-Fi can put three 22 MHz channels there. The band at 13 553 to 13 567 kHz is 14 kHz wide, room for one narrow channel and nothing else. As a share of its own centre frequency, though, the 2.4 GHz band is only 4.08 per cent, and the 24 GHz band 1.04 per cent: the higher bands are wider in hertz, not in proportion, and that absolute width is what carries data.

Distinctions

Lower frequencies (VLF to VHF)Higher frequencies (UHF to EHF)
WavelengthMetres to kilometresCentimetres to millimetres
AntennaLargeSmall, fits on a device
PropagationGround and sky waves; bends around obstaclesLine of sight; blocked by walls and leaves
Range for a given powerLongerShorter
Bandwidth availableLittleMuch
Typical useBroadcasting, navigation, submarinesCellular, Wi-Fi, sensor networks, radar, satellite
munotes.in555

Frequencies for Radio Transmission

AllocationLicence
WhoITU Radio Regulations, then the NFAPThe WPC Wing
What it saysWhich services may use a bandWhether this transmitter may operate
ScopeA band, worldwide or by RegionOne user or one class of equipment
ExemptionNot applicablePossible: ISM and other low-power use
A licensed bandAn ISM band
Who may transmitThe licenseeAnyone within the rules
Protection from interferenceYes, by the allocationNone: services "must accept harmful interference"
ExampleA cellular operator's 900 MHz carrier2 400 to 2 500 MHz: Wi-Fi, Bluetooth, 802.15.4
Design consequencePlan for the assigned channelSpread, hop, sense the channel, retry

What it does not mean

A band number is not a channel. ITU bands are decades of spectrum; channels are the small slices an allocation and a standard carve out of them.

ISM is not "free spectrum". It is spectrum where interference is permitted and unprotected, and where power and behaviour are still regulated by the administration.

An allocation is not a permission. The NFAP allocates bands to services; transmitting still needs a licence from the WPC Wing unless the requirement is exempted.

Higher frequency does not mean better. It means more bandwidth and smaller antennas, at the cost of range and penetration.

The three Regions are not time zones or continents in the ordinary sense. They are the ITU's own division for allocation, and the NFAP warns that "regions" without a capital R in its text do not mean them.

Quick revision

  • lambda = c / f; c is about 3 × 10^8 m/s. 1 MHz is 300 m, 868 MHz is 345 mm, 2.45 GHz is 122 mm.
  • ITU-R V.431: band N runs from 0.3 × 10^N to 3 × 10^N Hz. VLF 3 to 30 kHz, LF 30 to 300 kHz, MF 300 kHz to 3 MHz, HF 3 to 30 MHz, VHF 30 to 300 MHz, UHF 300 MHz to 3 GHz, SHF 3 to 30 GHz, EHF 30 to 300 GHz; metric names myriametric to millimetric.
  • Lower: longer range, better penetration, big antennas, little bandwidth. Higher: small antennas, much bandwidth, line of sight.
  • Radio Regulations: an ITU treaty; three Regions, India in Region 3; bands allocated to services (41 in the NFAP).
  • NFAP-2022 (WPC Wing, DoT): national plan, built on the Radio Regulations 2020, up to 3000 GHz; a licence from the WPC Wing is still needed unless exempted.
  • ISM bands (5.150): 40.66-40.70 MHz, 902-928 MHz (Region 2), 2 400-2 500 MHz, 5 725-5 875 MHz, 24-24.25 GHz, plus 13 553-13 567 kHz and 26 957-27 283 kHz; (5.138): 6 765-6 795 kHz, 61-61.5 GHz, 122-123 GHz, 244-246 GHz. Services there must accept harmful interference.
  • Program: quarter-wave 74.9 m at 1 MHz against 30.6 mm at 2 450 MHz; the 2.4 GHz band is 100 MHz wide, 4.08 per cent of its centre frequency.
munotes.in556

Frequencies for Radio Transmission

Test yourself

1. State the relation between frequency and wavelength and compute the wavelength at 900 MHz and at 2.4 GHz. A radio wave travels at the speed of light, so its wavelength is the speed of light divided by its frequency. At 900 MHz the wavelength is about 3 × 10^8 divided by 9 × 10^8, that is about 0.33 m or 33 cm; at 2.4 GHz it is about 0.125 m or 12.5 cm.

2. Give the ITU's band nomenclature from VLF to EHF with the frequency range of each. Band number N runs from 0.3 × 10^N to 3 × 10^N Hz. VLF is 3 to 30 kHz (myriametric waves), LF 30 to 300 kHz (kilometric), MF 300 kHz to 3 MHz (hectometric), HF 3 to 30 MHz (decametric), VHF 30 to 300 MHz (metric), UHF 300 MHz to 3 GHz (decimetric), SHF 3 to 30 GHz (centimetric) and EHF 30 to 300 GHz (millimetric).

3. Why do short-range wireless devices use UHF and SHF rather than HF or VHF? Because the antenna must be a fraction of a wavelength and a small device has no room for a large one: a quarter-wave antenna is about 75 m at 1 MHz and 749 mm at 100 MHz, but only 86 mm at 868 MHz and 31 mm at 2.45 GHz. The higher bands also have far more bandwidth in hertz, so they can carry more data and more channels, and the licence-free ISM bands most short-range equipment uses lie there. The cost is shorter range and weaker penetration through walls and foliage.

4. Who regulates the use of radio frequencies, internationally and in India? Internationally, the International Telecommunication Union, through the Radio Regulations, a treaty signed by its Member States, which govern spectrum and satellite orbits and are revised at World Radiocommunication Conferences; for allocation the world is divided into three Regions, and India is in Region 3. In India, the Wireless Planning and Coordination Wing of the Department of Telecommunications, Ministry of Communications, publishes the National Frequency Allocation Plan, most recently NFAP-2022, based on the Radio Regulations and covering frequencies up to 3000 GHz, and issues the licences without which spectrum may not be used unless the requirement is exempted.

5. What are the ISM bands, and what condition attaches to them? They are bands designated for industrial, scientific and medical applications, that is for equipment that radiates radio energy for purposes other than communication. Footnote 5.150 of the Radio Regulations designates, among others, 40.66 to 40.70 MHz, 902 to 928 MHz in Region 2, 2 400 to 2 500 MHz, 5 725 to 5 875 MHz and 24 to 24.25 GHz. The condition is that radiocommunication services operating within these bands must accept harmful interference caused by ISM applications; they are not protected, which is the price of using them without an individual licence.

munotes.in557

Frequencies for Radio Transmission

6. Why does the fact that 2 400 to 2 500 MHz is an ISM band shape the design of IEEE 802.15.4? Because the band is shared with Wi-Fi, Bluetooth, microwave ovens and industrial equipment, and no user of it is protected from interference. So 802.15.4 spreads each symbol over 32 chips, which lets a receiver decode through interference; it defines sixteen channels so a network can move away from a busy one; its annex identifies the four channels that fall in the guard bands between Wi-Fi channels; and its successors hop between channels. The band's width, 100 MHz, is what makes sixteen 2 MHz channels possible in the first place.

Contents This chapter on its own page

munotes.in558

Chapter Seventy-Five

Signals: Amplitude, Frequency and Phase

Syllabus topic Module 2, "Wireless Transmission: Signals"

In one line

Every radio signal is built from sine waves, and a sine wave has exactly three things that can be changed: its amplitude, its frequency and its phase; a signal can be drawn against time, against frequency (its spectrum, since Fourier showed any periodic signal is a sum of sine waves) or as a point whose distance is amplitude and whose angle is phase, and the width of its spectrum, its bandwidth, is what limits how fast it can carry data.

In the wording a student can write in an examination: a sine wave is written g(t) = A sin(2 pi f t + phi), where A is the amplitude, f the frequency in hertz, and phi the phase. The period is T = 1 / f. These three parameters are the only handles a transmitter has: changing A is amplitude modulation, changing f is frequency modulation, changing phi is phase modulation ([Modulation: ASK, FSK and PSK]). A signal may be represented in three domains: the time domain, amplitude against time, which shows the waveform; the frequency domain, amplitude against frequency, which shows which sine waves it is made of; and the phase domain or phase state diagram, in which a signal is one point, its distance from the origin the amplitude and its angle the phase, resolved into an in-phase (I) and a quadrature (Q) component.

Fourier's result is that any periodic signal g(t) with period T can be written as a sum of sine and cosine waves at multiples of the fundamental frequency f = 1 / T (its harmonics). A square wave is the sum of its odd harmonics with amplitudes falling as 1/n. A digital signal is therefore made of infinitely many harmonics, and any real channel, which passes only a finite band, rounds its edges. Bandwidth is the width of the band of frequencies a signal occupies or a channel passes, in hertz. Nyquist's limit says a noiseless channel of bandwidth B carries at most 2B symbols per second, so the bit rate is 2B times the bits per symbol.

The sine wave and its three parameters

Everything a radio transmits is built from this one function of time:

g(t) = A sin(2 pi f t + phi)

  • A, the amplitude, is how far the wave swings from zero. It sets the power, and therefore the range.
  • f, the frequency, is how many cycles pass in a second, in hertz. Its reciprocal is the period T = 1 / f, the time of one cycle. The frequency also fixes the wavelength ([Frequencies for Radio Transmission]).
  • phi, the phase, is where in its cycle the wave is at t = 0, in radians or degrees. Alone, a phase is invisible; against a reference, it carries information.
munotes.in559

Signals: Amplitude, Frequency and Phase

The three are independent, which is why they are the three axes of modulation. A transmitter that can change all three at once, as QAM does, sends more bits per symbol than one that changes only one ([Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM]).

Three ways of drawing a signal

Three panels. Top left, the time domain: a dashed square wave over one period with the partial sums of 1, 3 and 10 harmonics drawn over it, the 10-harmonic curve close to the square with ripples near the edges. Top right, the frequency domain: a line spectrum with vertical lines at f, 3f, 5f, 7f and 9f, their heights falling as 1 over n. Bottom left, the phase domain: axes labelled I and Q with a dashed unit circle and four points, amplitude 1 at 0 degrees, amplitude 0.5 at 0 degrees, amplitude 1 at 90 degrees and amplitude 1 at 180 degrees, each joined to the origin

Figure 75.1 One signal in the time, frequency and phase domains

The time domain plots amplitude against time. It is what an oscilloscope shows, and it makes the shape of a waveform obvious: a square wave looks square.

The frequency domain plots amplitude against frequency. A pure sine wave is one line; a square wave is a comb of lines at its odd harmonics. It makes the bandwidth obvious, and with it whether a channel can carry the signal, and whether the signal will interfere with its neighbours.

The phase domain (a phase state diagram, or constellation) draws each signal as one point: its distance from the origin is the amplitude, its angle is the phase. Written with I and Q, the point is (A cos phi, A sin phi), and a modulation scheme is just a set of allowed points ([Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM]). It makes the choices a modulator has obvious, and how far apart they are, which is how well noise is tolerated.

The three are the same signal. Which one to draw depends on the question: what does it look like, what does it occupy, or what can it mean.

Fourier: every periodic signal is a sum of sine waves

Fourier's theorem: any periodic signal with period T can be written as a constant plus a sum of sine and cosine waves whose frequencies are multiples of the fundamental f = 1 / T. Those multiples are the harmonics.

For a square wave of amplitude 1 and frequency f, the series is:

g(t) = (4 / pi) [ sin(2 pi f t) + sin(2 pi 3 f t) / 3 + sin(2 pi 5 f t) / 5 + ... ]

Only odd harmonics appear, and their amplitudes fall as 1 / n. Two consequences follow, and the program measures both.

A digital signal is wide. Square pulses need harmonics without end. Cut the series off at the ninth harmonic and the shape is recognisable but rounded; a channel that passes only up to 3f gives something closer to a sine wave than a square. This is why a bit rate needs bandwidth, and why the edges of a received digital signal are always rounded.

munotes.in560

Signals: Amplitude, Frequency and Phase

The overshoot does not go away. Adding harmonics reduces the error everywhere except at the jumps, where the partial sum overshoots by a fixed fraction, about 17.9 per cent of the half-amplitude, however many terms are added. The ripple narrows but does not shrink. That is the Gibbs phenomenon, and the program watches it settle.

Bandwidth, and what it allows

Bandwidth has two senses, and both appear in this book.

  • Of a signal: the width of the band of frequencies its spectrum occupies, in hertz. A square wave clipped to its first five harmonics occupies 9f.
  • Of a channel: the width of the band the channel passes without serious attenuation, in hertz. The 2.4 GHz ISM band is 100 MHz wide; one 802.15.4 channel takes about 2 MHz of it.

Nyquist's limit joins the two: a noiseless channel of bandwidth B can carry at most 2B symbols per second. With one bit per symbol that is 2B bits a second; with a modulation that carries k bits per symbol it is 2Bk. Noise sets a further limit, Shannon's, which is why real systems do not reach Nyquist's.

The 802.15.4 radio of [The 802.15.4 Physical Layer] is a worked example: 62.5 ksymbol/s carrying 4 bits per symbol is 250 kb/s, spread over 32 chips per symbol, so 2 Mchip/s, in a channel about 2 MHz wide. The spreading buys robustness with bandwidth, not with rate.

Signals, computed

The program tabulates one sine wave and the effect of doubling its amplitude, shifting its phase by a quarter turn and doubling its frequency; builds the square wave from 1, 2, 3, 5, 10 and 50 harmonics and measures the error away from the jumps and the overshoot at them; turns four channel bandwidths into symbol and bit rates; and prints four signals as amplitude and phase and as I and Q.

# A signal in three ways of looking at it: the sine wave's three parameters,
# Fourier's result that a square wave is a sum of odd harmonics, and what
# bandwidth does to a pulse.
import math

# 1. The sine wave g(t) = A sin(2 pi f t + phi). Three parameters, three effects.
A, f, phi = 1.0, 1000.0, math.pi / 2
print("g(t) = A sin(2 pi f t + phi) with A = %.1f, f = %.0f Hz, phi = pi/2" % (A, f))
print("  period T = 1/f = %.3f ms; the wave repeats %d times a second" % (1000 / f, f))
print("   t (us)   A=1 phi=0   A=2 phi=0   A=1 phi=pi/2   A=1 f=2000")
for t in (0, 125e-6, 250e-6, 375e-6, 500e-6):
    row = (math.sin(2 * math.pi * f * t), 2 * math.sin(2 * math.pi * f * t),
           math.sin(2 * math.pi * f * t + phi), math.sin(2 * math.pi * 2 * f * t))
    print("  %7.0f " % (t * 1e6) + "".join("%12.3f" % v for v in row))

# 2. Fourier: a square wave of frequency f is the sum of its odd harmonics,
#    (4/pi) sum over odd n of sin(2 pi n f t) / n. Adding harmonics one at a
#    time shows the shape appearing, and the error falling.
def partial(t, f, terms):
    return 4 / math.pi * sum(math.sin(2 * math.pi * n * f * t) / n
                             for n in range(1, 2 * terms, 2))

square = lambda t, f: 1.0 if (t * f) % 1.0 < 0.5 else -1.0
print("\nA square wave as a sum of odd harmonics (Fourier):")
print("  harmonics   highest frequency   worst error   average error   overshoot")
for terms in (1, 2, 3, 5, 10, 50):
    samples = [i / 20000 for i in range(20000)]
    errs = [abs(partial(t, 1.0, terms) - square(t, 1.0)) for t in samples]
    peak = max(partial(t, 1.0, terms) for t in samples)
    inner = [e for t, e in zip(samples, errs) if min((t % 0.5), 0.5 - (t % 0.5)) > 0.02]
    print("  %9d %18s %13.3f %15.3f %11.1f%%"
          % (terms, "%d f" % (2 * terms - 1), max(inner), sum(inner) / len(inner), 100 * (peak - 1)))

# 3. Bandwidth. A channel that passes only the harmonics below its bandwidth
#    turns square pulses into rounded ones; the bit rate a channel carries is
#    limited by how many harmonics survive (Nyquist: 2B symbols per second).
print("\nWhat bandwidth allows, for a channel B hertz wide (Nyquist: 2B symbols/s):")
print("     B        symbols/s   with 1 bit/symbol   with 4 bits/symbol (16-QAM)")
for B in (3e3, 200e3, 2e6, 20e6):
    print("  %7s %12.0f %19s %25s"
          % ("%g kHz" % (B / 1e3) if B < 1e6 else "%g MHz" % (B / 1e6), 2 * B,
             "%g kb/s" % (2 * B / 1e3), "%g kb/s" % (8 * B / 1e3)))

# 4. The phase domain: the same sine drawn as one point, amplitude and phase.
print("\nThe phase domain: four signals as (amplitude, phase), and as I and Q")
for name, amp, deg in (("carrier", 1.0, 0), ("half as strong", 0.5, 0),
                       ("quarter turn late", 1.0, 90), ("inverted", 1.0, 180)):
    rad = math.radians(deg)
    print("  %-18s amplitude %.2f, phase %3d degrees -> I = %5.2f, Q = %5.2f"
          % (name, amp, deg, amp * math.cos(rad), amp * math.sin(rad)))
munotes.in561

Signals: Amplitude, Frequency and Phase

g(t) = A sin(2 pi f t + phi) with A = 1.0, f = 1000 Hz, phi = pi/2
  period T = 1/f = 1.000 ms; the wave repeats 1000 times a second
   t (us)   A=1 phi=0   A=2 phi=0   A=1 phi=pi/2   A=1 f=2000
        0        0.000       0.000       1.000       0.000
      125        0.707       1.414       0.707       1.000
      250        1.000       2.000       0.000       0.000
      375        0.707       1.414      -0.707      -1.000
      500        0.000       0.000      -1.000      -0.000

A square wave as a sum of odd harmonics (Fourier):
  harmonics   highest frequency   worst error   average error   overshoot
          1                1 f         0.840           0.293        27.3%
          2                3 f         0.684           0.164        20.0%
          3                5 f         0.535           0.111        18.8%
          5                9 f         0.266           0.067        18.2%
         10               19 f         0.180           0.039        18.0%
         50               99 f         0.050           0.008        17.9%

What bandwidth allows, for a channel B hertz wide (Nyquist: 2B symbols/s):
     B        symbols/s   with 1 bit/symbol   with 4 bits/symbol (16-QAM)
    3 kHz         6000              6 kb/s                   24 kb/s
  200 kHz       400000            400 kb/s                 1600 kb/s
    2 MHz      4000000           4000 kb/s                16000 kb/s
   20 MHz     40000000          40000 kb/s               160000 kb/s

The phase domain: four signals as (amplitude, phase), and as I and Q
  carrier            amplitude 1.00, phase   0 degrees -> I =  1.00, Q =  0.00
  half as strong     amplitude 0.50, phase   0 degrees -> I =  0.50, Q =  0.00
  quarter turn late  amplitude 1.00, phase  90 degrees -> I =  0.00, Q =  1.00
  inverted           amplitude 1.00, phase 180 degrees -> I = -1.00, Q =  0.00
munotes.in562

Signals: Amplitude, Frequency and Phase

The three parameters. At 1 kHz the period is 1 ms, so a quarter of a period is 250 microseconds and the wave reaches 1.000 there. Doubling A doubles every value. Adding pi/2 to the phase shifts the whole wave a quarter period earlier, so it starts at 1.000: a quarter turn in phase is a quarter period in time. Doubling f makes the wave complete two cycles in the same millisecond, so it is back at zero at 250 microseconds. Three parameters, three visibly different signals.

Fourier and Gibbs. One harmonic, the fundamental alone, is a sine wave, wrong by up to 0.840 away from the edges. Three harmonics (up to 5f) bring the worst error to 0.535, ten harmonics (19f) to 0.180, fifty (99f) to 0.050: the average error falls from 0.293 to 0.008, so more bandwidth really does square the corners. The overshoot, though, stops falling: 27.3 per cent with one harmonic, 18.8 per cent with three, 18.0 per cent with ten and 17.9 per cent with fifty. That last figure is the Gibbs constant, and it is why a receiver looking at a sharp edge sees a ring that no amount of bandwidth removes.

Bandwidth into bits. A 3 kHz channel, the width of an old telephone circuit, carries at most 6,000 symbols a second, 6 kb/s with one bit per symbol. A 200 kHz channel, the width of a GSM carrier ([The GSM Radio Interface: Carriers, the TDMA Frame and Bursts]), carries 400 kb/s at one bit per symbol. A 20 MHz channel carries 40 Mb/s, or 160 Mb/s with 16-QAM's 4 bits per symbol. Every jump in rate in the history of radio is one of these two numbers going up: more hertz, or more bits per symbol.

munotes.in563

Signals: Amplitude, Frequency and Phase

The phase domain. The carrier itself sits at I = 1, Q = 0. Halving the amplitude moves it half way in. A quarter turn of phase moves it to I = 0, Q = 1, the quadrature axis. Inverting it, a half turn, moves it to I = -1: that is binary phase shift keying's second symbol, as far from the first as the amplitude allows, which is why BPSK is the most robust keying there is.

Distinctions

Amplitude AFrequency fPhase phi
What it isThe size of the swingCycles per secondWhere in the cycle at t = 0
Measured inVolts, or a relative unitHertzRadians or degrees
Changing it givesAmplitude modulation (ASK)Frequency modulation (FSK)Phase modulation (PSK)
In the phase domainDistance from the originNot shown (the diagram is at one frequency)Angle
Time domainFrequency domainPhase domain
AxesAmplitude against timeAmplitude against frequencyI against Q
A sine wave isA curveOne lineOne point
A square wave isA squareLines at f, 3f, 5f, ...Not usually drawn
Makes obviousThe waveform's shapeBandwidth and interferenceThe symbols and their distance apart
Signal bandwidthChannel bandwidth
Belongs toWhat is transmittedWhat carries it
Example802.15.4's spread signal, about 2 MHzOne 802.15.4 channel, 5 MHz from its neighbours
If too largeInterferes with neighbouring channelsNot applicable
If too smallNot applicableRounds the pulses and limits the rate

What it does not mean

Phase alone carries nothing. It is measured against a reference; a receiver must recover that reference, or use the difference between consecutive symbols.

A square wave is not made of squares. It is made of sine waves; the squareness is what the high harmonics add.

More harmonics do not remove the overshoot. They narrow the ripple at a jump; its height settles at about 17.9 per cent.

Bandwidth is not bit rate. It is hertz; the bit rate depends also on the bits per symbol and on the noise.

The frequency domain is not a different signal. It is the same signal described by its components.

Quick revision

  • g(t) = A sin(2 pi f t + phi); T = 1 / f. Three parameters: amplitude, frequency, phase, and therefore three modulations: ASK, FSK, PSK.
  • Three domains: time (amplitude against time), frequency (the spectrum; bandwidth is visible), phase (a point at distance A and angle phi, or I = A cos phi, Q = A sin phi).
  • Fourier: a periodic signal is a sum of harmonics of f = 1 / T. Square wave = (4/pi) times the sum over odd n of sin(2 pi n f t) / n.
  • A digital signal needs infinite bandwidth; a real channel rounds it. Gibbs: overshoot settles at about 17.9 per cent, however many harmonics.
  • Bandwidth (hertz): of a signal, what it occupies; of a channel, what it passes. Nyquist: at most 2B symbols/s, so 2Bk bits/s with k bits per symbol.
  • Program: worst error 0.840 / 0.535 / 0.180 / 0.050 with 1 / 3 / 10 / 50 harmonics; 3 kHz gives 6 kb/s, 20 MHz gives 40 Mb/s (160 with 16-QAM).
munotes.in564

Signals: Amplitude, Frequency and Phase

Test yourself

1. Write the equation of a sine wave and explain each parameter. A sine wave is g(t) = A sin(2 pi f t + phi). A is the amplitude, the size of the swing, which sets the signal's power; f is the frequency in hertz, the number of cycles per second, whose reciprocal T = 1 / f is the period; and phi is the phase, the position in the cycle at time zero, measured in radians or degrees. These three are the only parameters of the wave, which is why the three basic modulation schemes change one each.

2. What are the time, frequency and phase domains, and what does each show? The time domain plots amplitude against time and shows the shape of the waveform, as an oscilloscope does. The frequency domain plots amplitude against frequency and shows the sine waves the signal is composed of, and therefore its bandwidth and whether it will interfere with neighbouring channels. The phase domain plots the signal as a single point whose distance from the origin is its amplitude and whose angle is its phase, equivalently as its in-phase and quadrature components, and shows the symbols a modulation scheme can send and how far apart they are.

3. State Fourier's result and apply it to a square wave. Any periodic signal of period T can be expressed as a constant plus a sum of sine and cosine waves whose frequencies are integer multiples, the harmonics, of the fundamental frequency f = 1 / T. For a square wave of amplitude 1, only the odd harmonics appear and their amplitudes fall as 1 / n: g(t) = (4 / pi)[sin(2 pi f t) + sin(2 pi 3 f t)/3 + sin(2 pi 5 f t)/5 + ...]. Truncating the series, as any real channel does, rounds the corners of the wave.

munotes.in565

Signals: Amplitude, Frequency and Phase

4. What is the Gibbs phenomenon, and what did the program find? It is the overshoot that appears at a discontinuity when a Fourier series is truncated: adding more harmonics narrows the ripple but does not reduce its height. The program measured the peak of the partial sum of a square wave of amplitude 1 and found overshoots of 27.3 per cent with one harmonic, 18.8 per cent with three, 18.0 per cent with ten and 17.9 per cent with fifty, while the average error away from the jumps fell from 0.293 to 0.008.

5. Define bandwidth and state Nyquist's limit. Bandwidth is the width, in hertz, of the band of frequencies that a signal occupies or that a channel passes. Nyquist's limit states that a noiseless channel of bandwidth B can carry at most 2B symbols per second, so the maximum bit rate is 2B times the number of bits per symbol. A 3 kHz channel therefore carries 6,000 symbols a second, which is 6 kb/s with one bit per symbol and 24 kb/s with four.

6. Why does a digital signal need more bandwidth than an analogue one of the same rate? Because a digital signal is made of sharp transitions, and by Fourier a sharp transition is a sum of harmonics extending to infinity: a square wave contains every odd multiple of its fundamental. A channel passes only a finite band, so it delivers a rounded approximation, and the narrower the channel the more rounded the pulses and the harder it is to tell one symbol from the next. The rate a channel can carry is therefore bounded, at 2B symbols per second, by its bandwidth.

Contents This chapter on its own page

munotes.in566

Chapter Seventy-Six

Antennas: Radiators, Dipoles and Radiation Patterns

Syllabus topic Module 2, "Wireless Transmission: Antennas"

In one line

An antenna turns a current into a radio wave and back; the ideal is the isotropic radiator, which sends equally in every direction and exists only as a reference, so real antennas are measured against it in dBi: a half-wave dipole, the length of half a wavelength, radiates nothing along its wire and most across it, which concentrates its power into 2.15 dBi of gain, and every antenna's pattern, gain and size follow from the wavelength it is built for.

In the wording a student can write in an examination: an antenna is the interface between a guided current and a radio wave; the same antenna transmits and receives with the same pattern (reciprocity). An isotropic radiator is a theoretical point source radiating equally in all directions; its power flux density at distance d is p / (4 pi d squared). It cannot be built, but it is the reference for gain, measured in dBi, decibels relative to isotropic. A half-wave dipole is a conductor half a wavelength long, fed at the centre (two arms of a quarter wavelength each); its pattern is a figure of eight in the plane containing the wire, with nulls along the wire and maxima across it, and a circle in the plane perpendicular to it (an omnidirectional, or doughnut-shaped, pattern). Its gain is 2.15 dBi, or 0 dBd. A quarter-wave monopole is half a dipole worked against a ground plane. A radiation pattern plots radiated power against direction, usually in two perpendicular planes; e.i.r.p., the equivalent isotropically radiated power, is the transmitted power multiplied by the antenna's gain. A receiving antenna's effective aperture is lambda squared over 4 pi times its gain, so for a given gain a higher frequency collects less power.

What an antenna does

A transmitter produces an alternating current; an antenna turns some of it into an electromagnetic wave that leaves the conductor. A receiving antenna does the reverse: the passing wave induces a current. The same structure does both, with the same directional behaviour, which is why a directional antenna that transmits toward a node also hears that node better.

Three quantities describe an antenna:

  • Its size relative to the wavelength. Radiation is efficient when the conductor is a sizeable fraction of a wavelength, which is why [Frequencies for Radio Transmission] mattered before this chapter: the wavelength decides how big the antenna must be.
  • Its pattern. Where the power goes.
  • Its gain. How much more power goes in the best direction than an isotropic radiator would send there.

The isotropic radiator

The isotropic radiator is a point that radiates equally in every direction. Spread a power p over a sphere of radius d and the power flux density is

munotes.in567

Antennas: Radiators, Dipoles and Radiation Patterns

s = p / (4 pi d squared)

which is ITU-R P.525's equation (3), where s is the "power flux-density (W/m2)". It is the inverse square law, and it is the whole of free-space loss.

No such antenna exists: a real radiator always favours some directions. But it is the right reference, because it is the only pattern with no preferred direction, and because the sums are easy. ITU-R P.525 uses it throughout, speaking of the "equivalent isotropically radiated power (e.i.r.p.) of the transmitter in the direction of the point in question". E.i.r.p. is the power an isotropic radiator would need to produce the same flux in the direction the real antenna favours: transmit power times gain.

The half-wave dipole

The practical reference antenna is the half-wave dipole: a straight conductor half a wavelength long, cut in the middle and fed there, so that each arm is a quarter wavelength. At that length the current forms a standing wave with a maximum at the feed and zeros at the ends, and the antenna presents a convenient impedance.

The CC2420 data sheet, which drives 802.15.4 radios, gives the rule directly: "The length of the /2-dipole antenna is given by: L = 14250 / f where f is in MHz, giving the length in cm. An antenna for 2450 MHz should be 5.8 cm. Each arm is therefore 2.9 cm."

Half of a wavelength at 2450 MHz is 6.1 cm, so the data sheet's 5.8 cm is about 95 per cent of it. Real conductors are shortened a little, because the wave travels slightly slower along a wire than in free space. The program computes both.

The monopole. "Monopole antennas are resonant antennas with a length corresponding to one quarter of the electrical wavelength (/4). They are very easy to design and can be implemented simply as a 'piece of wire' or even integrated into the PCB." A monopole is one arm of a dipole working against a ground plane, which mirrors it; the CC2420 sheet notes that a monopole is single-ended and so needs a balun between it and the radio's differential output, while "A differential antenna like a dipole would be the easiest to interface not needing a balun".

On a mote. Telos carries neither: "Telos uses an internal 2.4GHz Planar Inverted Folded Antenna (PIFA) built into the printed circuit board and tuned to match the radio circuitry. An optional SMA coax connection may be used instead of the internal antenna. Integration of the antenna lowers the overall cost of the mote since no expensive external antennae are needed." A folded, printed antenna fits the board and costs nothing to make; it pays in gain and in a pattern that the board and its battery distort.

munotes.in568

Antennas: Radiators, Dipoles and Radiation Patterns

Radiation patterns

A radiation pattern plots the power radiated against direction. Space has two angles, so a pattern is a surface; on paper it is drawn as two perpendicular slices, usually the vertical (elevation) plane and the horizontal (azimuth) plane, or, for a dipole, the plane containing the wire and the plane across it.

Three polar plots. Left, the isotropic radiator: a perfect circle around a point, the same in every direction, 0 dBi. Centre, a half-wave dipole in the plane containing its wire, drawn as a vertical bar: a figure of eight with nulls along the wire and maxima across it. Right, the same dipole in the plane across the wire, dashed, a circle, with a directional antenna's pattern over it: one large main lobe to one side and small side lobes. A note says the pattern is the same for transmitting and receiving, and that gain in dBi compares the best direction with an isotropic radiator, the dipole being 2.15 dBi

Figure 76.1 Radiation patterns: isotropic, a half-wave dipole in two planes, and a directional antenna

For the half-wave dipole, the relative power at an angle theta from the wire is

(cos((pi/2) cos theta) / sin theta) squared

which is zero along the wire (theta = 0 or 180 degrees) and greatest across it (theta = 90 degrees). In the plane containing the wire, that traces the figure of eight; in the plane across the wire, the power is the same in every direction, so it traces a circle. Together they make the familiar doughnut. That is what omnidirectional means in practice: uniform in one plane, not in space.

How to read a pattern. The main lobe is the direction of greatest radiation; side lobes are the smaller maxima; a null is a direction of no radiation. Beam width is usually quoted where the power has fallen to half, that is by 3 dB. A dipole's null matters in a sensor network: a mote standing upright is deaf to a mote directly above it.

Gain, dBi and dBd

Gain is the ratio of the power an antenna sends in its best direction to the power an isotropic radiator would send there for the same input. It is a consequence of the pattern and not of any amplification: an antenna has no power of its own, so gain in one direction is paid for by loss in another.

  • dBi: gain relative to an isotropic radiator.
  • dBd: gain relative to a half-wave dipole. Since the dipole itself is 2.15 dBi, dBi = dBd + 2.15.

The 2.15 dBi is not a convention; it is the average of the dipole's own pattern over the sphere, and the program computes it by integration, arriving at 1.641, which is 2.15 dB.

Gain and range. Received power falls as the square of distance in free space, so a gain of 6 dB, a factor of four in power, doubles the range at which a given power arrives; 14 dBi multiplies it by five. Gain and coverage move the other way: a beam of gain G covers roughly 41253 / G square degrees, so 14 dBi lights up about 4 per cent of the sky. [Directional Antennas, Sectorisation, Diversity and Spatial Reuse] is about spending that trade deliberately.

munotes.in569

Antennas: Radiators, Dipoles and Radiation Patterns

The receiving side: effective aperture

An antenna's ability to collect power is its effective aperture, an area. ITU-R P.525 gives it for the isotropic case in equation (4): the received power is the flux density times "a: effective aperture of a receiving isotropic antenna (m2)", and that aperture is lambda squared over 4 pi. For an antenna of gain G it is G times that.

The consequence is important and often surprising: for the same gain, a higher frequency collects less power, because the aperture shrinks with the square of the wavelength. That is the wavelength term in free-space loss ([Signal Propagation: Ranges, Path Loss and How a Signal Travels]), and part of why a 2.4 GHz link is shorter than an 868 MHz one at the same power.

Antennas, computed

The program computes dipole and monopole lengths for five bands and checks the data sheet's rule; follows an isotropic milliwatt out to 100 m and computes what an isotropic antenna collects at 2450 MHz; tabulates the half-wave dipole's pattern and integrates it over the sphere to get its gain; and turns gain into range and into beam area.

# Antennas: the isotropic radiator as the reference, dipole lengths, the
# radiation pattern of a half-wave dipole, and what gain means in decibels.
import math

C = 299792458.0

# 1. Dipole and monopole lengths. The CC2420 data sheet gives the half-wave
#    dipole as L = 14250 / f cm with f in MHz; that is 0.95 of a half wavelength,
#    the usual shortening for the wave's slower speed along a real conductor.
print("Antenna lengths, and the data sheet's rule L = 14250 / f cm:")
print("   frequency   wavelength   half-wave   each arm   quarter-wave   data sheet")
for mhz in (433.92, 868.3, 915.0, 2450.0, 5800.0):
    lam = C / (mhz * 1e6) * 100                        # in cm
    print("  %7.1f MHz %10.1f cm %9.1f cm %8.1f cm %12.1f cm %10.1f cm"
          % (mhz, lam, lam / 2, lam / 4, lam / 4, 14250 / mhz))
print("  the rule is %.3f of a half wavelength" % (14250 / 2450 / (C / 2450e6 * 100 / 2)))

# 2. What an isotropic radiator means. ITU-R P.525: the power flux density at
#    distance d is p / (4 pi d^2), and the effective aperture of an isotropic
#    receiving antenna is lambda^2 / (4 pi).
print("\nAn isotropic radiator of 1 mW (0 dBm), and what an isotropic antenna collects:")
print("   distance   flux density (W/m2)   received at 2450 MHz    in dBm")
lam = C / 2450e6
for d in (1, 10, 30, 100):
    s = 1e-3 / (4 * math.pi * d ** 2)
    pr = s * lam ** 2 / (4 * math.pi)
    print("  %6d m %20.3e %20.3e %9.1f" % (d, s, pr, 10 * math.log10(pr / 1e-3)))

# 3. The radiation pattern of a half-wave dipole, in the plane containing the
#    wire: the power at an angle theta from the wire, relative to the maximum.
#    (cos(pi/2 cos theta) / sin theta)^2, the standard result, computed here.
def dipole(theta):
    if theta % math.pi == 0:
        return 0.0
    return (math.cos(math.pi / 2 * math.cos(theta)) / math.sin(theta)) ** 2

print("\nHalf-wave dipole, power against the angle from the wire:")
print("   angle   relative power   in dB   compare: isotropic")
for deg in (0, 15, 30, 45, 60, 75, 90):
    p = dipole(math.radians(deg))
    print("  %4d deg %13.3f %8s %18.3f"
          % (deg, p, "%.1f" % (10 * math.log10(p)) if p > 0 else "none", 1.0))

# The gain of the dipole: its peak power divided by the average over the sphere.
steps = 20000
avg = sum(dipole(math.pi * (i + 0.5) / steps) * math.sin(math.pi * (i + 0.5) / steps)
          for i in range(steps)) / steps * math.pi / 2
gain = 1.0 / avg
print("  averaged over the whole sphere, the dipole's gain is %.3f, that is %.2f dBi"
      % (gain, 10 * math.log10(gain)))

# 4. Gain in decibels, and what it buys. Doubling the range needs 6 dB.
print("\nWhat antenna gain buys, at 2450 MHz with 0 dBm into the antenna:")
print("   gain    e.i.r.p.   range for the same received power (relative)")
for dbi in (0, 2.15, 6, 9, 14):
    print("  %5.2f dBi %7.2f dBm %36.2f x" % (dbi, dbi, 10 ** (dbi / 20)))
print("  a beam with gain G covers roughly 41253 / G square degrees of sky:")
for dbi in (2.15, 6, 9, 14):
    g = 10 ** (dbi / 10)
    print("    %5.2f dBi: gain %5.2f, about %6.0f square degrees, %.1f%% of the sphere"
          % (dbi, g, 41253 / g, 100 / g))
munotes.in570

Antennas: Radiators, Dipoles and Radiation Patterns

Antenna lengths, and the data sheet's rule L = 14250 / f cm:
   frequency   wavelength   half-wave   each arm   quarter-wave   data sheet
    433.9 MHz       69.1 cm      34.5 cm     17.3 cm         17.3 cm       32.8 cm
    868.3 MHz       34.5 cm      17.3 cm      8.6 cm          8.6 cm       16.4 cm
    915.0 MHz       32.8 cm      16.4 cm      8.2 cm          8.2 cm       15.6 cm
   2450.0 MHz       12.2 cm       6.1 cm      3.1 cm          3.1 cm        5.8 cm
   5800.0 MHz        5.2 cm       2.6 cm      1.3 cm          1.3 cm        2.5 cm
  the rule is 0.951 of a half wavelength

An isotropic radiator of 1 mW (0 dBm), and what an isotropic antenna collects:
   distance   flux density (W/m2)   received at 2450 MHz    in dBm
       1 m            7.958e-05            9.482e-08     -40.2
      10 m            7.958e-07            9.482e-10     -60.2
      30 m            8.842e-08            1.054e-10     -69.8
     100 m            7.958e-09            9.482e-12     -80.2

Half-wave dipole, power against the angle from the wire:
   angle   relative power   in dB   compare: isotropic
     0 deg         0.000     none              1.000
    15 deg         0.043    -13.7              1.000
    30 deg         0.175     -7.6              1.000
    45 deg         0.394     -4.0              1.000
    60 deg         0.667     -1.8              1.000
    75 deg         0.904     -0.4              1.000
    90 deg         1.000      0.0              1.000
  averaged over the whole sphere, the dipole's gain is 1.641, that is 2.15 dBi

What antenna gain buys, at 2450 MHz with 0 dBm into the antenna:
   gain    e.i.r.p.   range for the same received power (relative)
   0.00 dBi    0.00 dBm                                 1.00 x
   2.15 dBi    2.15 dBm                                 1.28 x
   6.00 dBi    6.00 dBm                                 2.00 x
   9.00 dBi    9.00 dBm                                 2.82 x
  14.00 dBi   14.00 dBm                                 5.01 x
  a beam with gain G covers roughly 41253 / G square degrees of sky:
     2.15 dBi: gain  1.64, about  25145 square degrees, 61.0% of the sphere
     6.00 dBi: gain  3.98, about  10362 square degrees, 25.1% of the sphere
     9.00 dBi: gain  7.94, about   5193 square degrees, 12.6% of the sphere
    14.00 dBi: gain 25.12, about   1642 square degrees, 4.0% of the sphere
munotes.in571

Antennas: Radiators, Dipoles and Radiation Patterns

Lengths. At 2450 MHz the wavelength is 12.2 cm, so a half-wave dipole is 6.1 cm with arms of 3.1 cm, and the data sheet's rule gives 5.8 cm: the rule is 0.951 of a half wavelength, the usual shortening. At 433.9 MHz the dipole is 34.5 cm, too long for a small device, which is why sub-gigahertz sensor nodes usually use a quarter-wave wire or a printed antenna instead.

The inverse square law. One milliwatt spread over a sphere gives 79.6 microwatts per square metre at 1 m and 8.0 nanowatts at 100 m. An isotropic antenna at 2450 MHz, whose aperture is 1.2 square centimetres, collects -40.2 dBm at 1 m and -80.2 dBm at 100 m. The CC2420's sensitivity is about -95 dBm ([The 802.15.4 Physical Layer]), so in free space, with plain antennas, the link closes with about 15 dB to spare at 100 m; indoors it does not, which is what the next chapters are about.

The dipole's pattern and gain. Across the wire the power is 1.000; at 60 degrees from the wire 0.667, which is -1.8 dB; at 30 degrees 0.175, that is -7.6 dB; at 15 degrees 0.043, -13.7 dB; and along the wire, nothing at all. Averaged over the whole sphere the dipole radiates about three fifths of its peak, so its gain is 1.641, and ten times the logarithm of that is 2.15 dBi: the number every antenna is quoted against, derived from the pattern.

What gain buys. 6 dBi doubles the range for the same received power, 14 dBi multiplies it by 5.01. The same 14 dBi narrows the coverage to about 1,642 square degrees, 4 per cent of the sphere. Gain is not power; it is aim.

munotes.in572

Antennas: Radiators, Dipoles and Radiation Patterns

Distinctions

Isotropic radiatorHalf-wave dipoleQuarter-wave monopole
ExistsNo: a referenceYesYes, over a ground plane
LengthA pointHalf a wavelength (two arms of a quarter)A quarter wavelength
PatternA sphereFigure of eight in the wire's plane, circle across itHalf the dipole's, above the plane
Gain0 dBi by definition2.15 dBi (0 dBd)About 5.15 dBi over a perfect plane
At 2450 MHzNot applicable6.1 cm (data sheet: 5.8 cm)3.1 cm
GainPower
What it isConcentration into a directionEnergy per second from the transmitter
UnitsdBi or dBddBm or W
Where it comes fromThe patternThe power amplifier
CostsCoverage in other directionsBattery
Togethere.i.r.p. = transmit power + gain, in dB
TransmittingReceiving
What the antenna doesTurns current into a waveTurns a wave into current
Described byGain and patternEffective aperture (and the same pattern)
Frequency dependenceSize falls with wavelengthAperture falls with the square of the wavelength

What it does not mean

Gain is not amplification. An antenna is passive: what it adds in one direction it takes from another.

Omnidirectional is not spherical. A dipole is uniform in one plane only, and has deep nulls along its wire.

A longer antenna is not always better. It must match the wavelength; a half-wave dipole at the wrong frequency radiates poorly.

Higher frequency does not collect more. For a given gain the aperture goes as the square of the wavelength, so a 2.4 GHz antenna collects a ninth of what an 868 MHz one of the same gain does.

A pattern is not a range map. It shows relative power by direction; the range also depends on power, sensitivity and what is in the way.

dBi and dBd are not interchangeable. They differ by 2.15 dB, which is a factor of 1.64 in power.

Quick revision

  • Antenna: current to wave and back; the same pattern transmitting and receiving.
  • Isotropic radiator: equal in all directions, s = p / (4 pi d squared); the reference for dBi; e.i.r.p. = power x gain.
  • Half-wave dipole: lambda / 2 long, two arms of lambda / 4; data sheet rule L = 14250 / f cm (f in MHz), 5.8 cm at 2450 MHz, about 0.95 of a half wavelength; pattern figure of eight in the wire's plane (nulls along the wire), circle across it; gain 2.15 dBi = 0 dBd.
  • Quarter-wave monopole: lambda / 4 over a ground plane; needs a balun with a differential radio. Telos uses an internal PIFA on the board.
  • Pattern: power against direction, drawn in two planes; main lobe, side lobes, nulls, beam width at 3 dB.
  • Effective aperture = G lambda squared / (4 pi): a higher frequency collects less.
  • Program: at 2450 MHz an isotropic antenna collects -40.2 dBm at 1 m, -80.2 dBm at 100 m; the dipole's gain integrates to 1.641, which is 2.15 dBi; 6 dBi doubles the range, 14 dBi gives 5.01 times it but covers only 4 per cent of the sphere.
munotes.in573

Antennas: Radiators, Dipoles and Radiation Patterns

Test yourself

1. What is an isotropic radiator, and why is it used although it cannot be built? It is a theoretical point source that radiates equally in every direction, so the power flux density at a distance d from a radiator of power p is p divided by 4 pi d squared. No real antenna has that pattern, since a real radiator always favours some directions, but it is the only pattern with no preferred direction, which makes it the natural reference: antenna gains are quoted in dBi, decibels relative to isotropic, and transmitted powers as equivalent isotropically radiated power.

2. Describe a half-wave dipole and compute its length at 900 MHz and 2.45 GHz. It is a straight conductor half a wavelength long, cut at the centre and fed there, so each arm is a quarter wavelength. At 900 MHz the wavelength is about 33.3 cm, so the dipole is about 16.7 cm with arms of 8.3 cm; at 2.45 GHz the wavelength is 12.2 cm, so the dipole is 6.1 cm with arms of 3.1 cm. Practical dipoles are cut about 5 per cent shorter, which is the CC2420 data sheet's rule of 14250 / f cm, giving 5.8 cm at 2450 MHz.

3. Draw and explain the radiation pattern of a half-wave dipole. In the plane containing the wire the pattern is a figure of eight: the radiated power is proportional to (cos((pi/2) cos theta) / sin theta) squared, where theta is the angle from the wire, so it is zero along the wire and greatest at right angles to it, falling to 0.667 of the maximum at 60 degrees and 0.175 at 30 degrees. In the plane perpendicular to the wire the power is the same in every direction, so the pattern is a circle. Together they form a doughnut shape around the wire, which is what is meant by calling a dipole omnidirectional.

4. Define antenna gain, and explain dBi and dBd. Gain is the ratio of the power an antenna radiates in its strongest direction to the power an isotropic radiator would send in that direction for the same input; it comes from the shape of the pattern, not from amplification, so gain in one direction is paid for by less power elsewhere. It is expressed in dBi, decibels relative to an isotropic radiator, or in dBd, relative to a half-wave dipole. Since the dipole's own gain is 2.15 dBi, a gain in dBi is the gain in dBd plus 2.15.

munotes.in574

Antennas: Radiators, Dipoles and Radiation Patterns

5. Show that a half-wave dipole's gain is 2.15 dBi. Its relative power at an angle theta from the wire is (cos((pi/2) cos theta) / sin theta) squared, with a peak of 1 across the wire. Averaging that over the whole sphere, weighting each angle by sin theta for the area of the corresponding ring, gives an average of about 0.609 of the peak. The gain is the peak divided by that average, which comes to 1.641, and ten times its logarithm to base ten is 2.15 dB, so the dipole's gain is 2.15 dBi.

6. What is effective aperture, and what does it imply about frequency? The effective aperture is the area from which a receiving antenna effectively collects power from a passing wave: the received power is the power flux density times the aperture. For an isotropic antenna it is the wavelength squared divided by 4 pi, and for an antenna of gain G it is G times that. Since it falls with the square of the wavelength, an antenna of the same gain at a higher frequency collects less power, which is one reason why higher-frequency links have shorter range at the same transmit power.

7. What antenna does a Telos mote use, and why? An internal 2.4 GHz planar inverted folded antenna built into the printed circuit board and tuned to the radio, with an optional SMA connector for an external antenna instead. It is used because integrating the antenna into the board removes the cost of an external antenna and makes the mote a single robust device; the price is lower gain and a pattern distorted by the board and its batteries, compared with a proper dipole.

Contents This chapter on its own page

munotes.in575

Chapter Seventy-Seven

Directional Antennas, Sectorisation, Diversity and Spatial Reuse

Syllabus topic Module 2, "Wireless Transmission: Antennas" (and the paired practical, "MANET Simulation with Directional Antenna: Simulate MANET using directional antennas and analyze improvements in spatial reuse and interference reduction")

In one line

A directional antenna does not create power, it aims it, and aiming is worth more than it looks: a narrow beam reaches further for the same transmitter, hears fewer interferers, and above all lets other pairs of nodes talk at the same time, which is spatial reuse; sectorising a site multiplies its capacity by the number of sectors, two antennas a few centimetres apart rescue a signal lost to fading, and a smart antenna steers its beam toward the wanted signal and a null toward the unwanted one.

In the wording a student can write in an examination: a directed (directional) antenna concentrates its radiation into a main lobe; its gain is the ratio of the power in that direction to an isotropic radiator's, so a narrower beam gives more gain, more range for the same power, and less interference to and from other directions. A sectorised antenna divides a site into sectors, typically three of 120 degrees or six of 60 degrees, each with its own directed antenna and its own channels, so one site can carry that many times the traffic; sectorisation is a standard step in raising cellular capacity ([Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing]).

Diversity uses more than one antenna to defeat fading ([Multipath, Fading and the Doppler Effect]), since two antennas a few centimetres apart rarely fade at the same instant. In switched (selection) diversity the receiver measures each antenna and uses the better one; in combining diversity it adds the signals from several antennas. Other forms are space, polarisation, frequency and time diversity. Smart (adaptive) antennas are arrays whose elements are fed with controlled phases, so the beam can be steered at a wanted signal and a null placed on an interferer, and the pattern changed as either moves.

Spatial reuse is what all of this is for: the same frequency used at the same moment by two pairs far enough apart, or aimed away from each other, so the capacity of an area grows with the number of simultaneous conversations rather than with the bandwidth alone.

A directed antenna: aiming, not amplifying

An antenna is passive, so its gain is aim, as [Antennas: Radiators, Dipoles and Radiation Patterns] showed: what is added in the main lobe is taken from everywhere else. Concentrating a whole sphere's radiation into a beam of b degrees would, in the ideal case, multiply the power in that beam by 360 / b in one plane, and since received power falls as the square of distance, it multiplies the range by the square root of the gain.

Four things change at once when the beam narrows:

  • Range grows for the same transmit power.
  • Interference caused falls: fewer nodes hear the transmission at all.
  • Interference suffered falls: a directional receiver hears only what lies in its beam.
  • Coverage falls: the nodes outside the beam are not reached, so the antenna must be aimed, by hand or by a smart antenna.
munotes.in576

Directional Antennas, Sectorisation, Diversity and Spatial Reuse

The CC2420 data sheet's list is the practical end of this: a dipole, a monopole, a helical antenna ("a good compromise in size critical applications") or a loop, each with a different pattern and size, and the choice is made by what has to be reached.

Sectorisation

A sectorised site is several directed antennas back to back, each covering its own sector and each given its own channels. Three antennas of 120 degrees, or six of 60, are the usual arrangements.

Two gains follow. Each antenna has gain (10 log of 360 divided by the sector's width, ideally: 4.77 dB for 120 degrees, 7.78 dB for 60), which extends the cell or lets the transmitter run at less power. And the site's capacity multiplies: with three sectors, three sets of channels can be in use at once where one was before. The cost is that a station near a boundary is served by one sector and interferes with the next, and that a moving station must be handed over between sectors as well as between cells ([Handover in GSM]).

Diversity

Fading is the enemy this answers. A signal arriving by several paths can cancel itself at one point in space and be strong a few centimetres away ([Multipath, Fading and the Doppler Effect]). Diversity is the use of two or more copies of the signal that are unlikely to fade at the same time.

DECT's own definition, from its overview standard: "antenna diversity: diversity implies that the Radio Fixed Part (RFP) for each bearer independently can select different antenna properties such as gain, polarization, coverage patterns and other features that may affect the practical coverage", with the note that "A typical example is space diversity, provided by two vertically polarized antennas separated by 10 cm to 20 cm."

Switched (selection) diversity picks the better antenna. DECT describes exactly how, and what it is worth: "Prolonged preamble transmissions are intended to be used in combination with a preamble switched antenna instant receiver selection diversity algorithm implemented at the receiving end. This algorithm implies that the receiver during a first part of the preamble makes a first link quality estimate using one antenna, and during a second part makes a second estimate using the other antenna, and then, for a third (the last) part of the preamble and the rest of the packet, selects the antenna which gave the highest quality estimate." The gain is large: "This algorithm can provide a performance improvement corresponding to 10 dB increased link budget in a mobile or moving environment." And it is cheap, which is the point: "Traditional means to provide this performance improvement requires two complete radio receivers. The prolonged preamble helps implementing low-cost and efficient means for the quality estimates needed in the algorithm."

munotes.in577

Directional Antennas, Sectorisation, Diversity and Spatial Reuse

DECT's link budget assumes it: its fading margin table notes that "Normally switched antenna diversity is implemented, whereby MF = 11 dB in a quasi-stationary environment."

Combining diversity uses all the copies at once instead of choosing: the signals from several antennas are added, in the best case weighted by their strength. It performs better than selection and costs more, since every antenna needs its own receiver chain.

Other kinds. Space diversity uses separated antennas; polarisation diversity uses two polarisations at one place; frequency diversity sends on two frequencies (frequency hopping is a moving form of it, [Frequency Hopping Spread Spectrum]); time diversity repeats or interleaves in time. All work for the same reason: two copies that fade independently.

Smart antennas

A smart or adaptive antenna is an array of elements whose signals are combined with controlled phases and amplitudes. Because the pattern is computed rather than built, it can be changed:

  • Beam steering: point the main lobe at the wanted station and follow it as it moves.
  • Null steering: place a null of the pattern on an interferer, which removes it far more effectively than any filter.
  • Per-user beams: a base station can hold several beams at once, giving each user the gain of a directional antenna and reusing channels within one cell (space division multiple access, in the terms of [Multiplexing: Space, Frequency, Time and Code]).

The cost is computation and calibration, and a sensor node has neither, which is why smart antennas belong to base stations and not to motes.

Spatial reuse: the practical's question

MU's practical asks for exactly the measurement this chapter is for: "Simulate MANET using directional antennas and analyze improvements in spatial reuse and interference reduction."

Spatial reuse means two transmissions on the same frequency at the same time, far enough apart or aimed away from each other that neither spoils the other. It is the only way the capacity of an area grows without more spectrum: a network of omnidirectional radios wastes most of every transmission, since the energy that reaches nodes other than the intended receiver does nothing but silence them ([Hidden and Exposed Terminals, and RTS and CTS] is the same problem seen from the MAC).

The program measures it. Sixty nodes stand in a 200 m square with an omnidirectional range of 90 m. Each round, every node offers a transmission to a neighbour; a transmission runs if no already running transmission reaches its receiver. With directional antennas, whether a transmission reaches a node depends on the beam as well as the distance, and a narrow beam also reaches further.

munotes.in578

Directional Antennas, Sectorisation, Diversity and Spatial Reuse

Beams, computed

# The practical's question, computed: what does narrowing the beam buy? A field
# of nodes with omnidirectional antennas, then with sectored ones, then with
# beams steered at the receiver; how many transmissions can run at once, and
# how much interference each one suffers.
import math
import random

rnd = random.Random(77)
SIDE, N, RANGE = 200.0, 60, 90.0
nodes = [(rnd.uniform(0, SIDE), rnd.uniform(0, SIDE)) for _ in range(N)]
dist = lambda a, b: math.hypot(a[0] - b[0], a[1] - b[1])
bearing = lambda a, b: math.atan2(b[1] - a[1], b[0] - a[0])

def within(beam, tx, rx, other):
    """Does a transmission from tx to rx also reach `other`? An omnidirectional
    antenna (beam = 360) reaches everything in range; a beam of `beam` degrees,
    aimed at rx, reaches only what lies inside it. A narrower beam also carries
    further for the same power: gain 360/beam, and range grows as its square root."""
    gain = 360.0 / beam
    reach = RANGE * gain ** 0.5
    if dist(tx, other) > reach:
        return False
    if beam >= 360:
        return True
    off = abs((bearing(tx, other) - bearing(tx, rx) + math.pi) % (2 * math.pi) - math.pi)
    return off <= math.radians(beam) / 2

def trial(beam, receiving_too, rounds=400):
    """Each round, every node offers one transmission to a neighbour in range.
    A transmission succeeds if no other running transmission reaches its
    receiver (and, when receiving_too, only transmissions inside the receiver's
    own beam can spoil it). Count how many run at once."""
    total, hurt = 0, 0
    for _ in range(rounds):
        pairs = []
        for i, tx in enumerate(nodes):
            near = [j for j, p in enumerate(nodes) if j != i and dist(tx, p) <= RANGE]
            if near:
                pairs.append((i, rnd.choice(near)))
        rnd.shuffle(pairs)
        running = []
        for tx, rx in pairs:
            clash = False
            for otx, orx in running:
                if orx == rx or otx == rx or otx == tx or orx == tx:
                    clash = True
                    break
                if within(beam, nodes[otx], nodes[orx], nodes[rx]):
                    if not receiving_too:
                        clash = True
                        break
                    off = abs((bearing(nodes[rx], nodes[otx]) - bearing(nodes[rx], nodes[tx])
                               + math.pi) % (2 * math.pi) - math.pi)
                    if off <= math.radians(beam) / 2:
                        clash = True
                        break
            if clash:
                hurt += 1
            else:
                running.append((tx, rx))
        total += len(running)
    return total / rounds, hurt / rounds

print("%d nodes over %.0f m by %.0f m, omnidirectional range %.0f m." % (N, SIDE, SIDE, RANGE))
print("A beam of b degrees has gain 360/b, so it reaches sqrt(360/b) times as far.")
print("\n  beam      reach   transmissions at once   blocked   only the sender aims")
for beam in (360, 180, 90, 60, 30, 15):
    both, hurt = trial(beam, True)
    one, _ = trial(beam, False)
    print("  %3d deg %7.0f m %19.1f %9.1f %19.1f"
          % (beam, RANGE * (360 / beam) ** 0.5, both, hurt, one))

# How many neighbours a transmission disturbs, which is what spatial reuse means.
print("\nNodes reached by one transmission (the neighbours it silences):")
for beam in (360, 180, 90, 60, 30, 15):
    counts = []
    for i, tx in enumerate(nodes):
        near = [j for j, p in enumerate(nodes) if j != i and dist(tx, p) <= RANGE]
        if not near:
            continue
        rx = near[0]
        counts.append(sum(1 for j, p in enumerate(nodes)
                          if j not in (i, rx) and within(beam, tx, nodes[rx], p)))
    print("  %3d deg: on average %.1f of %d nodes" % (beam, sum(counts) / len(counts), N - 2))

# Sectorisation: one site with k sectors reuses its channels k times.
print("\nSectorisation at one site, three channels to give out:")
for k in (1, 3, 6):
    print("  %d sector(s) of %3d degrees: %d simultaneous channels, each antenna %.2f dBi over isotropic"
          % (k, 360 // k, 3 * k, 10 * math.log10(k)))
munotes.in579

Directional Antennas, Sectorisation, Diversity and Spatial Reuse

60 nodes over 200 m by 200 m, omnidirectional range 90 m.
A beam of b degrees has gain 360/b, so it reaches sqrt(360/b) times as far.

  beam      reach   transmissions at once   blocked   only the sender aims
  360 deg      90 m                 4.9      55.1                 4.9
  180 deg     127 m                 9.5      50.5                 5.4
   90 deg     180 m                13.8      46.2                 5.8
   60 deg     220 m                16.9      43.1                 6.7
   30 deg     312 m                20.7      39.3                10.4
   15 deg     441 m                21.9      38.1                14.6

Nodes reached by one transmission (the neighbours it silences):
  360 deg: on average 20.2 of 58 nodes
  180 deg: on average 20.6 of 58 nodes
   90 deg: on average 19.8 of 58 nodes
   60 deg: on average 15.7 of 58 nodes
   30 deg: on average 8.6 of 58 nodes
   15 deg: on average 4.2 of 58 nodes

Sectorisation at one site, three channels to give out:
  1 sector(s) of 360 degrees: 3 simultaneous channels, each antenna 0.00 dBi over isotropic
  3 sector(s) of 120 degrees: 9 simultaneous channels, each antenna 4.77 dBi over isotropic
  6 sector(s) of  60 degrees: 18 simultaneous channels, each antenna 7.78 dBi over isotropic

Spatial reuse. With omnidirectional antennas, 4.9 transmissions run at once out of the roughly 60 offered. Halving the beam to 180 degrees nearly doubles that to 9.5; 90 degrees gives 13.8, 60 degrees 16.9, and 30 degrees 20.7, more than four times the omnidirectional figure. The blocked count falls in step, from 55.1 to 39.3. That is the practical's answer: narrowing the beam from 360 to 30 degrees roughly quadrupled the number of conversations the same field could hold at the same instant.

munotes.in580

Directional Antennas, Sectorisation, Diversity and Spatial Reuse

Both ends must aim. The last column repeats the experiment with the beam only at the sender, the receiver still listening in every direction. At 90 degrees that gives 5.8 simultaneous transmissions against 13.8 when both ends aim. Most of the gain comes from the receiver ignoring what is not in front of it, which is worth knowing before buying antennas for one end of a link.

The surprise in the middle. The second table counts how many nodes one transmission reaches, that is how many it silences. Going from 360 to 180 degrees does not reduce it: 20.2 against 20.6. The beam covers half the angle but reaches 1.41 times as far, and in a field of this density the extra area exactly cancels the narrower angle. Only from about 60 degrees does the count fall (15.7), and at 15 degrees it is 4.2. The lesson is that gain is not free: a narrow beam that is allowed to shout further can interfere with as many nodes as a wide one. Spatial reuse improves from 180 degrees anyway, because the nodes reached are in a different place for each transmission, but the naive claim that a narrow beam always disturbs fewer neighbours is false, and the model shows why.

Sectorisation. At one site with three channels, one omnidirectional antenna supports three simultaneous channels; three sectors of 120 degrees support nine, with 4.77 dB of gain each; six sectors of 60 degrees support eighteen, with 7.78 dB. Capacity multiplies with the sectors, which is why the first answer to a congested cell is to sectorise it.

Distinctions

OmnidirectionalDirectedSectorisedSmart (adaptive)
PatternUniform in one planeOne main lobe, fixedSeveral fixed lobes from one siteComputed, steerable
AimingNone neededBy hand, onceBy constructionAutomatic, per user
Gain0 to about 2 dBiUp to tens of dBi4.77 dB (120 deg), 7.78 dB (60 deg) ideallyAs the array allows
SuitsNodes that move or have many neighboursFixed linksBase stationsBase stations
CostNoneMust be pointedSeveral antennas and feedsArray, processing, calibration
Switched (selection) diversityCombining diversity
What it doesMeasures each antenna, uses the bestAdds the signals from all antennas
HardwareOne receiver and a switchOne receiver chain per antenna
DECT's figureAbout 10 dB of link budgetNot specified there
ComplexityLow: DECT does it in the preambleHigher
More bandwidthMore spatial reuse
GivesA faster linkMore links at once
CostsSpectrum, which is allocatedAntennas and aiming
LimitRegulationGeometry and interference
ProgramNot modelled4.9 transmissions at 360 deg, 20.7 at 30 deg
munotes.in581

Directional Antennas, Sectorisation, Diversity and Spatial Reuse

What it does not mean

A directional antenna does not add power. It takes power from the directions it does not serve.

A narrower beam does not always disturb fewer nodes. If its extra gain is spent on range, it can reach as many: the program's 180-degree beam silenced as many nodes as the omnidirectional one.

Diversity is not repetition. The copies must fade independently, which is why the antennas are separated, or the polarisations or frequencies differ.

Sectorisation is not cell splitting. Sectors divide one site's coverage by angle; splitting adds new sites ([Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing]).

Smart antennas are not for motes. They need an array and continuous computation; a sensor node has a printed antenna and a battery.

Spatial reuse is not free capacity. It needs the geometry to cooperate, and a MAC that lets two nearby transmissions proceed instead of deferring to each other.

Quick revision

  • Directed antenna: one main lobe; gain about 360 / beam; range grows as the square root of the gain; must be aimed.
  • Sectorised: three sectors of 120 degrees or six of 60 degrees at one site; capacity multiplies by the sectors; ideal gains 4.77 dB and 7.78 dB.
  • Diversity: several copies unlikely to fade together. Switched (selection): measure and pick the better antenna; DECT does it in a prolonged preamble and calls it worth "10 dB increased link budget", with a fading margin of 11 dB assumed. Combining: add them all, better but one receiver per antenna. Kinds: space (10 to 20 cm apart), polarisation, frequency, time.
  • Smart (adaptive) antennas: arrays that steer the beam and place nulls; several beams at once (space division); for base stations, not motes.
  • Spatial reuse: same frequency, same moment, different places or directions.
  • Program (the practical): simultaneous transmissions 4.9 / 9.5 / 13.8 / 16.9 / 20.7 / 21.9 at 360 / 180 / 90 / 60 / 30 / 15 degrees; only 5.8 at 90 degrees if the receiver does not aim; nodes silenced per transmission 20.2 at 360 degrees and 20.6 at 180, falling to 4.2 at 15.

Test yourself

1. How does a directional antenna improve a wireless network, and what does it cost? It concentrates the radiated power into a main lobe, so for the same transmitter it reaches further, it is heard by fewer nodes in other directions and, used at the receiver, it hears fewer interferers. That allows more transmissions to share the same frequency at the same time, which is spatial reuse. The costs are that it must be aimed, that nodes outside the beam are not served, and, for an adaptive antenna, the array and the processing needed to steer it.

munotes.in582

Directional Antennas, Sectorisation, Diversity and Spatial Reuse

2. What is sectorisation and what does it achieve? Sectorisation replaces one omnidirectional antenna at a site with several directed ones, typically three covering 120 degrees each or six covering 60 degrees, each with its own channels. Each antenna has gain, ideally 4.77 dB for 120 degrees and 7.78 dB for 60, which extends the reach or saves power, and the site can carry as many simultaneous channel sets as it has sectors, tripling or multiplying by six the capacity of the same site. The costs are interference between neighbouring sectors and handovers between them.

3. Explain switched and combining diversity, and say why diversity works at all. Diversity supplies two or more copies of the signal that are unlikely to fade at the same moment, for example from two antennas 10 to 20 cm apart, since multipath fading varies over a fraction of a wavelength. In switched or selection diversity the receiver estimates the quality of each antenna, as DECT does during a prolonged preamble, and then uses the better one for the rest of the packet; it needs only one receiver and a switch and DECT credits it with about 10 dB of link budget. In combining diversity the signals from all the antennas are added, ideally weighted by their strength, which performs better but needs a receiver chain per antenna.

4. What is a smart antenna? An array of antenna elements whose signals are combined with controlled phases and amplitudes, so that the radiation pattern is computed rather than fixed. It can steer its main lobe toward a wanted station and follow it as it moves, place a null on an interferer, and form several beams at once so that one cell can reuse a channel for users in different directions. It requires an array, continuous signal processing and calibration, so it belongs to base stations rather than to sensor nodes.

5. Define spatial reuse, and explain how directional antennas increase it. Spatial reuse is the use of the same frequency at the same time by two or more pairs of stations that are far enough apart, or aimed away from each other, for neither to spoil the other. Directional antennas increase it because a transmission confined to a narrow beam reaches fewer other nodes, and a directional receiver ignores transmissions outside its own beam, so two pairs that would have had to take turns can proceed together. In the chapter's simulation of 60 nodes, the number of transmissions running at once rose from 4.9 with omnidirectional antennas to 13.8 with 90-degree beams and 20.7 with 30-degree beams.

munotes.in583

Directional Antennas, Sectorisation, Diversity and Spatial Reuse

6. In the program, why did narrowing the beam from 360 to 180 degrees not reduce the number of nodes a transmission silenced? Because the model gives a narrower beam the gain it deserves, and gain spent on range works against the narrower angle. A 180-degree beam has twice the gain of an omnidirectional antenna and so reaches about 1.41 times as far; the area it covers is half the angle but twice the reach squared, which in a uniform field is the same area. So it silenced 20.6 nodes on average against 20.2. Spatial reuse still improved, because the nodes silenced differ from transmission to transmission, but the interference each transmission causes only falls once the beam is narrow enough, about 60 degrees and below in this field.

7. Why does the receiver's antenna matter as much as the sender's? Because a transmission is spoiled when an unwanted signal arrives at the receiver, and a directional receiver simply does not hear signals outside its beam. In the program, with 90-degree beams at both ends 13.8 transmissions ran at once; with the beam at the sender only, and the receiver listening in every direction, that fell to 5.8. Most of the gain from directional antennas therefore comes from the receiving end.

Contents This chapter on its own page

munotes.in584

Chapter Seventy-Eight

Signal Propagation: Ranges, Path Loss and How a Signal Travels

Syllabus topic Module 2, "Wireless Transmission: Signal propagation"

In one line

A transmission has no single range: it can be decoded out to one distance, can still ruin somebody else's reception further out, and is sensed as a busy channel out to a third; the power it delivers falls with the square of distance in free space and faster near the ground, and on the way it is blocked, reflected, refracted, scattered and diffracted, so what arrives is never what was sent.

In the wording a student can write in an examination: a transmission has three ranges. Within the transmission range a receiver can decode the frame (its received power is above the receiver's sensitivity). Within the wider interference range the signal is too weak to be decoded but strong enough to spoil another reception. Within the detection (carrier sense) range a receiver senses the channel as busy. Path loss in free space is L = 20 log10(4 pi d / lambda) or, in practical units, L = 32.4 + 20 log10 f(MHz) + 20 log10 d(km), so it grows by 6 dB for each doubling of distance. Near the ground a reflected path arrives as well, and beyond a crossover distance the two nearly cancel, so the loss grows as the fourth power of distance (12 dB per doubling); in general L(d) = L(d0) + 10 n log10(d / d0), with the path-loss exponent n between about 2 and 5.

A wave reaches a receiver in three ways: as a ground wave (following the earth's curvature, for long waves below about 2 MHz), as a sky wave (reflected by the ionosphere, for short waves, which is how HF crosses continents), and by line of sight (straight, for VHF and above, limited by the radio horizon, about 4/3 of the geometric horizon because the atmosphere bends the ray slightly). On the way, a wave suffers blocking or shadowing (an obstacle absorbs it), reflection (off surfaces large compared with the wavelength), refraction (bending as it passes through a medium of different density), scattering (off objects of the order of the wavelength, such as rain or foliage) and diffraction (bending around an edge, which is why a receiver behind a building is not in total silence).

Three ranges, not one

Drawing a circle around a node and calling it its range hides three different questions:

  • Transmission range. How far a frame can be decoded: the received power is at or above the receiver's sensitivity (the CC2420's is "-95 dBm" typical, against the standard's requirement of -85 dBm).
  • Interference range. How far the same transmission can ruin someone else's reception. A receiver needs the wanted signal to exceed the unwanted one by some margin (a signal-to-interference ratio), so a transmitter too far away to be decoded can still be close enough to destroy a weak wanted signal. This range is larger than the transmission range.
  • Detection or carrier-sense range. How far the channel is sensed busy, set by the CCA threshold ([The 802.15.4 Physical Layer]). On the CC2420 the "Carrier sense level" is "-77 dBm" typical, which is 18 dB above its sensitivity, so this range is shorter than the transmission range.
munotes.in585

Signal Propagation: Ranges, Path Loss and How a Signal Travels

That last fact is worth pausing on. The textbook picture has detection wider than transmission, so that carrier sense protects every receiver. On this radio it is the other way round: a node can be decodable but not sensed. Carrier sense therefore misses senders that a receiver can hear, and the MAC has to recover with acknowledgements and retries ([CSMA-CA, Data Transfer and Frames in 802.15.4]). It is the hidden terminal problem of [Hidden and Exposed Terminals, and RTS and CTS] arriving through the radio's own thresholds rather than through geometry.

The edges are not sharp either. Zuniga and Krishnamachari measured what is actually there: "three distinct reception regions in a wireless link: connected, transitional, and disconnected. The transitional region is often quite significant in size, and is generally characterized by high-variance in reception rates and asymmetric connectivity." And its size matters: "in dense deployments such as those envisioned for sensor networks, a large number of the links in the network (even higher than 50%) can be unreliable due to the transitional region." A range is a probability, not a boundary.

Path loss: free space and the ground

Free space is the reference: energy spreads over a sphere and nothing else happens. ITU-R P.525 gives the loss between isotropic antennas as 20 log10(4 pi d / lambda), or 32.4 + 20 log10 f + 20 log10 d with f in MHz and d in km. [Radio Technology in WSNs: The Sensor Radio and Its Link Budget] worked this into a link budget; the fact to carry forward is 6 dB per doubling of distance.

Near the ground a second copy arrives, reflected from the surface, having travelled slightly further. At short distances the two paths differ enough in length to arrive at random relative phases; beyond a crossover distance of about 4 pi ht hr / lambda (with ht and hr the antenna heights) the reflected path is almost exactly out of phase with the direct one, and the two largely cancel. Beyond that point the received power falls as the fourth power of distance: 12 dB per doubling. The program computes the crossover for a pair of motes.

In general the log-distance model is used: L(d) = L(d0) + 10 n log10(d / d0). The exponent n is 2 in free space, about 3 in a cluttered outdoor site, 4 near the ground, and 4 to 6 indoors through walls. A sensor network on the ground lives at the wrong end of that range, which is why its ranges are tens of metres and not hundreds.

munotes.in586

Signal Propagation: Ranges, Path Loss and How a Signal Travels

Three ways a wave travels

  • Ground wave. Below about 2 MHz a vertically polarised wave follows the curvature of the earth, held to the surface by the conducting ground. It reaches beyond the horizon, which is why long and medium wave broadcasting works at night and day over hundreds of kilometres.
  • Sky wave. Between about 2 and 30 MHz the ionosphere reflects waves back to earth, and several such bounces carry a short-wave signal across the world. The ionosphere changes with the sun, so the usable frequencies change by day and night.
  • Line of sight. Above about 30 MHz a wave travels essentially straight and passes through the ionosphere rather than being reflected, so transmitter and receiver must be able to "see" each other. The radio horizon is a little further than the geometric one, because the atmosphere's density falls with height and refracts the ray downward; the usual rule is to compute the geometric horizon with the earth's radius multiplied by 4/3.

Everything in this book, from 802.15.4 at 2.4 GHz to GSM at 900 MHz and satellites at gigahertz, is line of sight. That is why heights, obstacles and the horizon matter, and why a cell tower is on a mast.

What gets in the way

A wave that leaves an antenna rarely arrives by one clean path. Five mechanisms change it, each defined by the size of the obstacle compared with the wavelength:

  • Blocking (shadowing). An obstacle absorbs the wave and casts a shadow: a wall, a hill, a person. The loss depends on the material and the thickness, and it is what makes an indoor path-loss exponent so large. Because obstacles come and go, the loss varies slowly around its average, which is long-term (slow) fading ([Multipath, Fading and the Doppler Effect]).
  • Reflection. Off surfaces large compared with the wavelength: walls, the ground, vehicles, water. The reflected wave arrives later and with a changed phase, which is where multipath comes from. 3GPP's own description of a mobile channel starts here: the paths are "further randomized by local reflections or diffractions" near the mobile station.
  • Refraction. A wave bends when it passes into a medium where it travels at a different speed, as light does in water. In the atmosphere the density falls with height, so a ray bends gently downward, which is the 4/3 earth rule.
  • Scattering. Off objects of about the size of the wavelength, or rough surfaces: rain, foliage, street furniture. One wave becomes many weak ones in many directions. At 2.4 GHz, where the wavelength is 12 cm, leaves scatter well, which is why a sensor network in a field behaves worse in summer.
  • Diffraction. At the edge of an obstacle the wave bends into the geometric shadow, so a receiver just behind a building still hears something. Diffraction is stronger at lower frequencies, which is another reason 868 MHz reaches around corners better than 2.4 GHz.
munotes.in587

Signal Propagation: Ranges, Path Loss and How a Signal Travels

The practical consequence is that line of sight is a physical statement, not a geometric one: a link can work without a clear sight line, through diffraction and reflections, and can fail with one, through fading.

Propagation, computed

The program compares free-space loss with the two-ray ground model for two motes 1.5 m above the ground and finds the crossover; computes the three ranges for three path-loss exponents; tabulates the geometric and radio horizons; and prices a few obstacles in range.

# How a signal travels: free space against the ground-reflected path, the three
# ranges a transmission has, and the horizon that line of sight allows.
import math

C = 299792458.0

# 1. Free-space loss, ITU-R P.525 equation (6): L = 32.4 + 20 log10 f(MHz) +
#    20 log10 d(km). Near the ground a second path arrives, reflected, and
#    where the two nearly cancel the loss grows as the fourth power of the
#    distance instead of the second (the two-ray, or plane-earth, model).
free = lambda f_mhz, d_km: 32.4 + 20 * math.log10(f_mhz) + 20 * math.log10(d_km)
tworay = lambda d_m, ht, hr: 40 * math.log10(d_m) - 20 * math.log10(ht * hr)

print("Free space against two paths, 2450 MHz, antennas 1.5 m and 1.5 m up:")
print("   distance   free space   two-ray (ground reflection)   difference")
for d in (1, 5, 10, 30, 100, 300, 1000):
    fs, tr = free(2450, d / 1000), tworay(d, 1.5, 1.5)
    print("  %7d m %10.1f dB %26.1f dB %11.1f dB" % (d, fs, tr, tr - fs))
cross = (4 * math.pi * 1.5 * 1.5) / (C / 2450e6)
print("  the two agree near d = 4 pi ht hr / lambda = %.0f m; beyond it the ground path dominates" % cross)

# 2. The three ranges. A transmission can be decoded out to the transmission
#    range; it can still spoil a reception further out (the interference
#    range, where it arrives within SIR dB of a wanted signal at sensitivity);
#    and carrier sense declares the channel busy only out to where the signal
#    passes the CCA threshold. The CC2420's own figures are used: 0 dBm out,
#    sensitivity -95 dBm typical, carrier sense level -77 dBm typical, and a
#    signal-to-interference ratio of 6 dB taken as the limit of reception.
P_TX, SENS, CCA, SIR = 0.0, -95.0, -77.0, 6.0        # dBm, dBm, dBm, dB
def range_for(limit_dbm, f_mhz=2450, n=3.0):
    """Distance at which the received power falls to limit_dbm, with a
    log-distance model of exponent n anchored on free space at 1 m."""
    return 10 ** ((P_TX - limit_dbm - free(f_mhz, 0.001)) / (10 * n))

print("\nThe three ranges (0 dBm, free space at 1 m, then an exponent n):")
print("      n   transmission   interference   carrier sense   interference/transmission")
for n in (2.0, 3.0, 4.0):
    tx, intf, cs = range_for(SENS, n=n), range_for(SENS - SIR, n=n), range_for(CCA, n=n)
    print("  %5.1f %12.1f m %13.1f m %14.1f m %20.2f" % (n, tx, intf, cs, intf / tx))
print("  a receiver at the edge of a link is spoiled by any transmitter inside its")
print("  interference range, which is wider than the range at which it can be heard.")

# 3. Line of sight and the radio horizon. The geometric horizon from a height h
#    is sqrt(2 R h); radio waves bend a little, so the usual rule multiplies the
#    earth's radius by 4/3.
R = 6371000.0
print("\nHow far is the horizon, and the radio horizon (earth radius x 4/3):")
print("   height   geometric   radio")
for h in (1.5, 10, 30, 100, 300):
    print("  %6.1f m %9.1f km %7.1f km"
          % (h, math.sqrt(2 * R * h) / 1000, math.sqrt(2 * (4 / 3) * R * h) / 1000))

# 4. What obstacles cost. Typical one-off losses, added to the path loss.
print("\nWhat one obstacle costs (illustrative figures), and the range it leaves at n = 3:")
budget = P_TX - SENS
for name, loss in (("nothing", 0), ("a wooden door", 4), ("a plasterboard wall", 6),
                   ("a brick wall", 12), ("a concrete floor", 20), ("a lift shaft", 30)):
    l1 = free(2450, 0.001)
    d = 10 ** ((budget - loss - l1) / 30)
    print("  %-20s %3d dB   range %6.1f m  (%.0f%% of the open figure)"
          % (name, loss, d, 100 * d / range_for(SENS, n=3.0)))
munotes.in588

Signal Propagation: Ranges, Path Loss and How a Signal Travels

Free space against two paths, 2450 MHz, antennas 1.5 m and 1.5 m up:
   distance   free space   two-ray (ground reflection)   difference
        1 m       40.2 dB                       -7.0 dB       -47.2 dB
        5 m       54.2 dB                       20.9 dB       -33.2 dB
       10 m       60.2 dB                       33.0 dB       -27.2 dB
       30 m       69.7 dB                       52.0 dB       -17.7 dB
      100 m       80.2 dB                       73.0 dB        -7.2 dB
      300 m       89.7 dB                       92.0 dB         2.3 dB
     1000 m      100.2 dB                      113.0 dB        12.8 dB
  the two agree near d = 4 pi ht hr / lambda = 231 m; beyond it the ground path dominates

The three ranges (0 dBm, free space at 1 m, then an exponent n):
      n   transmission   interference   carrier sense   interference/transmission
    2.0        550.6 m        1098.6 m           69.3 m                 2.00
    3.0         67.2 m         106.5 m           16.9 m                 1.58
    4.0         23.5 m          33.1 m            8.3 m                 1.41
  a receiver at the edge of a link is spoiled by any transmitter inside its
  interference range, which is wider than the range at which it can be heard.

How far is the horizon, and the radio horizon (earth radius x 4/3):
   height   geometric   radio
     1.5 m       4.4 km     5.0 km
    10.0 m      11.3 km    13.0 km
    30.0 m      19.6 km    22.6 km
   100.0 m      35.7 km    41.2 km
   300.0 m      61.8 km    71.4 km

What one obstacle costs (illustrative figures), and the range it leaves at n = 3:
  nothing                0 dB   range   67.2 m  (100% of the open figure)
  a wooden door          4 dB   range   49.4 m  (74% of the open figure)
  a plasterboard wall    6 dB   range   42.4 m  (63% of the open figure)
  a brick wall          12 dB   range   26.7 m  (40% of the open figure)
  a concrete floor      20 dB   range   14.5 m  (22% of the open figure)
  a lift shaft          30 dB   range    6.7 m  (10% of the open figure)
munotes.in589

Signal Propagation: Ranges, Path Loss and How a Signal Travels

Free space against the ground. At 1 m the two-ray formula is meaningless (it is derived for distances beyond the crossover, and gives a negative loss), and out to about 100 m it predicts less loss than free space, because the reflected ray still adds. The crossover for these antennas is 231 m; beyond it the ground path turns against the direct one, and at 1 km the two-ray loss is 12.8 dB worse than free space. Every doubling then costs 12 dB instead of 6, so range grows only as the fourth root of power: to double the range of a ground-level link one must raise the power sixteen times.

The three ranges. At n = 3 the transmission range is 67.2 m, the interference range 106.5 m, and the carrier-sense range only 16.9 m. The interference range is 1.58 times the transmission range, and the ratio falls toward 1.41 as n rises, because a steeper slope compresses every distance. But the carrier-sense range is a quarter of the transmission range, purely because the CC2420's CCA threshold sits 18 dB above its sensitivity. A node listening before transmitting therefore cannot hear most of the senders whose frames it could have decoded.

The horizon. From 1.5 m, mote height, the radio horizon is 5.0 km; from a 30 m mast, 22.6 km; from 300 m, 71.4 km. A link between two stations sees the sum of their horizons, so two 30 m masts can reach about 45 km if nothing intervenes. That is the shape of cellular planning ([Cellular Systems: Cells, Clusters and Frequency Reuse]).

munotes.in590

Signal Propagation: Ranges, Path Loss and How a Signal Travels

Obstacles. With a 95 dB budget and n = 3, an open range of 67.2 m falls to 42.4 m through a plasterboard wall (6 dB), 26.7 m through a brick wall (12 dB), 14.5 m through a concrete floor (20 dB) and 6.7 m into a lift shaft (30 dB). Each 6 dB halves the range at n = 2 but only cuts it to 63 per cent at n = 3: a steeper exponent makes obstacles hurt less in distance terms, because the signal was falling fast anyway.

Distinctions

Transmission rangeInterference rangeDetection (carrier sense) range
Defined byReceiver sensitivityThe signal-to-interference ratio a reception needsThe CCA threshold
What happens thereThe frame can be decodedAnother reception is spoiledThe channel is reported busy
Program (n = 3)67.2 m106.5 m16.9 m
Usually assumedSmallestLargestIn between
On a CC2420MiddleLargestSmallest, because CCA sits 18 dB above sensitivity
Ground waveSky waveLine of sight
FrequenciesBelow about 2 MHzAbout 2 to 30 MHzAbove about 30 MHz
How it travelsFollows the earth's surfaceReflected by the ionosphereStraight, to the radio horizon
ReachHundreds of kmIntercontinental, and variableThe horizon, extended by height
Used byLong and medium wave broadcastingShort wave, amateur radioEverything in this book
MechanismObstacle compared with the wavelengthEffect
Blocking (shadowing)Large and absorbingA shadow; slow variation as obstacles change
ReflectionLarge and smoothA delayed copy; the source of multipath
RefractionA medium of changing densityThe path bends; the 4/3 earth rule
ScatteringComparable to the wavelength, or roughMany weak copies in many directions
DiffractionAn edgeThe wave bends into the shadow

What it does not mean

A range is not a circle. It is a probability that falls with distance, with a wide transitional region in which links are unreliable and often asymmetric.

Carrier sense does not protect every receiver. With a CCA threshold above sensitivity, a node can fail to sense a sender whose frames it could decode.

Free space is not the usual case. Near the ground the exponent is 3 to 4, and indoors 4 to 6; free space is the best case and the reference.

Line of sight is not required for a link. Diffraction and reflections carry signals into shadows; conversely a clear line does not guarantee a link, because fading can cancel it.

A higher path-loss exponent is not only bad news. It shortens every range, including the interference range, which is why dense networks in cluttered places can reuse space well.

munotes.in591

Signal Propagation: Ranges, Path Loss and How a Signal Travels

Quick revision

  • Three ranges: transmission (above sensitivity), interference (spoils another reception; wider), detection (CCA threshold). Program at n = 3: 67.2 m, 106.5 m, 16.9 m. CC2420: sensitivity -95 dBm, carrier sense -77 dBm, so carrier sense is the shortest.
  • Zuniga and Krishnamachari: connected, transitional, disconnected regions; in dense deployments over 50 per cent of links can be unreliable.
  • Free space: L = 32.4 + 20 log10 f(MHz) + 20 log10 d(km), 6 dB per doubling. Two-ray: beyond 4 pi ht hr / lambda (231 m for two motes at 1.5 m, 2450 MHz) the loss goes as d to the fourth, 12 dB per doubling. General: L(d) = L(d0) + 10 n log10(d/d0), n = 2 to 6.
  • Travel: ground wave (below 2 MHz, follows the earth), sky wave (2 to 30 MHz, ionosphere), line of sight (above 30 MHz; radio horizon with earth radius times 4/3: 5.0 km from 1.5 m, 22.6 km from 30 m).
  • In the way: blocking/shadowing, reflection (large smooth surfaces), refraction (changing density), scattering (objects near the wavelength: rain, leaves), diffraction (edges; stronger at lower frequencies).
  • Obstacles at n = 3: plasterboard 6 dB leaves 63 per cent of the range; brick 12 dB, 40 per cent; concrete floor 20 dB, 22 per cent.

Test yourself

1. Distinguish the transmission, interference and detection ranges. Within the transmission range the received power exceeds the receiver's sensitivity, so a frame can be decoded. Within the interference range the signal is too weak to be decoded but still strong enough, relative to a wanted signal, to spoil another node's reception; since a receiver needs the wanted signal to exceed the interferer by a margin, this range is wider than the transmission range. Within the detection or carrier-sense range the signal exceeds the clear channel assessment threshold, so a node senses the channel as busy; on a CC2420, whose carrier sense level is -77 dBm against a sensitivity of -95 dBm, this range is the shortest of the three.

2. Give the free-space path-loss formula and explain what it implies. Between isotropic antennas the free-space basic transmission loss is 20 log10(4 pi d / lambda), or in practical units 32.4 + 20 log10 f + 20 log10 d with the frequency in megahertz and the distance in kilometres. It implies that the loss rises by 6 dB for every doubling of distance, and by 6 dB for every doubling of frequency, so a 2.4 GHz link loses about 9 dB more than an 868 MHz link over the same distance.

munotes.in592

Signal Propagation: Ranges, Path Loss and How a Signal Travels

3. Why does a link near the ground lose power faster than free space predicts? Because a second copy of the signal arrives after reflecting from the ground. Beyond a crossover distance of roughly 4 pi ht hr / lambda, where ht and hr are the antenna heights, the reflected ray arrives almost exactly out of phase with the direct ray and largely cancels it, so the received power falls with the fourth power of distance rather than the second: 12 dB per doubling instead of 6. For two motes 1.5 m above the ground at 2450 MHz, the crossover is about 231 m.

4. Describe the three ways a radio wave can travel from transmitter to receiver. As a ground wave, below about 2 MHz, following the curvature of the earth and reaching far beyond the horizon. As a sky wave, roughly 2 to 30 MHz, reflected back to earth by the ionosphere, possibly several times, giving intercontinental reach that varies with the time of day. And by line of sight, above about 30 MHz, travelling essentially straight, so the two stations must be within the radio horizon, which is about 4/3 of the geometric horizon because the atmosphere refracts the ray downward.

5. Explain blocking, reflection, refraction, scattering and diffraction. Blocking or shadowing is absorption by an obstacle large compared with the wavelength, casting a radio shadow. Reflection is the bouncing of the wave from a large smooth surface, producing delayed copies and hence multipath. Refraction is the bending of the wave as it passes through a medium whose density changes, as in the atmosphere. Scattering is the dispersal of the wave by objects of about the size of the wavelength or by rough surfaces, such as rain or foliage, into many weak waves. Diffraction is the bending of the wave around an edge into the geometric shadow, which is why reception behind an obstacle is not zero and why lower frequencies reach around corners better.

6. A CC2420 has a sensitivity of -95 dBm and a carrier sense level of -77 dBm. What follows for a CSMA MAC? The carrier-sense range is much shorter than the transmission range, in the chapter's model 16.9 m against 67.2 m at a path-loss exponent of 3. A node performing clear channel assessment therefore fails to detect transmitters whose frames it could have received, so it may transmit into an ongoing reception. Carrier sense alone cannot prevent these collisions; the MAC must rely on acknowledgements and retransmissions, and the effect adds to the hidden terminal problem that geometry already causes.

Contents This chapter on its own page

munotes.in593

Chapter Seventy-Nine

Multipath, Fading and the Doppler Effect

Syllabus topic Module 2, "Wireless Transmission: Signal propagation"

In one line

A radio signal reaches the receiver by several paths of different lengths, so copies arrive at different times and phases: added together they can reinforce or cancel, which is fading, and they smear each symbol into the next, which is intersymbol interference; when either end moves, every path's frequency shifts a little, which is the Doppler effect, and the channel that was good a moment ago is gone.

In the wording a student can write in an examination: multipath propagation is the arrival of the same transmitted signal by several paths, reflected, scattered and diffracted, with different delays and amplitudes. The spread of those delays is the delay spread; the last significant echo sets how far a symbol is smeared, causing intersymbol interference (ISI), so the symbol rate a channel can carry without an equaliser is limited by it. Copies that arrive in phase add and copies out of phase cancel, so the received amplitude varies: this is fading. Short-term (fast) fading varies over a fraction of a wavelength and is described by a Rayleigh distribution when no path dominates, or Rician when one does (a line of sight); long-term (slow) fading, or shadowing, is the slower variation as obstacles come and go. If the delay spread is small compared with the symbol time, all frequencies fade together (flat fading); if not, some frequencies fade and others do not (frequency-selective fading), and the width over which the channel is roughly constant is the coherence bandwidth. Movement adds the Doppler shift, which 3GPP writes as fd = v divided by the wavelength, with v "representing the vehicle speed"; its reciprocal is roughly the coherence time, how long the channel stays the same.

The answers are a fade margin, diversity ([Directional Antennas, Sectorisation, Diversity and Spatial Reuse]), equalisation, spread spectrum ([Spread Spectrum and Direct Sequence]) and multi-carrier modulation ([Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM]).

Multipath: the same signal, several times

3GPP describes the channel a mobile actually sees: "Radio propagation in the mobile radio environment is described by highly dispersive multipath caused by reflection and scattering. The paths between base station and MS may be considered to consist of large reflectors and/or scatterers some distance to the MS, giving rise to a number of waves that arrive in the vicinity of the MS with random amplitudes and delays."

And near the mobile it gets worse: "Close to the MS these paths are further randomized by local reflections or diffractions. Since the MS will be moving, the angle of arrival must also be taken into account, since it affects the doppler shift associated with a wave arriving from a particular direction."

munotes.in594

Multipath, Fading and the Doppler Effect

To make that simulable, the standard reduces the channel to a list of taps: "1) a discrete number of taps, each determined by their time delay and their average power; 2) the Rayleigh distributed amplitude of each tap, varying according to a doppler spectrum S(f)." Its Annex C gives three such lists, for a rural area, an urban area and hilly terrain, and the program uses them as they stand.

Delay spread and intersymbol interference

The delay spread is how far apart in time the echoes arrive, measured as the power-weighted standard deviation of the tap delays. It matters because a symbol smeared by more than its own duration runs into its neighbour: intersymbol interference.

The rule of thumb is that without an equaliser a channel can carry about 1 / (10 times the delay spread) symbols a second. Above that the receiver must equalise: estimate the channel from a known training sequence and undo the smearing. GSM's normal burst carries a 26-bit training sequence in its middle for exactly this reason ([The GSM Radio Interface: Carriers, the TDMA Frame and Bursts]).

Three ways out, and all are used:

  • Equalisation, as GSM does.
  • Spread spectrum, where a receiver can separate the echoes by their delays and even add them together, which is what a rake receiver in UMTS does ([The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover]).
  • Slower symbols on many carriers, which is OFDM: each carrier's symbol is long compared with the delay spread, so the smearing is small.

Fading

Short-term (fast) fading is what happens when the copies add. Move a receiver a few centimetres and the relative phases change, so the sum swings between constructive and destructive. It varies over a fraction of a wavelength, which at 2.4 GHz is a matter of centimetres.

When many paths arrive and none dominates, the amplitude follows a Rayleigh distribution, which is what 3GPP's taps use: "the Rayleigh distributed amplitude of each tap". When one path dominates, usually a line of sight, the distribution is Rician, and the fades are shallower. 3GPP's rural profile marks its first tap "RICE" and the rest "CLASS", because in the country there is usually a direct path.

Long-term (slow) fading, or shadowing, is the slower variation of the average level as the path passes behind buildings and trees. It is usually modelled as log-normal: the loss in decibels varies normally about the path-loss prediction.

Flat against frequency-selective. If the delay spread is much smaller than a symbol, every frequency in the signal fades together: flat fading. If not, the channel's response varies across the signal's band, so some frequencies are deeply faded and others are not: frequency-selective fading. The band over which the channel is roughly constant is the coherence bandwidth, about the reciprocal of the delay spread. A narrowband signal in a selective channel can be wiped out entirely; a wideband one loses only part of its spectrum, which is another argument for spreading.

munotes.in595

Multipath, Fading and the Doppler Effect

The fade margin. Because fades are random, a link is designed with extra power in hand. DECT's own link budget names the number: "Normally switched antenna diversity is implemented, whereby MF = 11 dB in a quasi-stationary environment." The program measures what margin a Rayleigh channel actually demands.

The Doppler effect

If the transmitter, the receiver or a reflector moves, each path's frequency shifts. 3GPP states the maximum, fd = v / lambda: the speed "in ms-1" divided by "the wavelength". A wave arriving from straight ahead is shifted up by fd, one from behind down by fd, and one from the side not at all, which is why a moving receiver sees a whole Doppler spectrum rather than a single shift; 3GPP defines two shapes, CLASS (the classical spectrum) and RICE.

Two consequences:

  • The carrier must be tracked. A receiver's frequency reference has to follow a shift that changes as the geometry changes.
  • The channel ages. The rate at which the fading pattern changes is set by fd, and the coherence time, roughly 1 / fd, is how long the channel stays much the same. A frame longer than the coherence time fades within itself, so interleaving and coding are needed to spread the damage.

Multipath, fading and Doppler, computed

The program computes the delay spread of 3GPP's three profiles and the symbol rate each allows without an equaliser; counts how many symbols an echo spills into for three systems; simulates Rayleigh fading to find the depth of fades and the margin an availability requires; and tabulates the Doppler shift and coherence time at five speeds.

# Multipath, fading and Doppler, from 3GPP TS 45.005's own channel profiles.
import math
import random

C = 299792458.0

# 1. Delay spread. Each profile is a list of taps: (relative delay in
#    microseconds, average relative power in dB). The delay spread is the
#    power-weighted standard deviation of the delays, and a rule of thumb makes
#    the symbol rate a channel can carry without an equaliser about 1/(10 s).
PROFILES = {
    "rural (RAx), 6 taps": [(0.0, 0.0), (0.1, -4.0), (0.2, -8.0), (0.3, -12.0),
                            (0.4, -16.0), (0.5, -20.0)],
    "urban (TUx), 12 taps": [(0.0, -4.0), (0.1, -3.0), (0.3, 0.0), (0.5, -2.6), (0.8, -3.0),
                             (1.1, -5.0), (1.3, -7.0), (1.7, -5.0), (2.3, -6.5), (3.1, -8.6),
                             (3.2, -11.0), (5.0, -10.0)],
    "hilly terrain (HTx), 12 taps": [(0.0, -10.0), (0.1, -8.0), (0.3, -6.0), (0.5, -4.0),
                                     (0.7, 0.0), (1.0, 0.0), (1.3, -4.0), (15.0, -8.0),
                                     (15.2, -9.0), (15.7, -10.0), (17.2, -12.0), (20.0, -14.0)],
}

def spread(taps):
    p = [10 ** (db / 10) for _, db in taps]
    t = [us for us, _ in taps]
    total = sum(p)
    mean = sum(pi * ti for pi, ti in zip(p, t)) / total
    var = sum(pi * (ti - mean) ** 2 for pi, ti in zip(p, t)) / total
    return mean, math.sqrt(var), max(t)

print("Delay spread of the propagation profiles of 3GPP TS 45.005, Annex C:")
print("  profile                        mean delay   delay spread   last echo   symbols/s without an equaliser")
for name, taps in PROFILES.items():
    mean, rms, last = spread(taps)
    print("  %-30s %7.2f us %11.2f us %9.1f us %19.0f" % (name, mean, rms, last, 1 / (10 * rms * 1e-6)))
print("  GSM sends 270,833 symbols/s, so urban and hilly echoes cross symbol boundaries: an equaliser is needed.")

# 2. Intersymbol interference: how many symbols an echo spills into.
print("\nHow far an echo reaches, in symbols, for three systems:")
for system, rate in (("GSM, 270.833 ksymbol/s", 270833), ("802.15.4, 62.5 ksymbol/s", 62500),
                     ("a 20 Msymbol/s link", 20e6)):
    print("  %-26s" % system + "".join(
        " %s: %5.2f symbols;" % (name.split(" (")[0], spread(taps)[1] * 1e-6 * rate)
        for name, taps in PROFILES.items()))

# 3. Fading. Rayleigh amplitudes, as 3GPP's taps use: the sum of many random
#    paths. How often is the signal far below its average, and what margin does
#    a given outage need?
rnd = random.Random(79)
# a Rayleigh amplitude is the length of a vector whose two components are
# independent and normal, so the power is the sum of their squares
samples = sorted((rnd.gauss(0, 1) ** 2 + rnd.gauss(0, 1) ** 2) / 2 for _ in range(200000))
mean = sum(samples) / len(samples)
print("\nRayleigh fading: how deep the fades are (200,000 samples, mean power %.3f):" % mean)
print("   fraction of the time   the power is below   that is a fade of")
for q in (0.5, 0.1, 0.05, 0.01, 0.001):
    v = samples[int(q * len(samples))]
    print("  %20.1f%% %19.4f %17.1f dB" % (100 * q, v / mean, 10 * math.log10(v / mean)))
print("  so a link that must work 99%% of the time needs about %.0f dB of fade margin"
      % -(10 * math.log10(samples[int(0.01 * len(samples))] / mean)))

# 4. Doppler. 3GPP: fd = v / lambda, with v the speed and lambda the wavelength.
#    The coherence time, how long the channel stays much the same, is about
#    1/fd; a frame longer than that fades within itself.
print("\nDoppler shift fd = v / lambda, and roughly how long the channel holds still:")
print("   speed          900 MHz            2450 MHz          coherence time at 2450 MHz")
for kmh in (1.4, 5, 50, 120, 300):
    v = kmh / 3.6
    f900, f2450 = v / (C / 900e6), v / (C / 2450e6)
    print("  %5.0f km/h %10.1f Hz %17.1f Hz %22.1f ms" % (kmh, f900, f2450, 1000 / f2450))
print("  a 802.15.4 frame of 133 octets lasts 4.256 ms, so below about %.0f km/h it sees one channel"
      % (3.6 * (1 / 0.004256) * (C / 2450e6)))
munotes.in596

Multipath, Fading and the Doppler Effect

Delay spread of the propagation profiles of 3GPP TS 45.005, Annex C:
  profile                        mean delay   delay spread   last echo   symbols/s without an equaliser
  rural (RAx), 6 taps               0.06 us        0.10 us       0.5 us             1023588
  urban (TUx), 12 taps              0.89 us        1.03 us       5.0 us               97466
  hilly terrain (HTx), 12 taps      2.70 us        5.10 us      20.0 us               19616
  GSM sends 270,833 symbols/s, so urban and hilly echoes cross symbol boundaries: an equaliser is needed.

How far an echo reaches, in symbols, for three systems:
  GSM, 270.833 ksymbol/s     rural:  0.03 symbols; urban:  0.28 symbols; hilly terrain:  1.38 symbols;
  802.15.4, 62.5 ksymbol/s   rural:  0.01 symbols; urban:  0.06 symbols; hilly terrain:  0.32 symbols;
  a 20 Msymbol/s link        rural:  1.95 symbols; urban: 20.52 symbols; hilly terrain: 101.96 symbols;

Rayleigh fading: how deep the fades are (200,000 samples, mean power 1.004):
   fraction of the time   the power is below   that is a fade of
                  50.0%              0.6965              -1.6 dB
                  10.0%              0.1044              -9.8 dB
                   5.0%              0.0510             -12.9 dB
                   1.0%              0.0095             -20.2 dB
                   0.1%              0.0009             -30.4 dB
  so a link that must work 99% of the time needs about 20 dB of fade margin

Doppler shift fd = v / lambda, and roughly how long the channel holds still:
   speed          900 MHz            2450 MHz          coherence time at 2450 MHz
      1 km/h        1.2 Hz               3.2 Hz                  314.7 ms
      5 km/h        4.2 Hz              11.4 Hz                   88.1 ms
     50 km/h       41.7 Hz             113.5 Hz                    8.8 ms
    120 km/h      100.1 Hz             272.4 Hz                    3.7 ms
    300 km/h      250.2 Hz             681.0 Hz                    1.5 ms
  a 802.15.4 frame of 133 octets lasts 4.256 ms, so below about 104 km/h it sees one channel
munotes.in597

Multipath, Fading and the Doppler Effect

Delay spread. The rural profile has a spread of 0.10 microseconds, the urban 1.03, and the hilly terrain 5.10, with echoes out to 20 microseconds, which is 6 km of extra path: a signal bouncing off a distant hillside. Without an equaliser those allow about 1,023,588, 97,466 and 19,616 symbols a second respectively. GSM sends 270,833 symbols a second, comfortably inside the rural figure and far outside the other two, which is why every GSM burst carries a training sequence and every GSM receiver has an equaliser.

How many symbols an echo crosses. At GSM's rate an urban echo spills over 0.28 of a symbol and a hilly one 1.38. At 802.15.4's 62.5 ksymbol/s the same channels give 0.06 and 0.32: the slow symbols of a sensor radio are largely immune, which is one reason the standard needs no equaliser. At 20 Msymbol/s the urban channel smears over 20 symbols and the hilly one over 102, which is why fast systems must use OFDM or a rake.

munotes.in598

Multipath, Fading and the Doppler Effect

How deep fades go. In a Rayleigh channel the power is below its mean 50 per cent of the time (a shallow 1.6 dB), below a tenth of the mean 10 per cent of the time (9.8 dB down), and 20.2 dB down 1 per cent of the time. So a link that must work 99 per cent of the time needs about 20 dB of fade margin, and 99.9 per cent needs 30 dB. That is the true cost of fading, and it dwarfs the few decibels an antenna or an amplifier buys; it is also why diversity, which DECT credits with about 10 dB, is worth so much.

Doppler. At walking pace, 5 km/h, the shift at 2450 MHz is 11.4 Hz and the channel holds still for about 88 ms. In a car at 120 km/h it is 272 Hz and 3.7 ms. An 802.15.4 frame of 133 octets lasts 4.256 ms, so below about 104 km/h a whole frame sees one channel: a sensor network, whose nodes usually do not move at all, lives in a channel that changes only when the world around it does. A GSM handset in a train does not: at 300 km/h the channel is new every 1.5 ms, which is shorter than a GSM frame, and the interleaving of [GSM Logical Channels and the Frame Hierarchy] exists to spread a codeword across those changes.

Distinctions

Short-term (fast) fadingLong-term (slow) fading
Caused byMultipath copies adding and cancellingObstacles shadowing the path
Varies overA fraction of a wavelength, centimetresMetres to tens of metres
DistributionRayleigh (no dominant path), Rician (one dominant)Log-normal
Answered byDiversity, spreading, coding and interleavingFade margin, power control, better siting
Flat fadingFrequency-selective fading
WhenDelay spread much less than the symbol timeDelay spread comparable to or more than it
EquivalentlySignal bandwidth less than the coherence bandwidthMore than it
What happensThe whole signal fades togetherParts of the spectrum fade, others do not
Answered byDiversity in time or spaceEqualisation, spreading, OFDM
Delay spreadDoppler shift
Comes fromPaths of different lengthMotion
Measured inMicrosecondsHertz
LimitsThe symbol rate (ISI)How long the channel lasts (coherence time)
Program0.10 / 1.03 / 5.10 us (rural / urban / hilly)3.2 Hz at 1.4 km/h, 681 Hz at 300 km/h, at 2450 MHz
munotes.in599

Multipath, Fading and the Doppler Effect

What it does not mean

Multipath is not only harmful. Copies can be combined: a rake receiver and MIMO turn several paths into a gain.

Rayleigh is not a worst case. It is the ordinary case when no path dominates; a Rician channel with a line of sight fades less.

A fade is not a loss of power at the transmitter. The power is being sent; it is cancelling itself at that point in space at that instant.

Doppler is not only a frequency error. The shift itself is usually easy to track; the damage is that the channel changes at that rate.

A high delay spread does not always need an equaliser. It depends on the symbol rate: the same hilly channel needs one at GSM's rate and not at 802.15.4's.

Quick revision

  • Multipath: copies by reflection, scattering, diffraction, with "random amplitudes and delays"; modelled as taps (delay, average power, Rayleigh amplitude, Doppler spectrum).
  • Delay spread: power-weighted spread of the delays. 3GPP: rural 0.10 us, urban 1.03 us, hilly 5.10 us (echoes to 20 us). Symbols a second without an equaliser, about 1/(10 times the spread): 1.02 M, 97 k, 20 k. GSM at 270.833 ksymbol/s needs an equaliser and a training sequence.
  • Fading: short-term (fast), over centimetres, Rayleigh (no dominant path) or Rician (line of sight); long-term (slow), shadowing, log-normal.
  • Flat against frequency-selective, decided by the delay spread against the symbol time; coherence bandwidth is about 1 / delay spread.
  • Fade margin: Rayleigh is 9.8 dB down 10 per cent of the time, 20.2 dB down 1 per cent, 30.4 dB down 0.1 per cent. DECT assumes 11 dB with switched diversity.
  • Doppler: fd = v / lambda; at 2450 MHz, 11.4 Hz at 5 km/h and 272 Hz at 120 km/h; coherence time about 1 / fd: 88 ms and 3.7 ms. An 802.15.4 frame (4.256 ms) sees one channel below about 104 km/h.
  • Answers: fade margin, diversity, equalisation, spread spectrum, OFDM, coding and interleaving.

Test yourself

1. What is multipath propagation, and how is a multipath channel modelled? It is the arrival of the same transmitted signal by several paths of different lengths, produced by reflection, scattering and diffraction, so that copies arrive with different delays, amplitudes and phases. For simulation the channel is reduced to a discrete number of taps, each with a time delay and an average power, with the amplitude of each tap varying randomly, usually with a Rayleigh distribution, according to a Doppler spectrum. 3GPP TS 45.005 gives such tap lists for rural, urban and hilly terrain.

munotes.in600

Multipath, Fading and the Doppler Effect

2. Define delay spread and explain how it limits the symbol rate. The delay spread is the spread of the arrival times of the multipath copies, computed as the power-weighted standard deviation of the tap delays. Echoes arriving later than one symbol period spill into the following symbol, causing intersymbol interference, so as a rule of thumb a channel can carry about one tenth of the reciprocal of the delay spread in symbols per second before an equaliser is needed. For 3GPP's urban profile, with a spread of 1.03 microseconds, that is about 97,000 symbols a second, well below GSM's 270,833, which is why GSM receivers equalise.

3. Distinguish short-term and long-term fading. Short-term or fast fading is the rapid variation of the received amplitude as multipath copies add and cancel; it changes over a fraction of a wavelength, so over centimetres at 2.4 GHz, and follows a Rayleigh distribution when no path dominates or a Rician one when there is a line of sight. Long-term or slow fading, also called shadowing, is the slower variation of the average received level as the path passes behind buildings, hills and trees; it is usually modelled as log-normal, that is normal in decibels.

4. What is the difference between flat and frequency-selective fading? If the delay spread is much smaller than the symbol duration, equivalently if the signal's bandwidth is smaller than the coherence bandwidth, then all the frequency components fade together and the fading is flat. If the delay spread is comparable to or larger than the symbol duration, the channel's response varies across the signal's band, so some frequencies are deeply attenuated while others are not, which is frequency-selective fading; it is combated by equalisation, spread spectrum or multi-carrier modulation.

5. How much fade margin does a Rayleigh channel require, and why? Because the received power varies at random, a link designed for the average fails whenever the signal fades below the receiver's sensitivity. In the chapter's simulation the power was 9.8 dB below the mean 10 per cent of the time, 20.2 dB below it 1 per cent of the time and 30.4 dB below it 0.1 per cent of the time, so a link that must work 99 per cent of the time needs about 20 dB of margin and one that must work 99.9 per cent about 30 dB. Diversity reduces the margin needed, which is why DECT assumes 11 dB with switched antenna diversity.

6. State the Doppler shift and explain coherence time. The maximum Doppler shift is fd = v / lambda, the speed of the mobile divided by the wavelength; paths arriving from ahead are shifted up by fd and those from behind down by fd, giving a spectrum rather than a single shift. The coherence time, roughly the reciprocal of fd, is how long the channel remains much the same. At 2450 MHz and 120 km/h, fd is 272 Hz and the coherence time about 3.7 ms, so a frame longer than that fades within itself and needs coding and interleaving; an 802.15.4 frame of 4.256 ms is inside the coherence time below about 104 km/h.

Contents This chapter on its own page

munotes.in601

Chapter Eighty

Multiplexing: Space, Frequency, Time and Code

Syllabus topic Module 2, "Wireless Transmission: Multiplexing"

In one line

Two transmissions can share one medium only if something keeps them apart, and there are exactly four things available: they can be in different places, on different frequencies, at different times, or spread by different codes, and every one of the four costs a guard space, whether a distance, a gap in the spectrum, an idle moment or a lower rate.

In the wording a student can write in an examination: multiplexing is the sharing of one transmission medium by several signals. There are four dimensions:

  1. Space division multiplexing (SDM): signals are separated by being in different places, so the same frequency may be reused at a distance. The guard space is a physical distance (the reuse distance of a cellular system) or a direction (a sectorised or smart antenna). Used by every cellular network ([Cellular Systems: Cells, Clusters and Frequency Reuse]).
  2. Frequency division multiplexing (FDM): the band is divided into carriers, each used continuously by one signal. The guard space is a guard band between carriers. Used by radio and television broadcasting, by GSM's 200 kHz carriers, and by 802.15.4's sixteen channels.
  3. Time division multiplexing (TDM): all signals use the whole band, but each in its own slot of a repeating frame. The guard space is a guard time for propagation delay and clock error. Used by GSM's eight slots per carrier and by the TDMA of [TDMA and Schedule-based MAC].
  4. Code division multiplexing (CDM): all signals use the whole band at the same time, each spread by a different code; a receiver recovers one by correlating with its code. The guard space is the code distance, paid for by bandwidth, since the rate falls by the spreading factor. Used by W-CDMA ([The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover]) and, for robustness rather than sharing, by 802.15.4.

Real systems combine them: GSM separates users by cell (space), carrier (frequency) and slot (time). When these schemes are used to give many users access to a shared medium, they are called FDMA, TDMA, CDMA and SDMA.

Space

The oldest and most valuable dimension. A frequency used in Mumbai can be used again in Pune, because the signal has faded to nothing in between. Its guard space is distance: the reuse distance, set by how much interference a receiver can tolerate.

Space division is what makes cellular networks possible, and the capacity of a city has almost nothing to do with the bandwidth allocated and almost everything to do with how small the cells are. The program counts it: the same 350 channels, in cells of 5 km radius, serve 50 calls; in cells of 250 m they serve 30,750.

munotes.in602

Multiplexing: Space, Frequency, Time and Code

Space can also be divided by direction rather than distance: a sectorised antenna serves three or six sectors from one site, and a smart antenna can aim a beam per user, which is SDMA in one cell ([Directional Antennas, Sectorisation, Diversity and Spatial Reuse]).

Frequency

The band is cut into carriers, each one used by a signal all the time. This is the scheme of broadcasting: every station has its own frequency and never stops.

Its guard space is the guard band. A modulated carrier is not a line but a band, and its spectrum has skirts; the guard band keeps a strong neighbour out of a weak receiver.

The 2.4 GHz band shows the trade. 802.15.4 puts sixteen 2 MHz channels 5 MHz apart, spending 3 MHz of every 5 as guard and leaving room at the band edges. Bluetooth puts 79 channels of 1 MHz at 1 MHz spacing, with no guard at all, and relies on hopping and short packets to survive the collisions. 802.11b's channels are 22 MHz wide but only 5 MHz apart, so they overlap: thirteen are defined and only three can be used at once, which is the arrangement [The 802.15.4 Physical Layer] had to plan around.

Advantages: simple, each signal is continuous, no synchronisation between users is needed. Costs: a user's carrier is idle when it has nothing to send; the guard bands are wasted spectrum; and a receiver needs a filter sharp enough to reject its neighbours.

Time

All signals use the whole band, one after another, in slots of a repeating frame. GSM's frame is 4.615 ms in 8 slots; a mobile transmits in one slot and is free for the other seven, which is how a handset can also listen to its neighbours and save power.

Its guard space is the guard time. Two bursts must not overlap even though the mobiles are at different distances and their clocks are not perfect. GSM's slot is 156.25 bit periods long and its normal burst carries 148 of them, leaving 8.25 bit periods, about 30.5 microseconds, as guard. In that time a radio wave travels 9.1 km, which is why a mobile further away than that needs timing advance, an instruction to transmit early ([The GSM Radio Interface: Carriers, the TDMA Frame and Bursts]).

Advantages: the whole band is available to each user in turn, so a bursty user can be given more slots; only one transmitter is active, so no intermodulation between users. Costs: everyone must be synchronised; the guard time is wasted; and a user waits for its slot, which is latency.

Code

All signals use the whole band at the same time, each multiplied by its own code. If the codes are orthogonal, a receiver that multiplies the sum by its own code and adds up the result recovers its own signal and sees zero from the others. W-CDMA uses exactly this: orthogonal variable spreading factor codes to separate the channels of one cell, and a scrambling code to separate the cells.

munotes.in603

Multiplexing: Space, Frequency, Time and Code

Its guard space is the distance between the codes, and it is paid for in bandwidth: a spreading factor of 8 means 8 chips per bit, so the bit rate is one eighth of the chip rate. In return, code division brings gifts the other dimensions do not: a signal below the noise can still be recovered, multipath copies can be separated by their delays and combined, and there is no hard limit on the number of users, only a rising noise floor as each one is added (soft capacity).

Costs: every transmitter must be heard at about the same power, or a near one drowns the far ones (the near-far problem), so fast power control is essential; and the codes must be kept orthogonal, which needs synchronisation.

Combining them

No real system uses one dimension alone.

  • GSM: space (cells), frequency (200 kHz carriers) and time (8 slots), plus frequency hopping.
  • UMTS: space (cells), code (W-CDMA), and frequency (5 MHz carriers).
  • 802.15.4: frequency (16 channels), time (the superframe and its slots), and code (32 chips per symbol, for robustness rather than for sharing users).
  • Wi-Fi: frequency (channels), time (CSMA/CA), and, in later versions, space (MIMO streams).

The order in which they are applied is part of the design: GSM first divides the world into cells, then the band into carriers, then each carrier into slots, and only then hops the carrier from frame to frame.

Multiplexing, computed

The program counts carriers in the 2.4 GHz band against what each standard defines; computes GSM's guard time and the distance a burst can travel within it; separates four users sent at once with Walsh codes of length 8; counts the calls a city can carry as its cells shrink; and states the guard space each dimension costs.

# The four dimensions a medium is shared in, each measured: space, frequency,
# time and code. The guard space each needs, and what it costs.
import math
import random

# 1. Frequency division: a band cut into carriers, with a guard band between
#    them and at the edges. How many fit in the 83.5 MHz of the 2.4 GHz band,
#    and how many each standard actually defines.
print("Frequency division in the 2 400 to 2 483.5 MHz band:")
print("  system      carrier   spacing   gap between carriers   fit in 83.5 MHz   the standard defines")
for name, width, spacing, defined in (("802.15.4", 2.0, 5.0, 16), ("Bluetooth", 1.0, 1.0, 79),
                                      ("802.11b", 22.0, 5.0, 13)):
    fit = int((83.5 - width) // spacing) + 1
    gap = spacing - width
    print("  %-11s %5.1f MHz %6.1f MHz %19s %14d %19d"
          % (name, width, spacing, "%.1f MHz" % gap if gap >= 0 else "they overlap", fit, defined))
print("  802.15.4 takes 16 of the 17 that would fit, leaving 5 MHz clear at each end of the band.")
print("  802.11b's channels are wider than their spacing, so only %d of its 13 can be used at once."
      % int(83.5 // 22.0))

# 2. Time division: a frame of slots, with a guard time for clock drift and
#    propagation. GSM's own numbers: a frame of 4.615 ms in 8 slots.
print("\nTime division, GSM's frame: 4.615 ms in 8 slots")
slot = 4.615 / 8
print("  one slot %.4f ms; a normal burst carries 148 bits of the 156.25 bit periods in a slot,"
      % slot)
print("  so %.2f bit periods, %.1f microseconds, are the guard time between bursts."
      % (156.25 - 148, (156.25 - 148) * slot * 1000 / 156.25))
print("  in that guard a signal travels %.1f km, which sets how far a mobile may be without timing advance"
      % ((156.25 - 148) * slot * 1e-3 / 156.25 * 299792458 / 1000))

# 3. Code division: two users separated by orthogonal codes rather than by
#    frequency or time. Walsh codes of length 8, sent at once and separated.
def walsh(n):
    if n == 1:
        return [[1]]
    small = walsh(n // 2)
    return [row + row for row in small] + [row + [-x for x in row] for row in small]

codes = walsh(8)
rnd = random.Random(80)
users = [(0, 1), (3, -1), (5, 1), (6, -1)]              # (code index, bit sent as +1 or -1)
air = [sum(bit * codes[c][i] for c, bit in users) for i in range(8)]
print("\nCode division with Walsh codes of length 8:")
print("  four users send at once; the air carries %s" % air)
for c, bit in users:
    got = sum(a * codes[c][i] for i, a in enumerate(air)) / 8
    print("  user with code %d sent %+d, and correlating recovers %+.0f" % (c, bit, got))
idle = 2
print("  a listener correlating with unused code %d recovers %+.0f: the others are invisible to it"
      % (idle, sum(a * codes[idle][i] for i, a in enumerate(air)) / 8))

# 4. Space division: the same frequency, time and code reused at a distance.
#    How many cells of radius r fit in a city of side L, and the capacity that
#    gives with a cluster of k cells sharing the band.
print("\nSpace division: a city 10 km square, cells of radius r, a cluster of 7:")
print("   cell radius   cells   channels a cell gets (of 350)   calls at once")
for r in (5.0, 2.0, 1.0, 0.5, 0.25):
    area = 2.598 * r ** 2                               # a hexagon of circumradius r
    cells = max(1, int(100 / area))
    per_cell = 350 // 7
    print("  %10.2f km %7d %30d %14d" % (r, cells, per_cell, cells * per_cell))

# 5. What each dimension costs: the guard space, as a share.
print("\nThe guard space each dimension needs:")
print("  frequency: 802.15.4 uses 2 MHz of every 5 MHz channel spacing, so %.0f%% is guard" % (100 * 3 / 5))
print("  time:      GSM leaves %.2f of 156.25 bit periods, %.1f%%, as guard" % (156.25 - 148, 100 * 8.25 / 156.25))
print("  code:      %d chips carry 1 bit here, so the rate falls to 1/%d of the chip rate" % (8, 8))
print("  space:     a cluster of 7 gives each cell 1/7 of the channels, %.0f%% of the band" % (100 / 7))
munotes.in604

Multiplexing: Space, Frequency, Time and Code

Frequency division in the 2 400 to 2 483.5 MHz band:
  system      carrier   spacing   gap between carriers   fit in 83.5 MHz   the standard defines
  802.15.4      2.0 MHz    5.0 MHz             3.0 MHz             17                  16
  Bluetooth     1.0 MHz    1.0 MHz             0.0 MHz             83                  79
  802.11b      22.0 MHz    5.0 MHz        they overlap             13                  13
  802.15.4 takes 16 of the 17 that would fit, leaving 5 MHz clear at each end of the band.
  802.11b's channels are wider than their spacing, so only 3 of its 13 can be used at once.

Time division, GSM's frame: 4.615 ms in 8 slots
  one slot 0.5769 ms; a normal burst carries 148 bits of the 156.25 bit periods in a slot,
  so 8.25 bit periods, 30.5 microseconds, are the guard time between bursts.
  in that guard a signal travels 9.1 km, which sets how far a mobile may be without timing advance

Code division with Walsh codes of length 8:
  four users send at once; the air carries [0, 0, 4, 0, 0, 4, 0, 0]
  user with code 0 sent +1, and correlating recovers +1
  user with code 3 sent -1, and correlating recovers -1
  user with code 5 sent +1, and correlating recovers +1
  user with code 6 sent -1, and correlating recovers -1
  a listener correlating with unused code 2 recovers +0: the others are invisible to it

Space division: a city 10 km square, cells of radius r, a cluster of 7:
   cell radius   cells   channels a cell gets (of 350)   calls at once
        5.00 km       1                             50             50
        2.00 km       9                             50            450
        1.00 km      38                             50           1900
        0.50 km     153                             50           7650
        0.25 km     615                             50          30750

The guard space each dimension needs:
  frequency: 802.15.4 uses 2 MHz of every 5 MHz channel spacing, so 60% is guard
  time:      GSM leaves 8.25 of 156.25 bit periods, 5.3%, as guard
  code:      8 chips carry 1 bit here, so the rate falls to 1/8 of the chip rate
  space:     a cluster of 7 gives each cell 1/7 of the channels, 14% of the band
munotes.in605

Multiplexing: Space, Frequency, Time and Code

Frequency. Seventeen 2 MHz channels spaced 5 MHz apart would fit in 83.5 MHz; 802.15.4 defines sixteen, leaving 5 MHz clear at each end, because the band's edges must not be splashed ([The 802.15.4 Physical Layer]). Bluetooth's 79 channels of 1 MHz have no guard band at all. 802.11b's 22 MHz channels at 5 MHz spacing overlap, so of thirteen only three are usable at once: a standard can define more channels than the band can hold, and the count that matters is the non-overlapping one.

munotes.in606

Multiplexing: Space, Frequency, Time and Code

Time. A GSM slot is 0.5769 ms, and 8.25 of its 156.25 bit periods, 30.5 microseconds, are guard. A wave travels 9.1 km in that time, so a burst from further away arrives more than a guard period late and would land in the next slot: timing advance exists for exactly this.

Code. Four users with Walsh codes 0, 3, 5 and 6 send +1, -1, +1 and -1 at the same instant, and the air carries the sum, [0, 0, 4, 0, 0, 4, 0, 0]. Correlating with each user's code recovers exactly the bit that user sent, and correlating with an unused code recovers zero: the other users are not noise to a receiver with the right code, they are invisible. That is the whole idea of code division, and it works only while the codes stay orthogonal.

Space. With 350 channels and a cluster of 7, every cell gets 50 channels whatever its size. A single 5 km cell over a 10 km city carries 50 calls; 1 km cells carry 1,900; 250 m cells carry 30,750. Nothing about the spectrum changed. This is why operators build more sites rather than asking for more spectrum, and why [Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing] is about capacity, not coverage.

What each costs. 802.15.4 gives up 3 MHz of every 5 to guard bands, 60 per cent of the spacing; GSM gives up 5.3 per cent of every slot to guard time; the length-8 Walsh code gives up seven eighths of the chip rate; a cluster of 7 gives each cell one seventh, 14 per cent, of the channels. Every dimension is paid for.

Distinctions

Space (SDM)Frequency (FDM)Time (TDM)Code (CDM)
Signals separated byPlace or directionCarrier frequencySlot in a frameSpreading code
Guard spaceReuse distance, beam widthGuard bandGuard timeCode distance (bandwidth)
Each user hasA placeA carrier, all the timeThe whole band, some of the timeThe whole band, all the time
SynchronisationNot neededNot neededEssentialEssential
Program's cost1/7 of the channels per cell3 of every 5 MHz5.3 per cent of a slot7 of every 8 chips
Used byEvery cellular network, sectors, MIMOBroadcasting, GSM carriers, 802.15.4 channelsGSM slots, TDMA MACsW-CDMA, 802.15.4's spreading
munotes.in607

Multiplexing: Space, Frequency, Time and Code

MultiplexingMultiple access
QuestionHow is the medium divided?How are the divisions given to users?
Decided byThe system's designA protocol, often dynamically
NamesFDM, TDM, CDM, SDMFDMA, TDMA, CDMA, SDMA

What it does not mean

There is no fifth dimension. Every scheme is one of these four or a combination; polarisation and MIMO streams are refinements of space.

Guard space is not waste to be eliminated. It is what makes the separation work; Bluetooth's zero guard band is paid for in collisions.

Code division does not give infinite capacity. Each extra user raises the noise floor for the rest, which is soft capacity, not free capacity.

TDM is not the same as packet switching. A time slot is reserved for a user whether or not it has data; a packet network gives the medium to whoever has something to send.

Defining more channels is not having more. 802.11b defines thirteen channels in a band that holds three non-overlapping ones.

Quick revision

  • Four dimensions: space (SDM), frequency (FDM), time (TDM), code (CDM); as access schemes, SDMA, FDMA, TDMA, CDMA.
  • Space: reuse at a reuse distance, or by direction (sectors, smart antennas). Capacity comes from smaller cells: program, 50 calls with 5 km cells against 30,750 with 250 m cells, same spectrum.
  • Frequency: carriers plus guard bands. 2.4 GHz: 802.15.4 16 channels of 2 MHz, 5 MHz apart; Bluetooth 79 of 1 MHz, no guard; 802.11b 13 defined, 3 usable.
  • Time: slots in a frame plus a guard time. GSM: 4.615 ms, 8 slots, slot 156.25 bit periods, burst 148, guard 8.25 (30.5 us, 9.1 km), hence timing advance.
  • Code: orthogonal codes, the whole band at once. Program: four Walsh-8 users recovered exactly, an unused code sees zero. Costs bandwidth (1 bit per 8 chips) and needs power control (near-far).
  • Combined: GSM = space + frequency + time (+ hopping); UMTS = space + code + frequency; 802.15.4 = frequency + time + spreading.
munotes.in608

Multiplexing: Space, Frequency, Time and Code

Test yourself

1. Name the four dimensions in which a medium can be multiplexed, and give the guard space each requires. Space, where signals are separated by being in different places or directions and the guard space is the reuse distance or the width of a beam; frequency, where the band is divided into carriers and the guard space is a guard band between them; time, where users take turns in slots of a frame and the guard space is a guard time covering propagation delay and clock error; and code, where all users occupy the whole band at once with different spreading codes and the guard space is the distance between the codes, paid for in bandwidth.

2. Compare frequency division and time division multiplexing. In frequency division each signal has its own carrier and uses it continuously, so no synchronisation between users is needed and receivers need only a filter, but guard bands are wasted, a user's carrier is idle when it has nothing to send, and the whole band is never available to one user. In time division every user has the whole band in turn, so capacity can be moved between users by giving them more slots and only one transmitter is active at a time, but all users must be synchronised, a guard time is wasted in every slot, and a user must wait for its slot.

3. How does code division multiplexing separate users, and what does it cost? Each user multiplies its data by a distinctive code, and all users transmit over the whole band at the same time; a receiver multiplies the received sum by the code of the wanted user and accumulates, which recovers that user's data while orthogonal codes contribute zero. It costs bandwidth, since a spreading factor of n means n chips per bit and a rate of one nth of the chip rate; it requires the codes to remain orthogonal, hence synchronisation; and it requires fast power control, because a nearby transmitter can drown distant ones, which is the near-far problem.

4. Why is GSM's guard time 8.25 bit periods, and what happens beyond the distance it allows? A slot is 156.25 bit periods long but the normal burst carries only 148, leaving 8.25 bit periods, about 30.5 microseconds, so that bursts from mobiles at different distances and with imperfect clocks do not overlap. A radio wave travels about 9.1 km in that time, so a mobile further away than that would have its burst arrive more than a guard period late and overlap the next slot; GSM therefore uses timing advance, instructing distant mobiles to transmit correspondingly early.

munotes.in609

Multiplexing: Space, Frequency, Time and Code

5. Why does making cells smaller increase capacity, and by how much? Because space division reuses the same channels in every cluster of cells, so the number of simultaneous calls is the number of cells times the channels per cell, and the channels per cell depend only on the cluster size, not on the cell's radius. Halving the radius quarters the area of a cell and so roughly quadruples the number of cells and the capacity. In the chapter's model of a 10 km city with 350 channels and a cluster of 7, one 5 km cell carries 50 calls and 250 m cells carry 30,750, with no change in spectrum.

6. How does GSM combine the dimensions of multiplexing? It divides space into cells, each reusing frequencies at a reuse distance; within a cell the band is divided by frequency into 200 kHz carriers; each carrier is divided by time into a frame of eight slots; and the carrier used may hop from frame to frame, which adds frequency diversity. A particular call is therefore identified by its cell, its carrier and its slot, and is separated from every other call in at least one of those dimensions.

Contents This chapter on its own page

munotes.in610

Chapter Eighty-One

Modulation: ASK, FSK and PSK

Syllabus topic Module 2, "Wireless Transmission: Modulation"

In one line

Modulation puts information onto a carrier by changing one of its three parameters, and the digital forms are the three keyings: amplitude shift keying switches the carrier on and off, frequency shift keying sends one tone for a 1 and another for a 0, and phase shift keying flips the phase, which keeps the amplitude constant and puts the two symbols as far apart as the power allows, and so survives about 3 dB more noise than either of the others.

In the wording a student can write in an examination: modulation is the process of varying a property of a high-frequency carrier in step with a lower-frequency signal. It is needed because an antenna must be a fraction of a wavelength, so baseband signals would need impossibly large antennas; because only a modulated signal can be placed in an allocated band, allowing frequency division multiplexing; and because the higher frequency propagates and can be filtered as required. Analogue modulation varies the carrier continuously: amplitude modulation (AM), frequency modulation (FM) and phase modulation (PM). Digital modulation, or keying, selects among discrete values:

  • Amplitude shift keying (ASK): the amplitude is switched, in the simplest case on for 1 and off for 0 (on-off keying). Simple and cheap, but vulnerable to noise and to fading, since noise is amplitude.
  • Frequency shift keying (FSK): one frequency for 1 and another for 0. Constant envelope, so it tolerates a cheap non-linear amplifier; needs more bandwidth, roughly 2(deviation + bit rate) by Carson's rule.
  • Phase shift keying (PSK): the phase is shifted, in binary PSK by 180 degrees. Constant envelope and the most robust of the three, needing about 3 dB less signal than ASK or FSK for the same error rate, at the cost of a receiver that must recover the phase (coherent detection) or compare successive symbols (differential PSK).

Why modulate at all

Three reasons, and each is decisive on its own.

  • The antenna. An efficient antenna is a fraction of a wavelength ([Antennas: Radiators, Dipoles and Radiation Patterns]). A 3 kHz voice signal has a wavelength of 100 km; a quarter-wave antenna would be 25 km long. Carried on a 900 MHz carrier, the same information needs an 8 cm antenna.
  • Sharing. Every unmodulated signal occupies the same low frequencies, so only one could use the medium. Modulation moves each to its own band, which is frequency division multiplexing ([Multiplexing: Space, Frequency, Time and Code]).
  • Propagation and regulation. The band decides the range, the penetration and the licence ([Frequencies for Radio Transmission]); a system must sit where it is allowed to sit.
munotes.in611

Modulation: ASK, FSK and PSK

Analogue and digital

Analogue modulation varies the carrier continuously with a continuous signal: AM varies the amplitude (broadcast radio on medium wave), FM the frequency (broadcast radio on VHF, which tolerates noise better because noise is mostly amplitude), PM the phase.

Digital modulation varies it in steps, one step per symbol, and is called keying. In this book everything after this chapter is digital, but the two are connected: a digital modulator is an analogue modulator driven by a signal with only a few values.

Two facts to carry: a keying scheme with M possible symbols carries log2 M bits per symbol, and by Nyquist a channel of bandwidth B carries at most 2B symbols a second ([Signals: Amplitude, Frequency and Phase]).

The three keyings

Four rows of waveforms over four bit periods, 1, 0, 1, 1, with the periods where the bit is 1 shaded. The first row is the unmodulated carrier, a steady sine wave. The second, amplitude shift keying, shows the carrier present for a 1 and absent for a 0. The third, frequency shift keying, shows two cycles per bit period for a 1 and one for a 0, at constant height. The fourth, phase shift keying, shows a wave of constant height that starts upward for a 1 and downward for a 0, inverting at each change of bit

Figure 81.1 The same bits keyed three ways

Amplitude shift keying. The carrier's amplitude takes one of two values; in on-off keying the second is zero. It is the cheapest thing that works: a transmitter needs only to switch an oscillator, which is why remote controls, doorbells and the simplest 433 MHz modules use it.

Its weaknesses are exactly the ones fading causes. Noise adds to amplitude, so a threshold detector is easily fooled; and a fade changes the amplitude, which is the very quantity carrying the information, so the receiver must keep re-estimating the threshold. 802.15.4's optional 868/915 MHz PHY uses ASK, in the form of parallel sequence spread spectrum, precisely because spreading protects it.

Frequency shift keying. Two tones, one per bit. The amplitude never changes, so the transmitter's amplifier can be run hard into saturation, where it is efficient, and a fade does not destroy the information. Detection can be non-coherent, comparing the energy at the two tones, which needs no phase reference and makes for a simple, cheap receiver.

The cost is bandwidth: the signal occupies the two tones plus the sidebands of each, roughly 2(deviation + rate) by Carson's rule. Minimum shift keying is the choice of the smallest separation that keeps the tones orthogonal, and is where [Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM] begins.

Phase shift keying. The phase is shifted between fixed values; in BPSK, between 0 and 180 degrees. The amplitude is constant, and, crucially, the two symbols sit at opposite ends of the constellation, the furthest apart two points of a given power can be ([Signals: Amplitude, Frequency and Phase] draws them). That geometry is worth about 3 dB.

The price is the phase reference. A coherent receiver must recover the carrier's phase, which costs circuitry and can fail. Differential PSK (DPSK) avoids it by encoding each bit as a change of phase rather than an absolute phase, so the receiver compares each symbol with the previous one; it is simpler and loses a little performance.

munotes.in612

Modulation: ASK, FSK and PSK

What matters when choosing

  • Robustness: how much noise the scheme survives, which is the distance between its symbols for a given power. PSK wins.
  • Bandwidth: how much spectrum a given bit rate needs. ASK and PSK are narrower than FSK.
  • Constant envelope: whether the amplitude is fixed. FSK and PSK are; ASK is not. A constant envelope allows a non-linear power amplifier, which is far more efficient, and efficiency is battery life, which is why GSM, DECT and 802.15.4 all avoid amplitude keying in their mandatory modes.
  • Receiver complexity: non-coherent FSK is simplest, ASK next, coherent PSK hardest.
  • Bits per symbol: all three above send one bit per symbol; the next chapter's schemes send two, three, four or more.

The keyings, computed

The program prints the carrier and the three keyed waveforms for the same four bits; computes the bandwidth each needs at 250 kb/s; sends 200,000 bits of each through Gaussian noise at five signal levels and counts the errors; and lists what real systems use.

# The three basic keyings: amplitude, frequency and phase shift keying. What
# each sends for a 1 and a 0, how wide each is, and how each stands up to noise.
import math
import random

BITS = [1, 0, 1, 1]
S = 6                                   # samples per symbol, one carrier cycle per symbol
print("The carrier, keyed three ways, for the bits %s (%d samples a symbol):" % (BITS, S))
schemes = (("carrier alone", lambda b, p: math.sin(2 * math.pi * p)),
           ("ASK: 1 on, 0 off", lambda b, p: (1.0 if b else 0.0) * math.sin(2 * math.pi * p)),
           ("FSK: 1 at 2f, 0 at f", lambda b, p: math.sin(2 * math.pi * p * (2 if b else 1))),
           ("PSK: 1 at 0, 0 at 180", lambda b, p: math.sin(2 * math.pi * p + (0 if b else math.pi))))
print("  %-23s %s" % ("bit", "  ".join("%-27s" % ("%d" % b) for b in BITS)))
for name, wave in schemes:
    cells = [" ".join("%+.2f" % wave(b, (i + 0.5) / S) for i in range(S)) for b in BITS]
    print("  %-23s %s" % (name, " |".join(cells)))
print("  ASK changes the height, FSK the number of cycles, PSK where the wave starts.")

# 1. Bandwidth. A keyed signal at R symbols a second occupies about R hertz
#    around its carrier for PSK and ASK (the main lobe is 2R wide); FSK also
#    spends the separation between its two tones. Carson's rule.
print("\nBandwidth needed, at 250 kbit/s with one bit per symbol:")
R = 250e3
print("  ASK and PSK: main lobe 2R = %.0f kHz" % (2 * R / 1e3))
for sep in (250e3, 500e3, 1e6):
    print("  FSK with tones %4.0f kHz apart: Carson's rule, 2(deviation + R) = %.0f kHz"
          % (sep / 1e3, 2 * (sep / 2 + R) / 1e3))

# 2. Robustness. Send many symbols through noise and count the errors. The
#    detectors are the textbook ones: ASK compares the amplitude with a
#    threshold, FSK compares the energy at the two tones, PSK compares the
#    phase (here, the sign of the correlation with the carrier).
rnd = random.Random(81)
def errors(scheme, snr_db, n=200000):
    sigma = math.sqrt(1 / (2 * 10 ** (snr_db / 10)))
    wrong = 0
    for _ in range(n):
        bit = rnd.getrandbits(1)
        if scheme == "ASK":                       # 1 -> amplitude sqrt(2), 0 -> 0
            r = (math.sqrt(2) if bit else 0.0) + rnd.gauss(0, sigma)
            wrong += (r > math.sqrt(2) / 2) != bool(bit)
        elif scheme == "FSK":                     # energy in the wrong tone must win
            a = (1.0 if bit else 0.0) + rnd.gauss(0, sigma)
            b = (0.0 if bit else 1.0) + rnd.gauss(0, sigma)
            wrong += (a > b) != bool(bit)
        else:                                     # PSK: +1 or -1
            r = (1.0 if bit else -1.0) + rnd.gauss(0, sigma)
            wrong += (r > 0) != bool(bit)
    return 100 * wrong / n

print("\nBit errors in 200,000 bits, at the same average energy per bit:")
print("   Eb/N0     ASK       FSK       PSK")
for snr in (0, 4, 8, 10, 12):
    print("  %5d dB %7.3f%% %8.3f%% %8.3f%%" % (snr, errors("ASK", snr), errors("FSK", snr), errors("PSK", snr)))
print("  PSK needs about 3 dB less than ASK and FSK for the same error rate: its two symbols")
print("  are twice as far apart for the same average power.")

# 3. Where each is used, and why a sensor radio and GSM chose what they did.
print("\nWhat real systems use:")
for system, scheme, why in (
        ("IEEE 802.15.4 at 2450 MHz", "O-QPSK with half-sine pulses", "constant envelope, 4 bits a symbol"),
        ("IEEE 802.15.4 at 868/915 MHz", "BPSK (mandatory)", "the most robust keying there is"),
        ("GSM", "GMSK, BT = 0.3", "constant envelope, narrow spectrum"),
        ("EDGE", "8PSK", "3 bits a symbol in the same 200 kHz"),
        ("a cheap ISM remote", "ASK (on-off keying)", "the simplest transmitter that works")):
    print("  %-30s %-28s %s" % (system, scheme, why))
munotes.in613

Modulation: ASK, FSK and PSK

The carrier, keyed three ways, for the bits [1, 0, 1, 1] (6 samples a symbol):
  bit                     1                            0                            1                            1
  carrier alone           +0.50 +1.00 +0.50 -0.50 -1.00 -0.50 |+0.50 +1.00 +0.50 -0.50 -1.00 -0.50 |+0.50 +1.00 +0.50 -0.50 -1.00 -0.50 |+0.50 +1.00 +0.50 -0.50 -1.00 -0.50
  ASK: 1 on, 0 off        +0.50 +1.00 +0.50 -0.50 -1.00 -0.50 |+0.00 +0.00 +0.00 -0.00 -0.00 -0.00 |+0.50 +1.00 +0.50 -0.50 -1.00 -0.50 |+0.50 +1.00 +0.50 -0.50 -1.00 -0.50
  FSK: 1 at 2f, 0 at f    +0.87 +0.00 -0.87 +0.87 +0.00 -0.87 |+0.50 +1.00 +0.50 -0.50 -1.00 -0.50 |+0.87 +0.00 -0.87 +0.87 +0.00 -0.87 |+0.87 +0.00 -0.87 +0.87 +0.00 -0.87
  PSK: 1 at 0, 0 at 180   +0.50 +1.00 +0.50 -0.50 -1.00 -0.50 |-0.50 -1.00 -0.50 +0.50 +1.00 +0.50 |+0.50 +1.00 +0.50 -0.50 -1.00 -0.50 |+0.50 +1.00 +0.50 -0.50 -1.00 -0.50
  ASK changes the height, FSK the number of cycles, PSK where the wave starts.

Bandwidth needed, at 250 kbit/s with one bit per symbol:
  ASK and PSK: main lobe 2R = 500 kHz
  FSK with tones  250 kHz apart: Carson's rule, 2(deviation + R) = 750 kHz
  FSK with tones  500 kHz apart: Carson's rule, 2(deviation + R) = 1000 kHz
  FSK with tones 1000 kHz apart: Carson's rule, 2(deviation + R) = 1500 kHz

Bit errors in 200,000 bits, at the same average energy per bit:
   Eb/N0     ASK       FSK       PSK
      0 dB  15.858%   15.820%    7.946%
      4 dB   5.649%    5.654%    1.286%
      8 dB   0.571%    0.612%    0.018%
     10 dB   0.075%    0.076%    0.001%
     12 dB   0.004%    0.002%    0.000%
  PSK needs about 3 dB less than ASK and FSK for the same error rate: its two symbols
  are twice as far apart for the same average power.

What real systems use:
  IEEE 802.15.4 at 2450 MHz      O-QPSK with half-sine pulses constant envelope, 4 bits a symbol
  IEEE 802.15.4 at 868/915 MHz   BPSK (mandatory)             the most robust keying there is
  GSM                            GMSK, BT = 0.3               constant envelope, narrow spectrum
  EDGE                           8PSK                         3 bits a symbol in the same 200 kHz
  a cheap ISM remote             ASK (on-off keying)          the simplest transmitter that works
munotes.in614

Modulation: ASK, FSK and PSK

The waveforms. For the bits 1, 0, 1, 1 the carrier is unchanged; ASK is present and then absent; FSK completes two cycles in a bit period instead of one; and PSK keeps its height and inverts, so the samples that were +0.50, +1.00, +0.50 become -0.50, -1.00, -0.50. Three different things done to one carrier, one each to its amplitude, frequency and phase.

Bandwidth. At 250 kb/s, ASK and PSK occupy a main lobe of 500 kHz. FSK with its tones 250 kHz apart needs 750 kHz by Carson's rule, and 1.5 MHz if the tones are 1 MHz apart. Wider tones are easier to tell apart and cost spectrum: that trade is the whole of FSK's design.

Robustness. At the same energy per bit, ASK and FSK make errors at almost exactly the same rate (5.649 and 5.654 per cent at 4 dB), while PSK makes them at 1.286 per cent, and at 8 dB the gap is 0.571 and 0.612 per cent against 0.018. Reading it the other way: PSK reaches a given error rate with about 3 dB less power, a factor of two. The reason is geometric. On-off keying's symbols are 0 and the peak; BPSK's are plus and minus the peak, twice as far apart for the same average power, since the transmitter is never idle.

munotes.in615

Modulation: ASK, FSK and PSK

What systems chose. 802.15.4's mandatory 868/915 MHz PHY is BPSK; its 2450 MHz PHY is O-QPSK with half-sine pulses, four bits a symbol and a constant envelope; GSM is GMSK with BT = 0.3; EDGE raises GSM's rate by moving to 8PSK in the same 200 kHz. Every one of them keeps the envelope constant. The exception is the cheap ISM remote control, where on-off keying survives because nothing about it has to be efficient.

Distinctions

ASKFSKPSK
What changesAmplitudeFrequencyPhase
Constant envelopeNoYesYes
Bandwidth at rate RAbout 2R2(deviation + R)About 2R
Noise toleranceLowestSame as ASKAbout 3 dB better
DetectionThreshold on amplitudeNon-coherent, by energy in each toneCoherent, or differential
Affected by fadingDirectly: the amplitude is the messageLittleLittle
Used byRemote controls; 802.15.4's optional ASK PHYPagers, simple radios, Bluetooth's GFSK802.15.4 BPSK and O-QPSK, EDGE's 8PSK
Analogue modulationDigital modulation (keying)
InputA continuous signalA stream of symbols
OutputA continuously varying carrierOne of a fixed set of carrier states
FormsAM, FM, PMASK, FSK, PSK and their combinations
Quality measured bySignal-to-noise ratioBit error rate
Coherent detectionDifferential detection
NeedsA recovered carrier phaseOnly the previous symbol
EncodesThe absolute phaseThe change of phase
PerformanceBestA little worse
ComplexityHigherLower

What it does not mean

Modulation is not encoding. Encoding maps bits to symbols or adds redundancy; modulation puts symbols onto a carrier.

On-off keying is not free of a carrier. It still needs one; it just switches it.

Constant envelope does not mean constant power at the receiver. The channel still fades; it means the transmitter's own amplitude does not carry information.

FSK's extra bandwidth is not waste. It buys a receiver that needs no phase reference, and tones far enough apart to survive drift.

PSK's 3 dB is not a property of phase. It is the distance between the symbols; any scheme that puts its symbols at opposite ends of the constellation earns it.

Quick revision

  • Why modulate: antenna size (a fraction of a wavelength), sharing the medium (FDM), and propagation and regulation.
  • Analogue: AM, FM, PM. Digital (keying): ASK, FSK, PSK. M symbols carry log2 M bits; Nyquist: 2B symbols a second.
  • ASK: amplitude switched, on-off keying; cheapest; hurt by noise and fading; not constant envelope.
  • FSK: two tones; constant envelope; non-coherent detection possible; bandwidth about 2(deviation + rate) (Carson).
  • PSK: phase shifted, BPSK by 180 degrees; constant envelope; symbols furthest apart, so about 3 dB better than ASK and FSK; needs coherent detection or differential encoding.
  • Program: at 250 kb/s, 500 kHz for ASK and PSK, 750 kHz for FSK with 250 kHz tones; errors at 8 dB, 0.571 / 0.612 / 0.018 per cent for ASK / FSK / PSK.
  • Real systems: 802.15.4 BPSK (868/915) and O-QPSK (2450), GSM GMSK BT = 0.3, EDGE 8PSK, cheap remotes on-off keying.
munotes.in616

Modulation: ASK, FSK and PSK

Test yourself

1. Why is modulation necessary in wireless communication? Because an efficient antenna must be a sizeable fraction of a wavelength, and a baseband signal of a few kilohertz has a wavelength of tens of kilometres, so its antenna would be impossibly large; modulating it onto a high-frequency carrier reduces the antenna to centimetres. Because unmodulated signals all occupy the same low frequencies and could not share a medium, while modulation moves each signal to its own band, allowing frequency division multiplexing. And because the choice of carrier frequency decides propagation, interference and what a regulator permits.

2. Distinguish analogue and digital modulation, and name the forms of each. Analogue modulation varies a parameter of the carrier continuously in proportion to a continuous message signal: amplitude modulation, frequency modulation and phase modulation. Digital modulation, or keying, sets the parameter to one of a finite set of values, one per symbol: amplitude shift keying, frequency shift keying and phase shift keying, together with their combinations such as QAM. Analogue quality is measured by the signal-to-noise ratio, digital quality by the bit error rate.

3. Describe ASK, FSK and PSK and compare them on bandwidth, robustness and complexity. In amplitude shift keying the carrier's amplitude takes one of two values, often on and off, which makes the simplest transmitter but suffers from noise and fading because the information is carried in the amplitude. In frequency shift keying one tone represents a 1 and another a 0, the amplitude is constant, and detection can be non-coherent by comparing the energy at the two tones, but the signal is wider, about 2(deviation + rate). In phase shift keying the phase is shifted, in the binary case by 180 degrees, the amplitude is constant, and the two symbols are as far apart as the power allows, so it needs about 3 dB less power than the other two for the same error rate, at the cost of recovering the carrier phase or using differential encoding.

4. Why does binary PSK outperform on-off keying by about 3 dB? Because the distance between the two symbols decides how easily noise turns one into the other. On-off keying's symbols are zero and the peak amplitude, so the transmitter is idle half the time and its average power buys a separation of one peak. BPSK transmits at full amplitude always, with symbols at plus and minus the peak, so for the same average power the separation is twice as large, which is a factor of four in energy and a factor of two, about 3 dB, in the signal-to-noise ratio needed. The chapter's simulation confirms it: at 8 dB, 0.571 per cent of bits were wrong with ASK and 0.018 per cent with PSK.

munotes.in617

Modulation: ASK, FSK and PSK

5. What is a constant envelope, and why does it matter for a battery-powered device? A modulation has a constant envelope when the amplitude of the transmitted carrier never changes, so only the frequency or phase carries information; FSK and PSK have one and ASK does not. It matters because a constant-envelope signal can be amplified by a non-linear power amplifier driven into saturation, which is far more efficient than a linear amplifier, and the power amplifier is the largest consumer in a radio. That is why GSM, DECT and IEEE 802.15.4 all use constant-envelope schemes.

6. What is differential phase shift keying, and why is it used? In differential PSK each bit is encoded as a change of phase from the previous symbol rather than as an absolute phase, so the receiver compares each symbol with its predecessor instead of recovering the carrier's absolute phase. It is used because coherent detection requires a carrier recovery circuit that costs complexity and can lose lock, especially on a fading channel; differential detection is simpler and more robust to a slowly changing phase, at the price of slightly worse error performance.

Contents This chapter on its own page

munotes.in618

Chapter Eighty-Two

Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM

Syllabus topic Module 2, "Wireless Transmission: Modulation"

In one line

Once one bit per symbol is not enough, a modulator can move the phase in smaller steps and never jump (MSK and GMSK, which keep the spectrum narrow and the envelope constant), place four, eight or sixty-four points in the plane and send several bits at a time (QPSK, 8PSK, QAM, at the price of packing the points closer together), or split the data across hundreds of slow carriers so that each symbol is long enough to ignore the echoes (OFDM).

In the wording a student can write in an examination: Minimum shift keying (MSK) is continuous-phase FSK with the smallest frequency separation that keeps the two tones orthogonal, a modulation index of 0.5; the phase advances or retreats by exactly 90 degrees per bit and never jumps, so the envelope is constant and the spectrum compact. Gaussian MSK (GMSK) passes the bit stream through a Gaussian filter before the modulator, which rounds the phase transitions further and narrows the spectrum at the cost of some intersymbol interference; GSM uses GMSK with a filter whose "BT = 0.3", at a symbol rate of "1 625/6 ksymb/s (i.e. approximately 270.833 ksymb/s)".

Quadrature phase shift keying (QPSK) sends two bits per symbol as one of four phases, doubling the rate in the same bandwidth for the same error performance per bit as BPSK. Offset QPSK (O-QPSK) delays the quadrature stream by half a symbol so that only one of the two components changes at a time, which limits the phase step to 90 degrees and keeps the envelope nearly constant; 802.15.4's 2450 MHz PHY uses O-QPSK with half-sine pulse shaping. Quadrature amplitude modulation (QAM) varies amplitude and phase together, giving 16, 64 or more points, hence 4, 6 or more bits per symbol, but placing them closer together, so each step up needs more signal-to-noise ratio. Orthogonal frequency division multiplexing (OFDM) divides the data over many narrow subcarriers, spaced so that each one's spectrum is zero at the others' centres (orthogonality); every subcarrier's symbol is long, so a multipath echo occupies a small fraction of it, and a guard interval (cyclic prefix) absorbs the rest. It underlies Wi-Fi, LTE, DVB-T and DAB.

MSK: a phase that turns

Ordinary FSK switches between two tones, and at each switch the phase can jump, which splashes energy into neighbouring channels. Continuous-phase FSK forbids the jump; MSK is the case where the two tones are as close as they can be while remaining orthogonal, a separation of half the bit rate.

Seen in the phase domain, MSK is elegant: over one bit period the phase advances by exactly +90 degrees for a 1 or -90 for a 0, smoothly. The amplitude never changes, the phase never jumps, and the spectrum is narrower than ordinary FSK's. It can also be seen as offset QPSK with half-sine pulses, which is why it sits between the two families.

munotes.in619

Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM

GMSK: GSM's modulation

GSM narrows the spectrum further by filtering the bit stream before it reaches the modulator. 3GPP specifies a Gaussian filter and fixes its shape by the product of its 3 dB bandwidth and the bit period: "BT = 0.3", where "B is the 3 dB bandwidth of the filter".

A narrower filter spreads each bit's influence over more than one bit period, which is deliberate intersymbol interference, and the equaliser in every GSM receiver undoes it ([Multipath, Fading and the Doppler Effect] explains why the receiver needs one anyway). What GSM buys is a spectrum tight enough to put carriers 200 kHz apart, and a constant envelope that lets the handset's amplifier run in saturation, which is battery life.

The symbol rate is fixed by the standard: "The modulating symbol rate is the normal symbol rate which is defined as 1/T = 1 625/6 ksymb/s (i.e. approximately 270.833 ksymb/s), which corresponds to 1 625/6 kbit/s (i.e. 270.833 kbit/s)". One bit per symbol, which is what EDGE later changed.

QPSK, offset QPSK and beyond

QPSK places four points on the circle, 90 degrees apart, so each symbol carries two bits. For the same energy per bit it has the same error rate as BPSK, while using half the bandwidth for a given bit rate: the two bits ride on the in-phase and quadrature components, which are independent. That is why QPSK, not BPSK, is the workhorse.

Offset QPSK fixes QPSK's one flaw. In plain QPSK both bits can change at once, which swings the phase by 180 degrees and takes the signal through the origin, so the envelope collapses to zero and a non-linear amplifier distorts it. Offsetting the quadrature stream by half a symbol means only one component changes at a time, so the phase moves by at most 90 degrees and the envelope stays nearly constant.

802.15.4 does exactly this, and adds shaping: "The chip sequences representing each data symbol are modulated onto the carrier using O-QPSK with half-sine pulse shaping. Even-indexed chips are modulated onto the in-phase (I) carrier and odd-indexed chips are modulated onto the quadrature-phase (Q) carrier." With half-sine pulses, O-QPSK is MSK.

8PSK puts eight points on the circle, three bits a symbol, and is how EDGE raises GSM's rate without new spectrum ([New Data Services: HSCSD, GPRS and EDGE]).

QAM: amplitude and phase together

Past eight points, the circle is crowded: the points are close together while the middle of the plane is empty. QAM uses the whole plane, varying amplitude as well as phase, and arranges its points in a square grid: 16-QAM has 16 points and 4 bits per symbol, 64-QAM has 64 and 6 bits.

munotes.in620

Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM

The cost is unavoidable. For a fixed average power, adding points brings them closer, and the nearest distance is what noise has to overcome; the program measures the shrinkage. The other cost is the amplitude: QAM does not have a constant envelope, so it needs a linear amplifier, which is less efficient, and it needs an accurate estimate of the channel's gain and phase. That is why QAM belongs to systems with power and processing, and why sensor radios stay with BPSK and O-QPSK.

Modern systems therefore adapt: a station near the base uses 64-QAM, one at the edge falls back to QPSK. The modulation is chosen per link, per moment.

OFDM: many slow carriers

The problem OFDM solves is the delay spread. A fast single-carrier link has symbols shorter than the echoes, so each symbol is smeared over its neighbours and an equaliser must undo it, which becomes harder as rates rise ([Multipath, Fading and the Doppler Effect] measured 20 symbols of smearing for a 20 Msymbol/s link in an urban channel).

OFDM turns the problem around. Instead of one carrier at 20 Msymbol/s, use 256 carriers at 78 ksymbol/s each. Every symbol is now 256 times longer, so the same echo covers a small fraction of it, and a guard interval (a cyclic prefix, a copy of the symbol's end placed at its start) absorbs what is left. Each subcarrier sees a channel that is essentially flat, so the equaliser becomes one complex multiplication per subcarrier.

The carriers are packed as closely as they can be without interfering: the spacing is the reciprocal of the symbol time, so each carrier's spectrum has a null at the centre of every other. That is the orthogonality in the name, and it is why OFDM is spectrally efficient as well as robust.

Its costs are a high peak-to-average power ratio, since hundreds of carriers can add in phase, which demands a linear amplifier and wastes battery, and sensitivity to frequency error and Doppler, which destroy the orthogonality. DVB-T, DAB, Wi-Fi from 802.11a onward, LTE and 5G all use it; sensor radios do not.

The modulations, computed

The program computes what a 200 kHz channel carries under each scheme; measures the nearest distance between constellation points at equal average power; sends 40,000 symbols of each through noise; computes OFDM's symbol length against an urban delay spread; and walks MSK's phase through five bits.

# The modulations that carry more than one bit a symbol, and the one that
# carries many at once: MSK and GMSK, QPSK and its offset form, QAM, and OFDM.
import cmath
import math
import random

# 1. Bits per symbol, and the rate each buys in a given channel. Nyquist:
#    a channel of B hertz carries 2B symbols a second.
print("What a 200 kHz channel (GSM's) can carry, by Nyquist's 2B symbols a second:")
print("  scheme        points   bits/symbol   symbols/s   bit rate")
for name, m in (("BPSK", 2), ("QPSK / 4-QAM", 4), ("8PSK", 8), ("16-QAM", 16), ("64-QAM", 64)):
    bits = math.log2(m)
    print("  %-14s %5d %12.0f %11s %10s"
          % (name, m, bits, "400 k", "%.0f kb/s" % (400 * bits)))
print("  GSM sends 270.833 ksymbol/s of GMSK, 1 bit a symbol; EDGE keeps the rate and uses 8PSK.")

# 2. The constellations, and how far apart their points are. For the same
#    average power, more points means closer points and less noise tolerated.
def constellation(name):
    if name == "BPSK":
        return [1, -1]
    if name == "QPSK":
        return [cmath.exp(1j * (math.pi / 4 + k * math.pi / 2)) for k in range(4)]
    if name == "8PSK":
        return [cmath.exp(1j * k * math.pi / 4) for k in range(8)]
    side = 4 if name == "16-QAM" else 8
    pts = [complex(i, q) for i in range(-side + 1, side, 2) for q in range(-side + 1, side, 2)]
    power = sum(abs(p) ** 2 for p in pts) / len(pts)
    return [p / math.sqrt(power) for p in pts]

print("\nAt the same average power, how close the nearest two points are:")
print("  scheme       points   nearest distance   against BPSK")
for name in ("BPSK", "QPSK", "8PSK", "16-QAM", "64-QAM"):
    pts = constellation(name)
    d = min(abs(a - b) for i, a in enumerate(pts) for b in pts[i + 1:])
    print("  %-12s %6d %18.3f %14.1f dB" % (name, len(pts), d, 20 * math.log10(d / 2)))

# 3. Errors through noise, measured, for the same energy per bit.
rnd = random.Random(82)
def symbol_errors(name, ebn0_db, n=40000):
    pts = constellation(name)
    bits = math.log2(len(pts))
    sigma = math.sqrt(1 / (2 * bits * 10 ** (ebn0_db / 10)))
    wrong = 0
    for _ in range(n):
        s = rnd.randrange(len(pts))
        r = pts[s] + complex(rnd.gauss(0, sigma), rnd.gauss(0, sigma))
        wrong += min(range(len(pts)), key=lambda k: abs(pts[k] - r)) != s
    return 100 * wrong / n

print("\nSymbols decoded wrongly (40,000 symbols, same energy per bit):")
print("   Eb/N0     BPSK     QPSK     8PSK   16-QAM   64-QAM")
for db in (4, 8, 12, 16):
    print("  %5d dB" % db + "".join("%9.3f" % symbol_errors(n, db)
                                    for n in ("BPSK", "QPSK", "8PSK", "16-QAM", "64-QAM")))

# 4. OFDM: many slow carriers instead of one fast one. Each carrier's symbol
#    is long compared with the delay spread, so the smearing is small.
print("\nOFDM: one fast carrier against many slow ones, over a 1.03 us delay spread (3GPP urban):")
print("  carriers   symbol time   spread as a share of a symbol   guard interval needed")
for n in (1, 16, 64, 256, 1024):
    total_rate = 20e6                                    # symbols a second in total
    sym = n / total_rate
    print("  %8d %11.2f us %30.1f%% %17.2f us" % (n, sym * 1e6, 100 * 1.03e-6 / sym, 1.03))
print("  with 256 carriers the echo covers 8 per cent of a symbol, so a short guard absorbs it.")

# 5. MSK and GMSK: the phase turns instead of jumping. How far the phase moves
#    in one symbol, and what GSM's filter does to the spectrum.
print("\nMSK: the phase advances a quarter turn (90 degrees) per bit, never jumping:")
phase = 0.0
for bit in (1, 0, 0, 1, 1):
    phase += 90 if bit else -90
    print("  bit %d -> phase %+4.0f degrees" % (bit, phase))
print("  GSM filters the bit stream with a Gaussian filter of BT = 0.3 before this, which")
print("  smooths the turns further and narrows the spectrum: that is GMSK.")
munotes.in621

Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM

What a 200 kHz channel (GSM's) can carry, by Nyquist's 2B symbols a second:
  scheme        points   bits/symbol   symbols/s   bit rate
  BPSK               2            1       400 k   400 kb/s
  QPSK / 4-QAM       4            2       400 k   800 kb/s
  8PSK               8            3       400 k  1200 kb/s
  16-QAM            16            4       400 k  1600 kb/s
  64-QAM            64            6       400 k  2400 kb/s
  GSM sends 270.833 ksymbol/s of GMSK, 1 bit a symbol; EDGE keeps the rate and uses 8PSK.

At the same average power, how close the nearest two points are:
  scheme       points   nearest distance   against BPSK
  BPSK              2              2.000            0.0 dB
  QPSK              4              1.414           -3.0 dB
  8PSK              8              0.765           -8.3 dB
  16-QAM           16              0.632          -10.0 dB
  64-QAM           64              0.309          -16.2 dB

Symbols decoded wrongly (40,000 symbols, same energy per bit):
   Eb/N0     BPSK     QPSK     8PSK   16-QAM   64-QAM
      4 dB    1.180    2.487   13.777   22.250   57.365
      8 dB    0.025    0.055    1.895    3.627   29.175
     12 dB    0.000    0.000    0.028    0.070    5.685
     16 dB    0.000    0.000    0.000    0.000    0.147

OFDM: one fast carrier against many slow ones, over a 1.03 us delay spread (3GPP urban):
  carriers   symbol time   spread as a share of a symbol   guard interval needed
         1        0.05 us                         2060.0%              1.03 us
        16        0.80 us                          128.8%              1.03 us
        64        3.20 us                           32.2%              1.03 us
       256       12.80 us                            8.0%              1.03 us
      1024       51.20 us                            2.0%              1.03 us
  with 256 carriers the echo covers 8 per cent of a symbol, so a short guard absorbs it.

MSK: the phase advances a quarter turn (90 degrees) per bit, never jumping:
  bit 1 -> phase  +90 degrees
  bit 0 -> phase   +0 degrees
  bit 0 -> phase  -90 degrees
  bit 1 -> phase   +0 degrees
  bit 1 -> phase  +90 degrees
  GSM filters the bit stream with a Gaussian filter of BT = 0.3 before this, which
  smooths the turns further and narrows the spectrum: that is GMSK.
munotes.in622

Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM

Four constellation diagrams side by side, all with the same dashed circle of average power: BPSK with two points on the horizontal axis and a nearest distance of 2.00; QPSK with four points at the diagonals, 1.41 apart; 8PSK with eight points around the circle, 0.77 apart; and 16-QAM with sixteen points in a square grid, 0.63 apart. A note says that more points means more bits a symbol and less room between them for noise

Figure 82.1 Four constellations at the same average power

munotes.in623

Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM

Bits and rates. A 200 kHz channel carries 400,000 symbols a second by Nyquist, so 400 kb/s with BPSK, 800 with QPSK, 1.2 Mb/s with 8PSK and 2.4 Mb/s with 64-QAM. Nothing about the spectrum changes; only how much each symbol means. GSM chose one bit per symbol and 270.833 ksymbol/s; EDGE kept the symbol rate and moved to 8PSK, tripling the rate in the same carrier.

What the points cost. At equal average power, BPSK's two points are 2.000 apart. QPSK's four are 1.414 apart, which is 3.0 dB closer, but each symbol carries two bits, so per bit nothing is lost, and that is why QPSK is free. 8PSK's are 0.765 apart, 8.3 dB down, for only 50 per cent more bits than QPSK: past four points the circle starts charging. 16-QAM is 10.0 dB down and 64-QAM 16.2 dB.

Errors. At 8 dB per bit, BPSK gets 0.025 per cent of symbols wrong and QPSK 0.055, essentially the same per bit; 8PSK gets 1.895 per cent wrong, 16-QAM 3.627 and 64-QAM 29.175. To reach the error rate BPSK has at 8 dB, 64-QAM needs about 16 dB and more. Every extra bit per symbol is bought with signal-to-noise ratio, which is why adaptive modulation exists.

OFDM. At 20 Msymbol/s in total, one carrier gives a symbol of 0.05 microseconds, and the urban delay spread of 1.03 microseconds is 2,060 per cent of it: hopeless without a powerful equaliser. With 64 carriers the symbol is 3.2 microseconds and the spread is 32 per cent of it; with 256 carriers, 12.8 microseconds and 8 per cent, so a guard interval of about a microsecond absorbs the echo and each subcarrier sees a flat channel. That single table is the reason every high-rate wireless system built since 1999 uses OFDM.

MSK's phase. Bits 1, 0, 0, 1, 1 take the phase to +90, 0, -90, 0 and +90 degrees: a quarter turn each way, never a jump, so the envelope is constant and the spectrum stays tight. GSM filters the stream before this to round the turns further, which is GMSK.

Distinctions

MSKGMSK
What it isContinuous-phase FSK at the minimum separationMSK with a Gaussian filter before the modulator
Phase per bitExactly 90 degrees, smoothlyThe same, rounded further
SpectrumCompactNarrower still
CostNone beyond FSKDeliberate intersymbol interference, needing an equaliser
Used byThe idea behind O-QPSK with half-sine pulsesGSM (BT = 0.3)
munotes.in624

Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM

BPSKQPSK8PSK16-QAM64-QAM
Points2481664
Bits a symbol12346
Nearest distance (equal power)2.0001.4140.7650.6320.309
Errors at 8 dB per bit0.025%0.055%1.895%3.627%29.175%
Constant envelopeYesOnly in the offset formNearlyNoNo
Single carrierOFDM
Symbol time at 20 Msymbol/s0.05 us12.8 us with 256 carriers
Urban echo (1.03 us)2,060 per cent of a symbol8 per cent
EqualiserComplex, over many symbolsOne multiplication per subcarrier
Peak to average powerLowHigh: needs a linear amplifier
Sensitive toDelay spreadFrequency error and Doppler

What it does not mean

More bits per symbol is not more capacity for free. Each step needs more signal-to-noise ratio; 64-QAM at the cell edge delivers nothing.

QPSK is not twice as fragile as BPSK. Per bit it performs the same; it is 8PSK and beyond that start paying.

OFDM does not remove multipath. It makes each subcarrier's channel flat, so that a simple correction suffices; the deep fades are still there, on some subcarriers, and coding across them is what saves the data.

GMSK's filter is not a mistake. The intersymbol interference it creates is the price of a narrow spectrum, and the equaliser was needed anyway.

Offset QPSK is not a different constellation. It has the same four points; only the timing of the two components differs, so the path between points avoids the origin.

Quick revision

  • MSK: continuous-phase FSK, modulation index 0.5; phase turns 90 degrees per bit, no jumps; constant envelope.
  • GMSK: MSK with a Gaussian filter, GSM's BT = 0.3, symbol rate 1625/6 ksymbol/s, about 270.833 ksymbol/s, 1 bit a symbol; narrow spectrum, 200 kHz carriers, needs an equaliser.
  • QPSK: 4 points, 2 bits a symbol, same error rate per bit as BPSK. O-QPSK: quadrature delayed by half a symbol, phase steps of at most 90 degrees, no path through the origin; 802.15.4 uses it with half-sine pulses, which makes it MSK.
  • QAM: amplitude and phase; 16-QAM 4 bits, 64-QAM 6 bits; not constant envelope; needs a linear amplifier and channel estimation; used adaptively.
  • OFDM: many orthogonal subcarriers, spacing = 1 / symbol time; long symbols plus a guard interval (cyclic prefix) defeat the delay spread; flat fading per subcarrier; high peak-to-average ratio, sensitive to frequency error. Wi-Fi, LTE, DVB-T, DAB.
  • Program: 200 kHz gives 400 k symbols/s, so 400 kb/s (BPSK) to 2.4 Mb/s (64-QAM); nearest distance 2.000 / 1.414 / 0.765 / 0.632 / 0.309; errors at 8 dB 0.025 / 0.055 / 1.895 / 3.627 / 29.175 per cent; OFDM with 256 carriers cuts the urban echo from 2,060 to 8 per cent of a symbol.
munotes.in625

Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM

Test yourself

1. What is MSK, and how does GSM's GMSK differ from it? Minimum shift keying is continuous-phase frequency shift keying with the smallest tone separation that keeps the two tones orthogonal, which makes the phase advance by exactly 90 degrees for a 1 or retreat by 90 for a 0 over each bit period, smoothly and without jumps, so the envelope is constant and the spectrum compact. GMSK passes the bit stream through a Gaussian filter before the modulator, rounding the phase transitions further and narrowing the spectrum; GSM specifies a filter with BT = 0.3, where B is the filter's 3 dB bandwidth and T the bit period. The filter spreads each bit over more than one bit period, which the receiver's equaliser must undo.

2. Why does QPSK give twice the bit rate of BPSK with no loss of error performance per bit? Because its four points are carried on two independent components, in phase and quadrature, each of which is effectively a BPSK signal. A symbol therefore carries two bits, and for a given energy per bit the distance between the points along each component is the same as BPSK's, so the error rate per bit is unchanged; the constellation's nearest distance falls by 3 dB, but each symbol carries twice as much, which exactly compensates.

3. What problem does offset QPSK solve, and how? In ordinary QPSK both bits of a symbol can change at once, which rotates the phase by 180 degrees and takes the signal through the origin, so the envelope momentarily collapses; a non-linear amplifier then distorts the signal and splashes energy into neighbouring channels. Offset QPSK delays the quadrature bit stream by half a symbol period so that only one component ever changes at a time, limiting the phase step to 90 degrees and keeping the envelope nearly constant; IEEE 802.15.4 uses this form with half-sine pulse shaping at 2450 MHz.

4. What does QAM gain and what does it cost? QAM varies both the amplitude and the phase, so its points fill the plane rather than a circle: 16-QAM carries 4 bits per symbol and 64-QAM carries 6, multiplying the bit rate in a given bandwidth. The cost is that for a fixed average power the points are closer together, so more signal-to-noise ratio is needed: the chapter measured nearest distances of 0.632 for 16-QAM and 0.309 for 64-QAM against 2.000 for BPSK, and error rates at 8 dB per bit of 3.627 and 29.175 per cent against 0.025. QAM also lacks a constant envelope, so it requires a linear amplifier and an accurate estimate of the channel.

munotes.in626

Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM

5. Explain OFDM and why it defeats multipath. OFDM divides the data across many narrow subcarriers, spaced by the reciprocal of the symbol time so that each one's spectrum is zero at the centres of the others, which is what makes them orthogonal. Because the data is spread over many carriers, each symbol lasts far longer than in a single-carrier system of the same total rate, so a multipath echo occupies only a small fraction of a symbol; a guard interval, usually a cyclic prefix, absorbs what remains, and each subcarrier then sees a channel that is flat and can be corrected by a single complex multiplication. In the chapter's example a 1.03 microsecond urban echo covers 2,060 per cent of a single-carrier symbol at 20 Msymbol/s but only 8 per cent of a symbol with 256 subcarriers.

6. What are the drawbacks of OFDM? Its peak-to-average power ratio is high, because many subcarriers can momentarily add in phase, so the transmitter needs a linear amplifier operated below its maximum, which wastes power; and it is sensitive to frequency offset and to Doppler, since either destroys the orthogonality of the subcarriers and creates interference between them. It also needs a guard interval, which is overhead, and more signal processing than a single carrier.

Contents This chapter on its own page

munotes.in627

Chapter Eighty-Three

Spread Spectrum and Direct Sequence

Syllabus topic Module 2, "Wireless Transmission: Spread spectrum"

In one line

Spread spectrum deliberately spends far more bandwidth than the data needs, multiplying each bit by a fast code, because at the receiver the same code collapses the wanted signal back into a narrow band while spreading everything else out: narrowband interference is averaged away, the transmitted signal can sit below the noise, multipath copies can be told apart, and several users with different codes can share the same band at the same time.

In the wording a student can write in an examination: in spread spectrum the transmitted signal occupies a bandwidth much greater than the minimum the data requires, and the spreading is done by a code known to the receiver. In direct sequence spread spectrum (DSSS) each data bit or symbol is multiplied by a faster chipping sequence (a pseudo-noise or PN code) of n chips, so the signal's bandwidth is multiplied by n and its power spectral density divided by n. The receiver multiplies by the same code and integrates (correlates), which restores the wanted signal to full strength while any signal uncorrelated with the code, including narrowband interference, stays spread and is largely averaged away. The processing gain is n, or 10 log10 n decibels.

The benefits are resistance to narrowband interference and jamming, a low power spectral density (the signal can be hidden in the noise, giving a low probability of interception), resistance to multipath (echoes that are delayed by more than a chip decorrelate and can be separated, or combined by a rake receiver), and code division multiple access, since users with different codes can share one band. The costs are the bandwidth itself, the need for synchronisation of the code at the receiver, and, in CDMA, the near-far problem, which demands power control. A good code has a sharp autocorrelation (a large peak at zero shift and small values elsewhere) and low cross-correlation with the other codes in use. IEEE 802.15.4 at 2450 MHz spreads each 4-bit symbol into 32 chips at 2.0 Mchip/s, a processing gain of about 15 dB.

Why spend bandwidth on purpose

Every other chapter in this part of the book treats bandwidth as precious. Spread spectrum spends it, and gets four things back.

  • Interference is averaged away. A narrowband interferer is not correlated with the code, so despreading spreads it out while collapsing the wanted signal. What reaches the decision is the interferer's power divided by roughly n.
  • The signal is quiet. The same power over n times the bandwidth is 1/n of the power spectral density, so the transmission can be below the noise floor of a receiver that does not know the code: hard to detect, hard to intercept, and gentle on other users of the band.
  • Multipath becomes separable. An echo delayed by more than one chip correlates poorly with the code at the right instant, so it behaves like unrelated noise rather than smearing the symbol. A rake receiver goes further and correlates at several delays at once, adding the echoes together as diversity.
  • Users can share. If two codes have low cross-correlation, two transmissions can occupy the same band at the same time and each receiver can pull out its own ([Multiplexing: Space, Frequency, Time and Code]).
munotes.in628

Spread Spectrum and Direct Sequence

For 802.15.4 the reason is the first two: its band is shared with Wi-Fi and microwave ovens, and its annex says the modulation "achieves low signal-to-noise ratio (SNR) and signal-to-interference ratio (SIR) requirements at the expense of a signal bandwidth that is significantly larger than the symbol rate."

Direct sequence: chips and codes

The transmitter takes each data bit, or in 802.15.4 each 4-bit symbol, and replaces it with a fixed sequence of chips sent at a much higher rate. Multiplying by the code (in the bipolar arithmetic of plus and minus one) is the whole operation:

  • Bit 1 is sent as the code.
  • Bit 0 is sent as the code inverted.

The receiver multiplies the incoming chips by the same code and adds them up over a symbol. If the code lines up, every chip contributes the same sign, so the sum is n times the amplitude; anything else contributes a mixture of signs and largely cancels. That one operation is despreading, detection and interference rejection at once.

What makes a good code. Two properties:

  • Autocorrelation: correlating the code with a shifted copy of itself should give a large value only at zero shift. That lets a receiver find where a symbol starts, and makes delayed echoes harmless.
  • Cross-correlation: correlating with another user's code should give nearly zero, so users do not see each other.

Barker codes are the classic example of the first property; 802.15.4's sixteen 32-chip sequences are built for the second, being cyclic shifts and conjugations of one another ([The 802.15.4 Physical Layer] verified that structure).

Spread spectrum, computed

The program prints the 11-chip Barker code and its autocorrelation; tabulates processing gain against code length; measures how much a narrowband jammer is weakened by despreading, for three codes and two kinds of jammer; and puts two users in one band with different codes.

# Spread spectrum, direct sequence: spread a bit over many chips, add a
# narrowband jammer and noise, then despread and see what survives.
import math
import random

BARKER11 = [1, -1, 1, 1, -1, 1, 1, 1, -1, -1, -1]      # the 11-chip Barker code
print("The 11-chip Barker code: %s" % BARKER11)

# 1. Why this code: its autocorrelation. A Barker code correlates to 11 with
#    itself and to at most 1 with any shift of itself, so a receiver can find
#    the start of a symbol as well as decode it.
def correlate(a, b):
    return sum(x * y for x, y in zip(a, b))

print("\nAutocorrelation, the code against shifted copies of itself:")
print("   shift  " + " ".join("%4d" % s for s in range(-5, 6)))
row = []
for s in range(-5, 6):
    shifted = [0] * 11
    for i in range(11):
        j = i + s
        if 0 <= j < 11:
            shifted[i] = BARKER11[j]
    row.append(correlate(BARKER11, shifted))
print("   value  " + " ".join("%4d" % v for v in row))
print("  a peak of 11 at zero shift, and never more than 1 away from it: that is what makes it a Barker code")

# 2. Processing gain: the spread signal occupies 11 times the bandwidth, and
#    despreading multiplies the wanted signal by 11 while narrowband
#    interference is spread out and averaged away.
print("\nProcessing gain of a code of length n: 10 log10(n)")
for n in (11, 15, 32, 64, 128, 1023):
    print("  n = %4d: %5.1f dB of gain; bandwidth %4d times the data rate" % (n, 10 * math.log10(n), n))
print("  802.15.4 at 2450 MHz spreads 32 chips to a symbol: %.1f dB." % (10 * math.log10(32)))

# 3. Spread, jam, despread. A narrowband jammer is a tone: over one symbol it
#    is nearly the same in every chip. Despreading multiplies by the code and
#    sums, which lifts the wanted signal by n and leaves the jammer summing
#    almost at random. Measured over many random jammer frequencies and phases.
rnd = random.Random(83)
def jam_ratio(code, spread_of_offsets=0.02, trials=20000):
    n = len(code)
    before, after = 0.0, 0.0
    for _ in range(trials):
        offset = rnd.uniform(0, spread_of_offsets)   # where the tone sits in the band
        phase = rnd.uniform(0, 2 * math.pi)
        jam = [math.cos(2 * math.pi * offset * k + phase) for k in range(n)]
        before += sum(j * j for j in jam) / n      # jammer power per chip, wanted power is 1
        after += (sum(j * c for j, c in zip(jam, code)) / n) ** 2
    return before / trials, after / trials

for code, name in ((BARKER11, "Barker 11"), ([1] * 32, "a code of all ones, 32 chips"),
                   ([1 if rnd.getrandbits(1) else -1 for _ in range(32)], "a random code, 32 chips")):
    b, a = jam_ratio(code)
    print("\n%s: jammer power per chip %.3f, after despreading %.4f" % (name, b, a))
    print("  the jammer is %.1f dB weaker relative to the wanted bit; 10 log10(n) would be %.1f dB"
          % (10 * math.log10(b / a), 10 * math.log10(len(code))))
print("  a code of all ones spreads nothing: it is the code's changes of sign that scatter the jammer.")

# the figure 10 log10(n) is the average over jammers anywhere in the band, not
# against one tone sitting on the carrier.
print("\nThe same codes against a jammer anywhere in the spread band:")
for code, name in ((BARKER11, "Barker 11"),
                   ([1 if rnd.getrandbits(1) else -1 for _ in range(32)], "a random code, 32 chips")):
    b, a = jam_ratio(code, spread_of_offsets=0.5)
    print("  %-28s %5.1f dB of rejection; 10 log10(n) = %.1f dB"
          % (name, 10 * math.log10(b / a), 10 * math.log10(len(code))))

# 4. Two users at once, separated by their codes rather than by frequency.
print("\nTwo users sharing the band with different codes:")
other = [1, 1, -1, 1, 1, 1, -1, -1, -1, 1, -1]
print("  their codes correlate to %+d out of 11, so each sees the other as %.1f%% of its own signal"
      % (correlate(BARKER11, other), 100 * abs(correlate(BARKER11, other)) / 11))
for a_bit, b_bit in ((1, 1), (1, -1)):
    air = [a_bit * x + b_bit * y for x, y in zip(BARKER11, other)]
    print("  A sends %+d and B sends %+d: A's receiver gets %+3d, B's gets %+3d"
          % (a_bit, b_bit, correlate(air, BARKER11), correlate(air, other)))
munotes.in629

Spread Spectrum and Direct Sequence

The 11-chip Barker code: [1, -1, 1, 1, -1, 1, 1, 1, -1, -1, -1]

Autocorrelation, the code against shifted copies of itself:
   shift    -5   -4   -3   -2   -1    0    1    2    3    4    5
   value     0   -1    0   -1    0   11    0   -1    0   -1    0
  a peak of 11 at zero shift, and never more than 1 away from it: that is what makes it a Barker code

Processing gain of a code of length n: 10 log10(n)
  n =   11:  10.4 dB of gain; bandwidth   11 times the data rate
  n =   15:  11.8 dB of gain; bandwidth   15 times the data rate
  n =   32:  15.1 dB of gain; bandwidth   32 times the data rate
  n =   64:  18.1 dB of gain; bandwidth   64 times the data rate
  n =  128:  21.1 dB of gain; bandwidth  128 times the data rate
  n = 1023:  30.1 dB of gain; bandwidth 1023 times the data rate
  802.15.4 at 2450 MHz spreads 32 chips to a symbol: 15.1 dB.

Barker 11: jammer power per chip 0.503, after despreading 0.0087
  the jammer is 17.6 dB weaker relative to the wanted bit; 10 log10(n) would be 10.4 dB

a code of all ones, 32 chips: jammer power per chip 0.499, after despreading 0.3335
  the jammer is 1.7 dB weaker relative to the wanted bit; 10 log10(n) would be 15.1 dB

a random code, 32 chips: jammer power per chip 0.501, after despreading 0.0341
  the jammer is 11.7 dB weaker relative to the wanted bit; 10 log10(n) would be 15.1 dB
  a code of all ones spreads nothing: it is the code's changes of sign that scatter the jammer.

The same codes against a jammer anywhere in the spread band:
  Barker 11                     10.5 dB of rejection; 10 log10(n) = 10.4 dB
  a random code, 32 chips       15.0 dB of rejection; 10 log10(n) = 15.1 dB

Two users sharing the band with different codes:
  their codes correlate to -1 out of 11, so each sees the other as 9.1% of its own signal
  A sends +1 and B sends +1: A's receiver gets +10, B's gets +10
  A sends +1 and B sends -1: A's receiver gets +12, B's gets -12
munotes.in630

Spread Spectrum and Direct Sequence

The Barker code's autocorrelation. Against itself the code correlates to 11; against every shift, to 0 or -1. That is the defining property of a Barker code, and it does two jobs: a receiver sliding the code along the incoming chips sees one unmistakable peak, so it finds the symbol boundary; and an echo arriving a chip or more late adds at most 1 instead of 11, so multipath is suppressed rather than accumulated.

munotes.in631

Spread Spectrum and Direct Sequence

Processing gain. Eleven chips give 10.4 dB, 32 give 15.1 dB, and the 1023-chip code of GPS gives 30.1 dB. The gain is the bandwidth multiplier expressed in decibels, and it is also the factor by which the transmitted power spectral density falls.

Despreading a jammer. Against a tone sitting near the carrier, the Barker code weakened the jammer by 17.6 dB, better than its nominal 10.4, while a random 32-chip code managed 11.7 dB against a nominal 15.1. The reason is that a slow tone is nearly a constant, so what matters is the code's own sum: the Barker code sums to +1 out of 11 and cancels a constant almost completely, while a random code can have a larger sum. A code of all ones, which is no spreading at all, gave 1.7 dB, and that is the control that proves the mechanism: it is the changes of sign in the code that scatter the jammer.

Measured properly, against a jammer anywhere in the spread band, both codes deliver exactly their processing gain: 10.5 dB for the Barker code against a nominal 10.4, and 15.0 dB for the 32-chip code against 15.1. So the rule 10 log10 n is an average over the band, not a promise about one particular interferer, and a system that must survive a tone at a known frequency should choose its code with that frequency in mind.

munotes.in632

Spread Spectrum and Direct Sequence

Two users. The two 11-chip codes here correlate to -1, so each user sees the other at about 9 per cent of its own signal. When both send +1 the air carries their sum and each receiver recovers +10 instead of +11; when they send opposite bits, each recovers +12 and -12. The bits come through, but the answer is not clean: real CDMA systems use longer codes, chosen to be orthogonal, and control the transmit powers so that no user arrives loud enough to swamp another.

Distinctions

NarrowbandDirect sequence spread spectrum
BandwidthAbout the symbol raten times greater
Power spectral densityHighDivided by n: can be below the noise
Narrowband interferenceDamages the signalReduced by about n after despreading
MultipathSmears symbolsEchoes past a chip decorrelate; a rake can combine them
SharingBy frequency or timeBy code as well
NeedsA filterThe code, and synchronisation to it
AutocorrelationCross-correlation
ComparesA code with a shifted copy of itselfTwo different codes
WantedA sharp peak at zero shift, small elsewhereNear zero always
BuysSymbol timing; multipath rejectionUsers sharing a band
ProgramBarker 11: 11 at zero shift, at most 1 elsewhereTwo codes: -1 of 11, so 9 per cent leakage
Processing gainWhat it is not
DefinitionThe bandwidth multiplier, 10 log10 n dBNot extra transmit power
Against a band-wide jammerDelivered exactly: 10.5 and 15.0 dB measured
Against one toneDepends on the code's spectrum there: 17.6 dB or 11.7 dB measuredNot a guarantee for every interferer

What it does not mean

Spreading does not add power. The same energy is spread over more hertz; the gain appears only after despreading, and only for signals that do not share the code.

A wider signal is not automatically spread spectrum. The bandwidth must come from a code the receiver knows, not from the data rate.

Processing gain is not error-correcting coding. It buys nothing against noise that is spread across the band in the same way as the signal; against white noise, a spread system and a narrowband one with the same energy per bit perform the same.

Low probability of interception is not encryption. Someone who knows the code hears everything; secrecy needs cryptography ([Keys and Link Security: Key Predistribution, SPINS and 802.15.4]).

Codes are rarely perfectly orthogonal in practice. Delays and imperfect synchronisation leave cross-correlation, which is why CDMA systems need power control.

Quick revision

  • Spread spectrum: bandwidth far beyond what the data needs, spread by a code the receiver knows.
  • DSSS: each bit or symbol becomes n chips; bandwidth x n, power spectral density / n; the receiver correlates with the same code. Processing gain = n = 10 log10 n dB.
  • Benefits: narrowband interference rejection, low power spectral density (low probability of interception), multipath resistance and rake combining, CDMA.
  • Costs: bandwidth, code synchronisation, the near-far problem (so power control).
  • A good code: sharp autocorrelation (Barker 11: 11 at zero shift, at most 1 elsewhere), low cross-correlation.
  • Program: gain 10.4 dB (n = 11), 15.1 dB (n = 32), 30.1 dB (n = 1023); against a band-wide jammer the measured rejection is 10.5 and 15.0 dB, matching the rule; against a tone on the carrier it was 17.6 dB (Barker) and 11.7 dB (a random code), and 1.7 dB for a code of all ones, which proves the sign changes do the work.
  • 802.15.4 at 2450 MHz: 32 chips a symbol, 2.0 Mchip/s, about 15 dB of gain.
munotes.in633

Spread Spectrum and Direct Sequence

Test yourself

1. What is spread spectrum, and why is bandwidth spent deliberately? It is any transmission whose bandwidth is much larger than the minimum the data requires, where the spreading is imposed by a code known to the receiver rather than by the data itself. The bandwidth buys four things: narrowband interference is spread out by the despreading operation that collapses the wanted signal, so it is reduced by the processing gain; the transmitted power is spread thinly, so the signal has a low power spectral density and can even sit below the noise, which makes it hard to detect and gentle on other users; multipath echoes delayed by more than a chip decorrelate and can be rejected or combined; and users with different codes can share one band.

2. Explain how direct sequence spreading and despreading work. The transmitter multiplies each data bit, in plus and minus one form, by a chipping sequence of n chips sent at n times the bit rate, so a 1 is sent as the code and a 0 as the code inverted; the signal now occupies n times the bandwidth. The receiver multiplies the incoming chips by the same code, aligned in time, and sums over the symbol: the wanted signal's chips all contribute the same sign and add to n times the amplitude, while any signal uncorrelated with the code contributes a mixture of signs and largely cancels. The ratio of the two effects is the processing gain, n, or 10 log10 n decibels.

3. What is processing gain, and what does it not protect against? Processing gain is the ratio of the spread bandwidth to the data bandwidth, equal to the number of chips per symbol, expressed as 10 log10 n decibels: 10.4 dB for 11 chips, 15.1 dB for the 32 chips of IEEE 802.15.4 at 2450 MHz. It protects against interference that is not correlated with the code, in particular narrowband interference, which is reduced by about that factor. It does not protect against white noise spread across the whole band in the same way as the signal, because such noise is not concentrated for the despreader to scatter; against white noise a spread and an unspread system with the same energy per bit perform the same.

munotes.in634

Spread Spectrum and Direct Sequence

4. What makes a good spreading code? Illustrate with the Barker code. It must have a sharp autocorrelation, so that correlating it with a shifted copy of itself gives a large value only at zero shift, and low cross-correlation with the other codes in use. The 11-chip Barker code correlates to 11 with itself and to 0 or at most 1 with every shift, which lets a receiver find the symbol boundary unmistakably and makes an echo arriving a chip or more late contribute at most 1 instead of 11. Low cross-correlation is what lets two users share a band; in the chapter's example two 11-chip codes correlated to only -1, so each user saw the other at about 9 per cent of its own signal.

5. In the program, why did a code of all ones give almost no protection? Because spreading works by the code's changes of sign. Despreading multiplies the received chips by the code and sums; for an interferer to cancel, the multiplication must give it a mixture of positive and negative contributions. A code of all ones multiplies every chip by the same value, so a slowly varying interferer adds up coherently just as the wanted signal does, and the ratio between them is unchanged: the measured rejection was 1.7 dB rather than the 15.1 dB its length would suggest.

6. Why did the measured rejection differ from 10 log10 n against a tone on the carrier? Because the processing gain is an average over interference spread across the band, not a guarantee for one frequency. A tone close to the carrier is nearly a constant over a symbol, so what decides its rejection is the sum of the code's chips: the Barker code sums to only 1 out of 11 and cancelled such a tone by 17.6 dB, better than its nominal 10.4, while a random 32-chip code with a larger sum achieved 11.7 dB rather than 15.1. Against a jammer anywhere in the band, both codes delivered their nominal gain, 10.5 and 15.0 dB.

Contents This chapter on its own page

munotes.in635

Chapter Eighty-Four

Frequency Hopping Spread Spectrum

Syllabus topic Module 2, "Wireless Transmission: Spread spectrum"

In one line

Frequency hopping spreads a signal by moving it: the transmitter changes carrier many times a second on a pseudo-random sequence the receiver follows, so a narrowband interferer or a fading channel can spoil only the hops that happen to land on it, and two systems sharing a band collide only occasionally instead of constantly.

In the wording a student can write in an examination: in frequency hopping spread spectrum (FHSS) the carrier frequency is changed repeatedly according to a hop sequence known to both ends; the signal therefore occupies a wide band over time although it is narrow at any instant. In slow frequency hopping one hop carries several symbols (GSM hops once per burst); in fast frequency hopping one symbol is spread over several hops, which gives frequency diversity within a symbol at the cost of a much faster synthesiser. The processing gain is the number of channels hopped over. The benefits are resistance to narrowband interference and jamming (only the hops that land on the interferer are hurt), frequency diversity against frequency-selective fading, and sharing, since two networks with different sequences collide only when their hops coincide. The costs are synchronisation (both ends must know the sequence and the time), the settling time of the synthesiser between hops, and collisions with other hoppers. GSM hops once per TDMA frame using a mobile allocation (MA) of up to 64 carriers, a hopping sequence number (HSN) shared by the cell and a mobile allocation index offset (MAIO) that differs per mobile, so mobiles in one cell never collide while mobiles in different cells collide only at random. DSSS spreads by multiplying with a fast code and is narrow in time but wide in frequency at every instant; FHSS is narrow at every instant and wide over time.

Hopping, and why it spreads

A direct sequence system is wide all the time. A hopping system is narrow at any instant, and wide only when watched over many hops. Both satisfy the definition of spread spectrum, because the bandwidth is decided by a code rather than by the data.

What hopping buys is avoidance by movement. An interferer on one channel, a channel in a deep fade, another network that is busy: all of them cost only the hops that land there. The rest of the sequence is untouched, so error-correcting coding across hops can repair the loss. That is the whole strategy, and it is why hopping is at its best when the enemy is narrow and fixed, as a microwave oven or another network's carrier is.

Slow and fast. If a hop lasts longer than a symbol, the hopping is slow: GSM, Bluetooth and TSCH are all slow hoppers, and a hop carries hundreds of symbols. If a symbol is spread over several hops, the hopping is fast: every symbol then samples several channels, so even a symbol whose hop lands in a fade is partly received elsewhere, which is diversity within the symbol. Fast hopping needs a synthesiser that settles in a fraction of a symbol, which is expensive, so it belongs to military systems.

munotes.in636

Frequency Hopping Spread Spectrum

The sequence. Both ends must generate the same sequence at the same time. It should look random, so that an adversary cannot predict the next channel and other systems collide only at chance; it should use every channel about equally; and, where several users share the band, sequences must be chosen so that they do not collide systematically. GSM's algorithm does all three, and the program runs it.

GSM's hopping, exactly as specified

3GPP TS 45.002 defines the mapping from a frame number to a carrier. Three parameters set it up:

  • MA, the mobile allocation: "Mobile allocation of radio frequency channels, defines the set of radio frequency channels to be used in the mobiles hopping sequence. The MA contains N radio frequency channels, where 1 <= N <= 64."
  • MAIO, the mobile allocation index offset, "(0 to N-1, 6 bits)".
  • HSN, the hopping sequence (generator) number, "(0 to 63, 6 bits)".

The algorithm gives "the index to an absolute radio frequency channel number (ARFCN) within the mobile allocation (MAI from 0 to N-1, where MAI=0 represents the lowest ARFCN in the mobile allocation".

With HSN = 0 the hopping is cyclic: "MAI = (FN + MAIO) modulo N", simply stepping through the list. With any other HSN the sequence is pseudo-random, computed from the frame number through three time parameters and a fixed table of 114 numbers, the RNTABLE, and finally offset by the MAIO.

The design has a neat property. Mobiles in the same cell are given the same HSN and different MAIOs, so their sequences are the same sequence shifted, and two mobiles never land on the same carrier in the same frame. Mobiles in different cells are given different HSNs, so their sequences are unrelated and collisions happen only at random, which spreads interference between cells evenly instead of concentrating it on one unlucky pair. The program verifies both.

Hopping, computed

The program implements GSM's generator with its own RNTABLE and prints the sequences for a cyclic and a pseudo-random case, checks that two MAIOs never collide and that the carriers are used evenly, measures what a jammer sitting on one of eight channels costs for different hop rates, and separates slow from fast hopping by counting symbols per hop.

munotes.in637

Frequency Hopping Spread Spectrum

# Frequency hopping: slow and fast, what a jammer on one channel can do, and
# GSM's own hopping sequence generator, run.
import random

# 1. GSM's algorithm, 3GPP TS 45.002 clause 6.2.3, with its RNTABLE.
RNTABLE = [48, 98, 63, 1, 36, 95, 78, 102, 94, 73, 0, 64, 25, 81, 76, 59, 124, 23, 104, 100,
           101, 47, 118, 85, 18, 56, 96, 86, 54, 2, 80, 34, 127, 13, 6, 89, 57, 103, 12, 74,
           55, 111, 75, 38, 109, 71, 112, 29, 11, 88, 87, 19, 3, 68, 110, 26, 33, 31, 8, 45,
           82, 58, 40, 107, 32, 5, 106, 92, 62, 67, 77, 108, 122, 37, 60, 66, 121, 42, 51, 126,
           117, 114, 4, 90, 43, 52, 53, 113, 120, 72, 16, 49, 7, 79, 119, 61, 22, 84, 9, 97,
           91, 15, 21, 24, 46, 39, 93, 105, 65, 70, 125, 99, 17, 123]

def mai(fn, n, hsn, maio):
    """The index into the mobile allocation for frame number fn."""
    if hsn == 0:                                    # cyclic hopping
        return (fn + maio) % n
    t1r, t2, t3 = (fn // (26 * 51)) % 64, fn % 26, fn % 51
    m = t2 + RNTABLE[(hsn ^ t1r) + t3]
    nbin = n.bit_length()
    mp, tp = m % (2 ** nbin), t3 % (2 ** nbin)
    s = mp if mp < n else (mp + tp) % n
    return (s + maio) % n

MA = [12, 20, 28, 36, 44, 52, 60, 68]               # 8 carriers a mobile may use
print("GSM hopping (TS 45.002, 6.2.3) over %d carriers %s" % (len(MA), MA))
for hsn, maio, label in ((0, 0, "HSN 0, MAIO 0: cyclic"), (13, 0, "HSN 13, MAIO 0"), (13, 3, "HSN 13, MAIO 3")):
    seq = [MA[mai(fn, len(MA), hsn, maio)] for fn in range(16)]
    print("  %-24s %s" % (label, " ".join("%3d" % f for f in seq)))
print("  two mobiles in one cell share the HSN and differ in MAIO, so they never collide:")
a = [mai(fn, len(MA), 13, 0) for fn in range(2000)]
b = [mai(fn, len(MA), 13, 3) for fn in range(2000)]
print("  in 2000 frames, MAIO 0 and MAIO 3 land on the same carrier %d times" % sum(x == y for x, y in zip(a, b)))

# how evenly the sequence uses the carriers
from collections import Counter
counts = Counter(mai(fn, len(MA), 13, 0) for fn in range(20000))
print("  over 20000 frames each of the 8 carriers is used %d to %d times" % (min(counts.values()), max(counts.values())))

# 2. A jammer parks on one channel. With hopping, only the hops that land on it
#    are lost; without, everything is.
print("\nOne channel of %d jammed, and what a frame loses:" % len(MA))
print("  hops per frame   frames lost outright   frames with some hops lost   bits lost on average")
for hops in (1, 4, 8, 26):
    trials, lost, damaged, total_bits = 20000, 0, 0, 0.0
    rnd = random.Random(84)
    for _ in range(trials):
        landed = [rnd.randrange(len(MA)) == 0 for _ in range(hops)]
        lost += all(landed)
        damaged += any(landed)
        total_bits += sum(landed) / hops
    print("  %14d %22.2f%% %28.2f%% %20.2f%%"
          % (hops, 100 * lost / trials, 100 * damaged / trials, 100 * total_bits / trials))
print("  hopping does not avoid the jammer; it spreads the damage so that coding can repair it.")

# 3. Slow against fast hopping: how many hops a symbol or a frame spans.
print("\nSlow and fast hopping:")
for name, hop_rate, symbol_rate in (("GSM, one hop a burst", 217.0, 270833.0),
                                    ("Bluetooth, 1600 hops/s", 1600.0, 1e6),
                                    ("a fast hopper", 20000.0, 10000.0)):
    per_hop = symbol_rate / hop_rate
    kind = "slow: many symbols a hop" if per_hop >= 1 else "fast: many hops a symbol"
    print("  %-24s %8.1f symbols a hop   %s" % (name, per_hop, kind))
munotes.in638

Frequency Hopping Spread Spectrum

GSM hopping (TS 45.002, 6.2.3) over 8 carriers [12, 20, 28, 36, 44, 52, 60, 68]
  HSN 0, MAIO 0: cyclic     12  20  28  36  44  52  60  68  12  20  28  36  44  52  60  68
  HSN 13, MAIO 0            20  60  68  28  68  28  12  36  68  12  20  12  44  28  44  52
  HSN 13, MAIO 3            44  20  28  52  28  52  36  60  28  36  44  36  68  52  68  12
  two mobiles in one cell share the HSN and differ in MAIO, so they never collide:
  in 2000 frames, MAIO 0 and MAIO 3 land on the same carrier 0 times
  over 20000 frames each of the 8 carriers is used 2455 to 2558 times

One channel of 8 jammed, and what a frame loses:
  hops per frame   frames lost outright   frames with some hops lost   bits lost on average
               1                  12.29%                        12.29%                12.29%
               4                   0.03%                        41.09%                12.38%
               8                   0.00%                        65.53%                12.47%
              26                   0.00%                        96.91%                12.50%
  hopping does not avoid the jammer; it spreads the damage so that coding can repair it.

Slow and fast hopping:
  GSM, one hop a burst       1248.1 symbols a hop   slow: many symbols a hop
  Bluetooth, 1600 hops/s      625.0 symbols a hop   slow: many symbols a hop
  a fast hopper                 0.5 symbols a hop   fast: many hops a symbol

The sequences. With HSN 0 the mobile walks its eight carriers in order: 12, 20, 28, 36, 44, 52, 60, 68, and round again. With HSN 13 the order is scrambled and repeats nothing obvious: 20, 60, 68, 28, 68, 28, 12, 36. Notice that the same carrier can come twice in a row and that 68 appears twice in the first five hops: a pseudo-random sequence is allowed to do that, and does.

munotes.in639

Frequency Hopping Spread Spectrum

No collisions within a cell. MAIO 3 produces the same sequence shifted three places along the allocation, and in 2,000 frames the two mobiles landed on the same carrier zero times. That is worth pausing on: two mobiles hopping pseudo-randomly over eight carriers would collide about a quarter of the time if their sequences were independent. GSM gets orthogonality inside a cell and randomness between cells out of one generator, by splitting the parameters into a shared HSN and a per-mobile MAIO.

Evenness. Over 20,000 frames the eight carriers were used between 2,455 and 2,558 times, against an ideal 2,500: even to within about 2 per cent. A hop sequence that favoured some carriers would waste the diversity that hopping exists to provide.

What a jammer costs. With one hop per frame, 12.29 per cent of frames land on the jammed channel and are lost outright. With four hops per frame, only 0.03 per cent of frames are wholly lost, but 41 per cent are damaged somewhere; with 26 hops, essentially every frame is damaged and none is lost. The average fraction of bits lost is the same 12.5 per cent in every case, which is the honest lesson: hopping does not reduce the interference, it redistributes it. What it buys is that the damage arrives as scattered errors inside many frames, which an error-correcting code can repair, instead of as whole frames destroyed, which it cannot.

Slow and fast. GSM hops 217 times a second and sends 270,833 symbols a second, so 1,248 symbols ride on each hop: slow hopping. Bluetooth's 1,600 hops a second at about 1 Msymbol/s gives 625 symbols a hop: also slow. Only a hopper faster than its own symbol rate, 20,000 hops a second against 10,000 symbols, gets several hops into one symbol, which is fast hopping.

Distinctions

Direct sequence (DSSS)Frequency hopping (FHSS)
At any instantWide: the whole spread bandNarrow: one channel
Over timeThe same bandThe whole hopping band
Spreading done byMultiplying by a fast codeChanging the carrier on a sequence
Processing gainChips per symbolNumber of channels
Against a narrowband jammerAverages it down by the gainAvoids it except on the hops that land there
Against another userSeen as noise, alwaysCollides only when hops coincide
NeedsChip-level synchronisationHop timing and the sequence
HardwareA correlatorA fast synthesiser
Examples802.15.4, W-CDMA, GPSGSM, Bluetooth, 802.15.4e TSCH
Slow hoppingFast hopping
DefinitionSeveral symbols per hopSeveral hops per symbol
DiversityAcross symbols, so coding is neededWithin a symbol
SynthesiserModestFast and expensive
ProgramGSM 1,248 symbols a hop; Bluetooth 6250.5 symbols a hop
Used byGSM, Bluetooth, TSCHMilitary systems
munotes.in640

Frequency Hopping Spread Spectrum

Same cellDifferent cells
HSNThe sameDifferent
MAIODifferent per mobileIrrelevant
ResultSequences shifted: never collideUnrelated: collide at random
Program0 collisions in 2,000 framesNot modelled

What it does not mean

Hopping does not dodge interference. The average fraction of time spent on a jammed channel is unchanged; the damage is spread out so that coding can repair it.

A hopping signal is not wideband at an instant. It is a narrowband signal that moves; its instantaneous spectrum is that of the underlying modulation.

A pseudo-random sequence is not a secret. GSM's generator is published; secrecy needs cryptography, and military hoppers keep their sequence key secret for that reason.

More hops per second is not automatically better. Each hop costs settling time, and the sequence must still be tracked; the gain comes from spreading damage, which a modest hop rate already achieves.

FHSS and DSSS are not rivals in every system. 802.15.4 spreads with a code, and its 2012 amendment hops as well; the two are combined in TSCH.

Quick revision

  • FHSS: the carrier moves on a pseudo-random hop sequence known to both ends; narrow at any instant, wide over time. Processing gain = number of channels.
  • Slow hopping: several symbols per hop (GSM 1,248, Bluetooth 625). Fast hopping: several hops per symbol; diversity inside a symbol, expensive synthesiser.
  • Benefits: interference and jamming resistance, frequency diversity, sharing with other hoppers. Costs: synchronisation, settling time, collisions.
  • GSM (TS 45.002, 6.2.3): MA (up to 64 carriers), HSN (0 to 63; 0 = cyclic, MAI = (FN + MAIO) mod N), MAIO (0 to N-1). Same cell: same HSN, different MAIO, so no collisions; different cells: different HSN, random collisions.
  • Program: 8 carriers used 2,455 to 2,558 times in 20,000 frames; MAIO 0 and 3 collided 0 times in 2,000; a jammer on 1 of 8 channels loses 12.29 per cent of frames outright at one hop a frame, but with 26 hops a frame no frame is lost and 96.91 per cent are merely damaged, the same 12.5 per cent of bits either way.
  • DSSS against FHSS: wide always against narrow and moving; correlator against synthesiser; averaging against avoidance.

Test yourself

1. What is frequency hopping spread spectrum, and how does it spread a signal? The transmitter changes its carrier frequency repeatedly according to a hop sequence known to the receiver, so that although the signal is narrowband at any instant, over many hops it occupies a wide band. The spreading is therefore in time rather than in instantaneous bandwidth, and the processing gain is the number of channels hopped over. Both ends must be synchronised to the sequence and to the hop timing.

munotes.in641

Frequency Hopping Spread Spectrum

2. Distinguish slow and fast frequency hopping. In slow hopping a hop lasts longer than a symbol, so several symbols, often hundreds, are sent on each carrier; GSM sends 1,248 symbols per hop and Bluetooth about 625. In fast hopping a single symbol is spread over several hops, so every symbol samples several channels and gains diversity within itself, but the frequency synthesiser must settle in a fraction of a symbol period, which is costly, so fast hopping is used mainly in military systems.

3. Explain GSM's hopping parameters and how two mobiles in one cell avoid each other. A mobile is given a mobile allocation, the set of up to 64 carriers it may use; a hopping sequence number from 0 to 63, which selects the pseudo-random generator, with 0 meaning cyclic hopping; and a mobile allocation index offset from 0 to N-1. For each frame the algorithm computes an index into the allocation from the frame number and the hopping sequence number and then adds the offset modulo N. Mobiles in the same cell share the hopping sequence number and are given different offsets, so their sequences are the same sequence shifted and they never occupy the same carrier in the same frame; mobiles in different cells use different hopping sequence numbers, so their sequences are unrelated and collisions occur only at random.

4. Does frequency hopping reduce the effect of a jammer? Explain with the chapter's measurements. It does not reduce the total interference: with one channel of eight jammed, about one eighth, 12.5 per cent, of the transmitted bits are hit whatever the hop rate. What it changes is the distribution. With one hop per frame, 12.29 per cent of frames land on the jammed channel and are destroyed entirely. With 26 hops per frame, no frame is entirely lost, although 96.91 per cent contain some damaged bits. Scattered bit errors within a frame can be corrected by an error-correcting code, while a wholly destroyed frame cannot, so hopping converts unrecoverable loss into recoverable loss.

5. Compare DSSS and FHSS. DSSS multiplies each symbol by a fast chipping code, so the signal is wide at every instant; the receiver correlates with the code, which averages narrowband interference down by the processing gain, and multipath echoes past a chip can be separated or combined. FHSS keeps the signal narrow at each instant and moves it between channels on a pseudo-random sequence, so an interferer or a fade costs only the hops that land on it. DSSS needs chip-level synchronisation and a correlator; FHSS needs hop timing and a fast synthesiser. IEEE 802.15.4, W-CDMA and GPS use DSSS; GSM, Bluetooth and the TSCH mode of 802.15.4e hop.

munotes.in642

Frequency Hopping Spread Spectrum

6. Why must a hop sequence use every channel about equally? Because the diversity that hopping provides comes from sampling many channels: if some were visited more often than others, a fade or an interferer on a favoured channel would cost more than its share, and the benefit of hopping would fall. The chapter's run of GSM's generator over 20,000 frames used each of eight carriers between 2,455 and 2,558 times against an ideal 2,500, which is even to within about 2 per cent.

Contents This chapter on its own page

munotes.in643

Chapter Eighty-Five

Cellular Systems: Cells, Clusters and Frequency Reuse

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: Cellular systems"

In one line

A cellular system gives up range on purpose: instead of one transmitter covering a city on all its channels, it uses many low-power base stations, each covering a small cell with a fraction of the channels, and reuses those channels in every cell far enough away that the interference is tolerable, so the capacity of the city depends on how small the cells are and not on how much spectrum was allocated.

In the wording a student can write in an examination: a cellular system divides the service area into cells, each served by a base station using a subset of the available channels at low power. A group of cells that together use all the channels once is a cluster, of size N; the pattern repeats, so each channel is reused in every cluster. On a hexagonal grid the only possible cluster sizes are N = i squared + i j + j squared for non-negative integers i and j, giving 1, 3, 4, 7, 9, 12, 13, 19 and so on (5, 8, 10, 11 and 14 are impossible). The reuse distance, the distance between two cells using the same channels, is D = R times the square root of 3N, where R is the cell radius; the ratio D / R is the co-channel reuse ratio. With six co-channel interferers in the first ring and a path-loss exponent n, the signal-to-interference ratio at the cell edge is about (3N) to the power n/2, divided by 6.

A small N reuses channels more often, giving more channels per cell and more capacity, but less distance between co-channel cells and so more interference; a large N is the reverse. Capacity comes from small cells: halving the cell radius quarters the area and roughly quadruples the number of cells, and so the number of simultaneous calls, with no extra spectrum. The costs are more base stations, more handovers as users cross cells, and more interference to manage.

Why cells

A single transmitter covering a city can use each channel once. Every conversation needs a channel, so the number of simultaneous calls equals the number of channels, and no amount of power changes that. In the 1940s that was the state of mobile radio: a city had a few dozen channels and a waiting list.

The cellular idea inverts it. Cover the city with many small cells, each with a low-power base station; give each cell only a fraction of the channels; and reuse the same channels in cells far enough apart. Capacity is then the number of cells times the channels per cell, and the number of cells is decided by how small they are made. The spectrum becomes almost irrelevant to capacity; the number of sites becomes everything.

munotes.in644

Cellular Systems: Cells, Clusters and Frequency Reuse

This is space division multiplexing ([Multiplexing: Space, Frequency, Time and Code]) applied to a whole city, and it rests on the very property that makes radio difficult: signals fade with distance, so what is loud here is negligible there.

The hexagon

Real coverage is a ragged blob, decided by terrain and buildings, but planning needs a shape that tiles the plane. Only three regular polygons do: the triangle, the square and the hexagon.

The hexagon wins because it wastes the least. Given a base station whose usable reach is a circle, the hexagon inscribed in that circle covers 83 per cent of it, against 64 per cent for the square and 41 per cent for the triangle, so fewer cells cover an area and neighbouring cells overlap least. It is also the closest regular tiling to a circle, and it gives every cell six equidistant neighbours, which makes the reuse geometry clean.

The hexagon is a model, not a claim about coverage; every real plan is drawn from measurements and adjusted after the site is built.

Clusters and reuse

A cluster is the group of N cells that share out all the channels; the pattern of clusters then repeats over the plane. If there are S channels in total, each cell gets about S / N.

Not every N is possible. To tile the plane with a repeating pattern of hexagons, one must step i cells in one direction and j cells in another at 60 degrees, and the cluster size that results is

N = i squared + i j + j squared

which gives 1, 3, 4, 7, 9, 12, 13, 16, 19, 21, 25, 27 and so on. 5, 8, 10, 11 and 14 cannot occur, which is a purely geometric fact and a favourite examination question.

The distance between the centres of two cells using the same channels, the reuse distance, follows from the same geometry:

D = R times the square root of 3N

so D / R is 3.00 for N = 3, 4.58 for N = 7 and 7.55 for N = 19.

Co-channel interference

Reuse buys capacity and pays in interference. A cell has six co-channel cells in its first ring, all at roughly the reuse distance D. A mobile at the edge of its own cell, at distance R from its base station, receives the wanted signal from R and the six interferers from about D, so with a path-loss exponent n:

S / I = (D / R) to the power n, divided by 6 = (3N) to the power n/2, divided by 6

munotes.in645

Cellular Systems: Cells, Clusters and Frequency Reuse

Two things follow. The ratio does not depend on the cell radius, so cells can be made as small as one likes without changing the interference, which is what makes cell splitting work ([Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing]). And it improves with the path-loss exponent: interference falls off faster than the wanted signal in a cluttered environment, so a city tolerates a smaller cluster than open country.

The program computes the table. An analogue system needed roughly 18 dB for acceptable speech, which is why N = 7 was the classic choice. Digital systems tolerate far less, because error-correcting coding, interleaving and equalisation recover what interference damages, so GSM commonly used clusters of 4 or 3, and a CDMA system reuses every channel in every cell, N = 1, treating other cells' signals as noise ([The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover]).

Cells, clusters and reuse, computed

The program lists the possible cluster sizes with their reuse distances; computes the signal-to-interference ratio for six cluster sizes and four path-loss exponents; counts the calls a 100 square kilometre city can carry as its cells shrink; counts the handovers a car then suffers; and compares the three tilings.

# Cells, clusters and frequency reuse: the hexagon, the cluster sizes the
# geometry allows, the reuse distance, the interference it leaves, and the
# capacity it buys.
import math

# 1. Which cluster sizes are possible. A hexagonal grid allows only
#    N = i^2 + i j + j^2 for non-negative integers i and j.
print("Cluster sizes a hexagonal grid allows, N = i*i + i*j + j*j:")
found = {}
for i in range(6):
    for j in range(6):
        n = i * i + i * j + j * j
        if n and n not in found:
            found[n] = (i, j)
for n in sorted(found)[:12]:
    i, j = found[n]
    print("  N = %2d  (i = %d, j = %d)   reuse distance D = R sqrt(3N) = %.2f R" % (n, i, j, math.sqrt(3 * n)))
print("  N = 5, 8, 10, 11 and 14 are impossible: no i and j give them.")

# 2. Co-channel interference. A cell has six co-channel neighbours in the first
#    ring, at about D. With a path-loss exponent n, the ratio of the wanted
#    signal (at the cell edge, distance R) to the sum of the six is:
#       S/I = (D/R)^n / 6 = (3N)^(n/2) / 6
print("\nSignal to co-channel interference at the cell edge, six interferers:")
print("   N    D/R    n = 2.0   n = 3.0   n = 3.5   n = 4.0")
for n in (3, 4, 7, 9, 12, 19):
    dr = math.sqrt(3 * n)
    row = "".join("%10.1f" % (10 * math.log10(dr ** e / 6)) for e in (2.0, 3.0, 3.5, 4.0))
    print("  %2d %6.2f %s   (dB)" % (n, dr, row))
print("  an analogue system wanted about 18 dB: N = 7 gives %.1f dB at n = 4 and %.1f dB at n = 3.5,"
      % (10 * math.log10(math.sqrt(21) ** 4 / 6), 10 * math.log10(math.sqrt(21) ** 3.5 / 6)))
print("  which is why 7 was the classic cluster and why smaller clusters waited for digital systems.")

# 3. Capacity. The same 350 channels, a city of 100 square km, cells of radius
#    R: how many calls at once, and how many sites it takes.
print("\n350 channels over 100 square km, hexagons of radius R:")
print("   R      cells   channels a cell (N = 7)   calls at once   sites needed")
for r in (8.0, 4.0, 2.0, 1.0, 0.5):
    area = 2.598 * r * r
    cells = max(1, round(100 / area))
    per = 350 // 7
    print("  %4.1f km %7d %24d %15d %14d" % (r, cells, per, cells * per, cells))

# 4. The trade the other way: smaller cells mean more handovers. A car at
#    50 km/h crossing cells of radius R.
print("\nA car at 50 km/h, and how often it changes cell:")
for r in (8.0, 4.0, 2.0, 1.0, 0.5, 0.25):
    minutes = (2 * r) / 50 * 60
    print("  R = %4.2f km: a cell is crossed in %5.2f minutes, so about %4.1f handovers an hour"
          % (r, minutes, 60 / minutes))

# 5. Why hexagons. Three shapes tile the plane; compare how well each covers a
#    circle of the same reach.
print("\nWhy hexagons: tiling the plane with cells of circumradius 1")
for name, sides in (("triangle", 3), ("square", 4), ("hexagon", 6)):
    area = 0.5 * sides * math.sin(2 * math.pi / sides)          # regular polygon, circumradius 1
    print("  %-9s area %.3f of the circle's %.3f, so %.0f%% of the reach is used"
          % (name, area, math.pi, 100 * area / math.pi))
print("  the hexagon wastes the least, and needs the fewest cells to cover an area.")
munotes.in646

Cellular Systems: Cells, Clusters and Frequency Reuse

Cluster sizes a hexagonal grid allows, N = i*i + i*j + j*j:
  N =  1  (i = 0, j = 1)   reuse distance D = R sqrt(3N) = 1.73 R
  N =  3  (i = 1, j = 1)   reuse distance D = R sqrt(3N) = 3.00 R
  N =  4  (i = 0, j = 2)   reuse distance D = R sqrt(3N) = 3.46 R
  N =  7  (i = 1, j = 2)   reuse distance D = R sqrt(3N) = 4.58 R
  N =  9  (i = 0, j = 3)   reuse distance D = R sqrt(3N) = 5.20 R
  N = 12  (i = 2, j = 2)   reuse distance D = R sqrt(3N) = 6.00 R
  N = 13  (i = 1, j = 3)   reuse distance D = R sqrt(3N) = 6.24 R
  N = 16  (i = 0, j = 4)   reuse distance D = R sqrt(3N) = 6.93 R
  N = 19  (i = 2, j = 3)   reuse distance D = R sqrt(3N) = 7.55 R
  N = 21  (i = 1, j = 4)   reuse distance D = R sqrt(3N) = 7.94 R
  N = 25  (i = 0, j = 5)   reuse distance D = R sqrt(3N) = 8.66 R
  N = 27  (i = 3, j = 3)   reuse distance D = R sqrt(3N) = 9.00 R
  N = 5, 8, 10, 11 and 14 are impossible: no i and j give them.

Signal to co-channel interference at the cell edge, six interferers:
   N    D/R    n = 2.0   n = 3.0   n = 3.5   n = 4.0
   3   3.00        1.8       6.5       8.9      11.3   (dB)
   4   3.46        3.0       8.4      11.1      13.8   (dB)
   7   4.58        5.4      12.1      15.4      18.7   (dB)
   9   5.20        6.5      13.7      17.3      20.8   (dB)
  12   6.00        7.8      15.6      19.5      23.3   (dB)
  19   7.55        9.8      18.6      22.9      27.3   (dB)
  an analogue system wanted about 18 dB: N = 7 gives 18.7 dB at n = 4 and 15.4 dB at n = 3.5,
  which is why 7 was the classic cluster and why smaller clusters waited for digital systems.

350 channels over 100 square km, hexagons of radius R:
   R      cells   channels a cell (N = 7)   calls at once   sites needed
   8.0 km       1                       50              50              1
   4.0 km       2                       50             100              2
   2.0 km      10                       50             500             10
   1.0 km      38                       50            1900             38
   0.5 km     154                       50            7700            154

A car at 50 km/h, and how often it changes cell:
  R = 8.00 km: a cell is crossed in 19.20 minutes, so about  3.1 handovers an hour
  R = 4.00 km: a cell is crossed in  9.60 minutes, so about  6.2 handovers an hour
  R = 2.00 km: a cell is crossed in  4.80 minutes, so about 12.5 handovers an hour
  R = 1.00 km: a cell is crossed in  2.40 minutes, so about 25.0 handovers an hour
  R = 0.50 km: a cell is crossed in  1.20 minutes, so about 50.0 handovers an hour
  R = 0.25 km: a cell is crossed in  0.60 minutes, so about 100.0 handovers an hour

Why hexagons: tiling the plane with cells of circumradius 1
  triangle  area 1.299 of the circle's 3.142, so 41% of the reach is used
  square    area 2.000 of the circle's 3.142, so 64% of the reach is used
  hexagon   area 2.598 of the circle's 3.142, so 83% of the reach is used
  the hexagon wastes the least, and needs the fewest cells to cover an area.
munotes.in647

Cellular Systems: Cells, Clusters and Frequency Reuse

Which clusters exist. 1, 3, 4, 7, 9, 12, 13, 16, 19, 21, 25 and 27, each with its (i, j). The gaps are the point: no arrangement of hexagons gives a cluster of 5, 8, 10, 11 or 14, so a planner choosing a reuse pattern chooses from this list and no other.

munotes.in648

Cellular Systems: Cells, Clusters and Frequency Reuse

Interference. At a path-loss exponent of 4, N = 3 gives 11.3 dB, N = 7 gives 18.7 dB and N = 19 gives 27.3 dB. At a gentler exponent of 3, the same clusters give 6.5, 12.1 and 18.6 dB: the environment matters as much as the plan. The classic analogue requirement of about 18 dB is met by N = 7 at n = 4 and not at n = 3.5 (15.4 dB), which is why real networks measured their propagation before choosing.

Capacity. With 350 channels and N = 7, every cell gets 50 channels whatever its size. One 8 km cell covering the city carries 50 calls. Cells of 2 km give 10 cells and 500 calls; 1 km gives 38 cells and 1,900 calls; 500 m gives 154 cells and 7,700 calls, a 154-fold increase from the same spectrum. The only thing that changed is the number of base stations, which is also 154.

The bill. A car at 50 km/h crosses an 8 km cell in 19.2 minutes, about 3 handovers an hour. In 500 m cells it crosses one every 1.2 minutes, 50 handovers an hour, and in 250 m cells, 100 an hour. Every one is a signalling exchange that can fail and drop the call ([Handover in GSM]), so the capacity of very small cells is paid for in mobility management, not only in sites.

The hexagon. A hexagon inscribed in the reach circle covers 83 per cent of its area, against 64 per cent for a square and 41 per cent for a triangle. That is the whole argument for the shape.

Distinctions

One large transmitterA cellular system
PowerHighLow per base station
Channels used at onceAll, onceAll, once per cluster, in every cluster
CapacityThe number of channelsCells times channels per cell
To add capacityNothing availableSplit the cells
CostOne siteMany sites, handovers, planning
Small cluster (N = 3)Large cluster (N = 19)
Channels per cellMany (S / 3)Few (S / 19)
Reuse distance3.00 R7.55 R
S / I at n = 411.3 dB27.3 dB
SuitsDigital systems with codingAnalogue systems, or noisy environments
munotes.in649

Cellular Systems: Cells, Clusters and Frequency Reuse

Cell radius RCluster size N
DecidesHow many cells fit in the areaHow many channels each cell gets
Affects interferenceNot at all: S / I is independent of RDirectly
Affects capacityQuadratically: halving R quadruples the cellsInversely
Affects handoversDirectly: smaller cells, more handoversNot at all

What it does not mean

Cells are not hexagons. The hexagon is a planning model; real coverage follows terrain, buildings and antenna patterns.

More spectrum is not the main route to capacity. Smaller cells are: the program's city gained 154 times the capacity with no extra spectrum.

A smaller cluster is not simply better. It gives each cell more channels and each mobile more interference; how far one can go depends on how much interference the air interface tolerates.

Cell radius does not change the interference ratio. S / I depends on D / R, and both scale together, which is exactly why splitting cells works.

N = 1 is not a contradiction. A CDMA system reuses every channel in every cell and lives with the interference, because its receivers are built to.

Quick revision

  • Cell: a small area served by one low-power base station with a subset of the channels. Cluster: N cells using all the channels once; the pattern repeats.
  • Possible cluster sizes: N = i squared + i j + j squared: 1, 3, 4, 7, 9, 12, 13, 16, 19, 21, 25, 27; 5, 8, 10, 11, 14 are impossible.
  • Reuse distance D = R times the square root of 3N: 3.00 R (N = 3), 4.58 R (N = 7), 7.55 R (N = 19).
  • Co-channel interference: six first-ring interferers, S / I = (3N) to the power n/2, divided by 6; independent of R. At n = 4: 11.3 / 18.7 / 27.3 dB for N = 3 / 7 / 19. Analogue wanted about 18 dB, hence N = 7; digital systems use 4, 3, or 1 (CDMA).
  • Capacity = cells x channels per cell. Program: 350 channels, N = 7, 100 square km: 50 calls with 8 km cells, 7,700 with 500 m cells, from 154 sites.
  • Costs: sites, and handovers: 3 an hour at R = 8 km against 100 an hour at R = 250 m for a car at 50 km/h.
  • Hexagon: covers 83 per cent of the reach circle, against 64 (square) and 41 (triangle).

Test yourself

1. Why does a cellular system have more capacity than a single powerful transmitter? Because a single transmitter can use each channel only once in the whole area, so the number of simultaneous calls equals the number of channels. A cellular system covers the area with many low-power cells, each using a fraction of the channels, and reuses the same channels in cells far enough apart for the interference to be acceptable. The capacity is then the number of cells multiplied by the channels per cell, and the number of cells can be increased by making them smaller, without any additional spectrum.

munotes.in650

Cellular Systems: Cells, Clusters and Frequency Reuse

2. Why is the hexagon used to model a cell? Because a planning model must tile the plane without gaps or overlaps, and only the triangle, the square and the hexagon do so. Of the three, the hexagon is closest to the circle that a base station's reach actually describes: inscribed in that circle it covers 83 per cent of its area, against 64 per cent for the square and 41 per cent for the triangle, so fewer cells cover the area and neighbours overlap least. It also gives every cell six equidistant neighbours, which makes the reuse geometry simple.

3. Which cluster sizes are possible, and why? Only those of the form N = i squared + i j + j squared for non-negative integers i and j, because a repeating hexagonal pattern is generated by moving i cells along one axis and j cells along another at 60 degrees. That gives 1, 3, 4, 7, 9, 12, 13, 16, 19, 21, 25, 27 and so on; the values 5, 8, 10, 11 and 14 cannot be produced by any pair, so no such cluster can be laid out on a hexagonal grid.

4. Derive the reuse distance and the signal-to-interference ratio. For a cluster of size N on a hexagonal grid with cell radius R, the distance between the centres of two co-channel cells is D = R times the square root of 3N. A mobile at the edge of its cell is at distance R from its own base station and at roughly D from each of the six co-channel base stations in the first ring, so with a path-loss exponent n the ratio of wanted to interfering power is (D / R) to the power n divided by 6, which is (3N) to the power n/2 divided by 6. It does not depend on R, so the same plan works at any cell size.

5. What does the choice of cluster size trade? A small cluster gives each cell a larger share of the channels, so more capacity per cell, but places co-channel cells closer together, so the signal-to-interference ratio is worse: at a path-loss exponent of 4, N = 3 gives 11.3 dB while N = 7 gives 18.7 dB and N = 19 gives 27.3 dB. Analogue systems needed about 18 dB and so used N = 7; digital systems, whose coding and equalisation tolerate more interference, use 4 or 3, and CDMA systems use N = 1.

munotes.in651

Cellular Systems: Cells, Clusters and Frequency Reuse

6. What does a network pay for smaller cells? More base stations, one per cell, with their sites, backhaul and maintenance: the chapter's city needed 154 of them for 500 m cells. More handovers, because users cross cells more often: a car at 50 km/h needs about 3 handovers an hour in 8 km cells but 50 an hour in 500 m cells and 100 an hour in 250 m cells, and each is a signalling exchange that can fail. And more careful interference planning, since neighbouring cells are closer together in every sense.

Contents This chapter on its own page

munotes.in652

Chapter Eighty-Six

Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: Cellular systems"

In one line

A congested cell can be given more channels (by allocating them dynamically instead of fixing them), or split into smaller cells, or sectorised, and each answer has a different bill: splitting costs sites, dynamic allocation costs signalling and planning, sectorisation actually loses trunking efficiency and pays only because a directional antenna hears fewer interferers and so allows a tighter reuse pattern, while in a CDMA system the cell has no fixed size at all and simply shrinks as it fills.

In the wording a student can write in an examination: the traffic a cell carries is measured in erlangs, one erlang being one channel occupied continuously, and the relation between channels, offered traffic and the blocking probability is given by the Erlang B formula. Because of trunking, a large pool of channels carries much more traffic per channel than a small one. Channels are given to cells by fixed channel allocation (FCA), where each cell has a permanent set; channel borrowing, where a congested cell borrows from a neighbour that is not using them; or dynamic channel allocation (DCA), where channels are assigned from a common pool as calls arrive, subject to interference constraints. FCA is simple and predictable and wastes capacity when demand moves; DCA follows demand at the cost of central control and signalling.

To add capacity to a busy area: cell splitting replaces one cell with several smaller ones, multiplying capacity by the number of new cells at the cost of new sites and more handovers; sectorisation divides a site into three or six sectors with directional antennas, which reduces the number of co-channel interferers a receiver sees and so allows a smaller cluster, though it divides each cell's channels and therefore loses trunking efficiency; umbrella cells (a large cell overlaying small ones) carry fast-moving users who would otherwise hand over constantly. In a CDMA system, where every user adds to the interference floor, the coverage radius falls as the load rises, which is cell breathing: coverage and capacity are the same resource.

How much traffic a cell carries: Erlang B

A cell with c channels does not carry c calls: calls arrive at random, and the question is how much traffic can be offered before too many find every channel busy. The standard model is Erlang B, which assumes calls arrive at random, last for random times, and are lost if no channel is free (there is no queue). It gives the blocking probability for c channels and an offered traffic of A erlangs, and it is computed by a simple recurrence that the program uses.

The important consequence is trunking efficiency. At 2 per cent blocking, one channel carries 0.02 erlangs, ten channels carry 5.08 (0.51 each), and a hundred carry 87.97 (0.88 each). Pooling channels is worth a great deal, and any scheme that chops a pool into small pieces pays for it. That single fact decides much of what follows.

munotes.in653

Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing

Allocating channels to cells

Fixed channel allocation. Each cell is given a permanent set, planned so that co-channel cells are at least the reuse distance apart. It is simple, needs no signalling, and is predictable. Its weakness is that demand moves: a business district is busy at eleven in the morning and empty at nine at night, while the channels stay where they were put.

Channel borrowing. A congested cell borrows a channel from a neighbour that is not using it, provided the interference constraints still hold. It is a small patch on fixed allocation, and it complicates the reuse plan: the lent channel is now unavailable in several cells.

Dynamic channel allocation. Channels live in a common pool, and one is assigned to each call when it arrives, chosen so that no co-channel cell within the reuse distance is using it. Demand is followed wherever it goes. The costs are a central controller or a distributed protocol, measurement and signalling for every call, and a plan that can no longer be drawn on paper.

The program measures what the difference is worth in a network where one cell at a time is busy.

Cell splitting

The brute-force answer: replace one cell of radius R with four of radius R/2, each with its own base station and its own full set of channels. The area of each is a quarter, so four cover the same ground, and the capacity is four times.

Nothing else changes: the interference ratio S / I depends on D / R, and both halve together ([Cellular Systems: Cells, Clusters and Frequency Reuse]), so the same cluster pattern still works. What changes is the bill: four sites instead of one, four times the backhaul and maintenance, and four times the handovers, since a moving user crosses cells twice as often in each direction.

In practice splitting is done where the traffic is, not uniformly, so a real network has 20 km cells in the countryside and 200 m cells in a shopping street, with overlaid large cells to catch the traffic the small ones cannot.

Sectorisation, and what it really buys

A site is sectorised by replacing its omnidirectional antenna with three antennas of 120 degrees, or six of 60, each with its own set of channels ([Directional Antennas, Sectorisation, Diversity and Spatial Reuse]).

The naive argument, that this triples the site's capacity, is wrong, and the program shows why. If the cell's 50 channels are divided into three sets of 16, each sector is a small trunk group: three sets of 16 carry 29.5 erlangs in all, against 40.3 erlangs for 50 channels in one pool. Sectorisation, taken by itself, loses capacity.

munotes.in654

Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing

What sectorisation actually does is reduce interference. A directional antenna facing one way hears only the co-channel interferers in that direction: with three sectors, about two of the six first-ring interferers instead of all six, and with six sectors, about one. The signal-to-interference ratio therefore improves by roughly a factor of three or six, which is enough to move to a smaller cluster. With N = 3 instead of N = 7, every site is given 116 channels instead of 50, and even after dividing them among the sectors the site carries 87.5 erlangs against 40.3.

So the chain is: directional antennas, fewer interferers, smaller cluster, more channels per site, more traffic, and the trunking loss is a tax paid along the way. That is why six sectors, which allow no smaller a cluster than three in this model, carry less than three sectors: the extra division costs more than the extra directivity gains.

Umbrella cells

Small cells are efficient for slow users and miserable for fast ones: a car crossing 500 m cells hands over every 1.2 minutes ([Cellular Systems: Cells, Clusters and Frequency Reuse] counted it). An umbrella cell, a large cell from a high site overlaying a carpet of small ones, solves this by holding the fast movers: a user whose measurements change quickly is handed up to the umbrella and stays there until it slows down. The umbrella carries few users but the most troublesome ones, and the small cells below carry the volume.

Cell breathing

In a CDMA system there is no fixed set of channels: every user transmits over the whole band, and each one raises the interference floor for the others ([Spread Spectrum and Direct Sequence]). A user at the cell edge is heard just above that floor, so as the cell fills, the floor rises and the edge moves inward. This is cell breathing, and it means a CDMA cell's radius is not a property of the base station but of the load, changing minute by minute.

Two consequences follow. Coverage and capacity are the same resource: an operator cannot plan them separately, and a cell that is planned to reach 2 km when empty reaches less when busy. And a neighbouring cell's load affects this cell's coverage, since interference crosses boundaries, so the whole network breathes together. The program computes the shrinkage for a cell of processing gain 128.

munotes.in655

Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing

Capacity, computed

The program tabulates Erlang B; compares fixed and dynamic allocation over 2,000 busy hours in which one of seven cells is busy at a time; prices splitting and sectorisation, both naively and through the interference argument; and computes a CDMA cell's radius as its load rises.

# Making a cellular network carry more: how channels are allocated, what cell
# splitting and sectorisation buy, and how a CDMA cell breathes.
import math
import random

# 1. Erlang B: how many calls a cell with c channels can carry at a given
#    blocking probability. B(c, A) = (A^c / c!) / sum(A^k / k!), computed by
#    the recurrence B(0, A) = 1, B(c, A) = A B(c-1, A) / (c + A B(c-1, A)).
def erlang_b(c, a):
    b = 1.0
    for k in range(1, c + 1):
        b = a * b / (k + a * b)
    return b

def capacity(c, target=0.02):
    lo, hi = 0.0, 10000.0
    for _ in range(80):
        mid = (lo + hi) / 2
        if erlang_b(c, mid) < target:
            lo = mid
        else:
            hi = mid
    return lo

print("Erlang B: traffic a cell can carry at 2 per cent blocking, and per channel:")
print("  channels   traffic (erlangs)   per channel   calls in the busy hour (3 min each)")
for c in (1, 5, 10, 20, 50, 100):
    a = capacity(c)
    print("  %8d %18.2f %13.2f %36.0f" % (c, a, a / c, a * 60 / 3))
print("  a few channels are used inefficiently; trunking makes big pools worth more per channel.")

# 2. Fixed against dynamic allocation. Seven cells share 350 channels; demand
#    moves about during the day. Fixed gives each cell 50; dynamic lends
#    channels to whoever needs them, up to what interference allows.
rnd = random.Random(86)
def day(fixed=True, hours=2000, per_cell=50, total=350):
    blocked, offered = 0, 0
    for _ in range(hours):
        # one cell is busy (a station, an office), the rest are quiet
        hot = rnd.randrange(7)
        demand = [rnd.gauss(70 if i == hot else 25, 8) for i in range(7)]
        for i, d in enumerate(demand):
            d = max(0.0, d)
            offered += d
            if fixed:
                blocked += d * erlang_b(per_cell, d)
            else:
                share = min(int(total * d / sum(max(0.0, x) for x in demand)), 120)
                blocked += d * erlang_b(max(1, share), d)
    return 100 * blocked / offered

print("\nSeven cells, 350 channels, one cell busy at a time (2000 busy hours):")
print("  fixed allocation, 50 channels each: %.2f%% of calls blocked" % day(True))
print("  dynamic allocation by demand:       %.2f%% of calls blocked" % day(False))

# 3. Cell splitting and sectorisation, as ways to add capacity to a hot cell.
print("\nAdding capacity to one congested cell of radius R with 50 channels:")
base = capacity(50)
print("  as it stands:                    %5.1f erlangs from 1 site" % base)
print("  split into 4 cells of R/2:       %5.1f erlangs from 4 sites (each still 50 channels)" % (4 * base))
for k, gain in ((3, 4.77), (6, 7.78)):
    per = 50 // 1                          # each sector keeps a full set in this model
    print("  %d sectors of %3d degrees:        %5.1f erlangs from 1 site, antennas %.2f dB"
          % (k, 360 // k, k * capacity(50 // k) * (k / k), gain))
print("  (sectors split the cell's channels, so each sector has fewer and trunking suffers:")
print("   3 sectors of 16 channels carry %.1f erlangs in all, against %.1f for 50 in one cell.)"
      % (3 * capacity(16), base))

# sectorisation pays through interference, not through trunking: a directional
# antenna hears fewer of the six co-channel neighbours, so the cluster can
# shrink and every site is given more channels.
print("\nWhat sectorisation really buys: fewer interferers, so a smaller cluster.")
print("  sectors   interferers seen   cluster it allows   channels a site   traffic carried")
for k, interferers, cluster in ((1, 6, 7), (3, 2, 3), (6, 1, 3)):
    per_site = 350 // cluster
    per_sector = per_site // k
    print("  %7d %18d %19d %17d %18.1f erlangs"
          % (k, interferers, cluster, per_site, k * capacity(per_sector)))

# 4. Cell breathing: in a CDMA cell every user raises the noise floor, so as
#    the cell fills, its edge moves in. With a processing gain G, a required
#    Eb/N0 and k users, the tolerable path loss falls.
print("\nCell breathing in a CDMA cell (processing gain 128, Eb/N0 needed 6 dB):")
G, need = 128, 10 ** (6 / 10)
print("   users   interference rise   cell radius (relative, path-loss exponent 3.5)")
for k in (1, 8, 16, 24, 30, 34):
    rise = 1 + (k - 1) / G * need          # others' power raises the floor
    radius = rise ** (-1 / 3.5)
    print("  %6d %19.2f dB %30.2f" % (k, 10 * math.log10(rise), radius))
print("  the cell shrinks as it fills: coverage and capacity are the same resource.")
munotes.in656

Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing

Erlang B: traffic a cell can carry at 2 per cent blocking, and per channel:
  channels   traffic (erlangs)   per channel   calls in the busy hour (3 min each)
         1               0.02          0.02                                    0
         5               1.66          0.33                                   33
        10               5.08          0.51                                  102
        20              13.18          0.66                                  264
        50              40.26          0.81                                  805
       100              87.97          0.88                                 1759
  a few channels are used inefficiently; trunking makes big pools worth more per channel.

Seven cells, 350 channels, one cell busy at a time (2000 busy hours):
  fixed allocation, 50 channels each: 10.32% of calls blocked
  dynamic allocation by demand:       0.31% of calls blocked

Adding capacity to one congested cell of radius R with 50 channels:
  as it stands:                     40.3 erlangs from 1 site
  split into 4 cells of R/2:       161.0 erlangs from 4 sites (each still 50 channels)
  3 sectors of 120 degrees:         29.5 erlangs from 1 site, antennas 4.77 dB
  6 sectors of  60 degrees:         21.8 erlangs from 1 site, antennas 7.78 dB
  (sectors split the cell's channels, so each sector has fewer and trunking suffers:
   3 sectors of 16 channels carry 29.5 erlangs in all, against 40.3 for 50 in one cell.)

What sectorisation really buys: fewer interferers, so a smaller cluster.
  sectors   interferers seen   cluster it allows   channels a site   traffic carried
        1                  6                   7                50               40.3 erlangs
        3                  2                   3               116               87.5 erlangs
        6                  1                   3               116               74.0 erlangs

Cell breathing in a CDMA cell (processing gain 128, Eb/N0 needed 6 dB):
   users   interference rise   cell radius (relative, path-loss exponent 3.5)
       1                0.00 dB                           1.00
       8                0.86 dB                           0.95
      16                1.66 dB                           0.90
      24                2.34 dB                           0.86
      30                2.79 dB                           0.83
      34                3.07 dB                           0.82
  the cell shrinks as it fills: coverage and capacity are the same resource.
munotes.in657

Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing

Trunking. At 2 per cent blocking, 10 channels carry 5.08 erlangs, 0.51 per channel; 50 carry 40.26, 0.81 each; 100 carry 87.97, 0.88 each. Doubling a pool more than doubles the traffic it carries, which is why anything that fragments channels, including sectorisation, must justify itself.

Dynamic allocation. With one cell of seven busy at a time, fixed allocation blocked 10.32 per cent of calls, because the busy cell ran out of its 50 while its neighbours sat idle. Dynamic allocation, giving each cell channels in proportion to its demand, blocked 0.31 per cent, a thirty-fold improvement, with no new sites and no new spectrum. That is the prize, and the price is the signalling and interference checking that every assignment needs.

Splitting and sectorising. The congested cell carries 40.3 erlangs. Split into four cells of half the radius, it carries 161.0, exactly four times, from four sites. Sectorised naively, it carries less: 29.5 erlangs with three sectors, 21.8 with six, because trunking efficiency falls. Sectorised properly, that is with the smaller cluster the directional antennas allow, one site carries 87.5 erlangs with three sectors: more than twice the original, from the same site. Six sectors carry 74.0, less than three, because the cluster cannot shrink further in this model and the extra division costs trunking.

Breathing. With a processing gain of 128 and a required Eb/N0 of 6 dB, one user sets the reference. Sixteen users raise the interference floor by 1.66 dB and shrink the radius to 0.90 of its empty value; thirty-four users raise it by 3.07 dB and shrink the radius to 0.82. A CDMA operator planning for coverage must therefore plan for a load, and a cell that loses coverage when busy pushes its edge users to neighbours, which then breathe in as well.

munotes.in658

Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing

Distinctions

Fixed allocationBorrowingDynamic allocation
Channels belong toA cell, permanentlyA cell, lent temporarilyA common pool
Follows demandNoA littleYes
SignallingNoneSomePer call
PlanningDrawn onceDrawn, with exceptionsContinuous, by rule
Program10.32 per cent blockedNot modelled0.31 per cent blocked
Cell splittingSectorisation
What changesSmaller cells, more sitesDirectional antennas at the same site
CapacityMultiplied by the number of new cells (4 for half the radius)Through a smaller cluster: 40.3 to 87.5 erlangs
TrunkingUnchangedWorse: channels are divided among sectors
HandoversMore, between cellsMore, between sectors as well
CostNew sites, backhaul, planningNew antennas and feeders at one site
A fixed-channel cell (GSM)A CDMA cell
CapacityHard: the channels run outSoft: quality falls as users are added
RadiusFixed by power and propagationFalls as the load rises (breathing)
NeighboursInterfere on co-channels onlyInterfere always
PlanningCoverage and capacity separatelyTogether: they are the same resource

What it does not mean

More channels in a cell is not proportionally more traffic. Trunking means the first channels are nearly useless and later ones are nearly fully used.

Sectorisation does not multiply capacity by the number of sectors. By itself it reduces capacity; the gain comes from the tighter reuse the directional antennas permit.

Cell splitting does not change the interference ratio. S / I depends on D / R, which splitting leaves unchanged; that is precisely why it works.

Dynamic allocation is not free. Every assignment needs interference information and signalling, and the network can no longer be planned on paper.

Cell breathing is not a fault. It is what soft capacity looks like: a CDMA system degrades gracefully rather than blocking abruptly.

Quick revision

  • Erlang B: blocking from channels and offered traffic; trunking: at 2 per cent blocking, 0.51 erlangs per channel with 10 channels, 0.88 with 100.
  • Allocation: fixed (simple, wastes when demand moves), borrowing, dynamic (follows demand, costs signalling). Program: 10.32 per cent blocked against 0.31 per cent.
  • Cell splitting: R to R/2 gives 4 cells, 4 times the capacity, 4 sites, more handovers; S / I unchanged.
  • Sectorisation: 3 or 6 directional antennas; fewer interferers (about 2 of 6, or 1), so a smaller cluster (7 to 3); naive: 40.3 to 29.5 erlangs (trunking lost); with the smaller cluster: 40.3 to 87.5 erlangs. Six sectors gave 74.0, less than three.
  • Umbrella cells: a large overlay carries fast movers, cutting handovers.
  • Cell breathing (CDMA): each user raises the interference floor, so the radius falls; program, gain 128, 6 dB needed: 0.90 of the radius at 16 users, 0.82 at 34. Coverage and capacity are one resource.
munotes.in659

Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing

Test yourself

1. What is trunking efficiency, and why does it matter for cellular planning? Trunking efficiency is the fact that a large pool of channels carries much more traffic per channel than a small pool at the same blocking probability, because random peaks in demand average out across a large group. At 2 per cent blocking, ten channels carry 5.08 erlangs, 0.51 per channel, while a hundred carry 87.97, 0.88 per channel. It matters because any scheme that divides a cell's channels into smaller groups, such as sectorisation, loses capacity before it gains any, and must be justified by some other benefit.

2. Compare fixed, borrowing and dynamic channel allocation. In fixed allocation each cell has a permanent set of channels planned to respect the reuse distance: simple, predictable and needing no signalling, but wasteful when demand moves between cells. In borrowing, a congested cell temporarily takes a channel from a neighbour that is not using it, which helps a little but complicates the reuse plan, since the borrowed channel is blocked in several cells. In dynamic allocation all channels are in a pool and one is assigned per call, subject to interference constraints, so demand is followed wherever it goes, at the cost of measurement, signalling and central or distributed control. The chapter's model blocked 10.32 per cent of calls with fixed allocation and 0.31 per cent with dynamic.

3. Explain cell splitting and say what it costs. A congested cell of radius R is replaced by several smaller cells, typically four of radius R/2, each with its own base station and a full set of channels, so the area's capacity is multiplied by four. The signal-to-interference ratio is unchanged, because it depends on the ratio of the reuse distance to the cell radius and both shrink together, so the same cluster plan still works. The costs are the new sites, with their backhaul, power and maintenance, and the extra handovers, since users cross cells twice as often in each direction.

4. Why does sectorisation increase capacity, given that it divides a cell's channels? Dividing the channels actually reduces the traffic a site can carry, because of trunking: in the chapter's model, 50 channels in one pool carry 40.3 erlangs but three sectors of 16 carry only 29.5. The gain comes from interference. A directional antenna receives co-channel interference from only the interferers in its own direction, about two of the six first-ring cells with three sectors instead of all six, so the signal-to-interference ratio improves and the network can use a smaller cluster, for example 3 instead of 7. Each site is then given far more channels, 116 instead of 50, and carries 87.5 erlangs, more than twice the original.

munotes.in660

Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing

5. What is an umbrella cell and what problem does it solve? It is a large cell, usually from a high site, overlaying an area covered by many small cells. It solves the handover problem of small cells: a fast-moving user crossing 500 m cells would hand over every minute or so, and each handover can fail. Such users are handed up to the umbrella cell, which they cross slowly, and are handed back down when they slow; the small cells below then carry the bulk of the stationary and slow traffic.

6. What is cell breathing, and why does it happen only in CDMA systems? In a CDMA system every user in the cell, and in neighbouring cells, transmits over the whole band, so each one raises the interference floor against which the others are received. A user at the cell edge is only just above that floor, so as the load grows the floor rises and the usable radius shrinks; the chapter's model shrank a cell to 0.90 of its empty radius at 16 users and 0.82 at 34. It does not happen in a fixed-channel system such as GSM, where a call either gets a channel or is blocked and the cell's radius does not depend on how many calls are in progress.

Contents This chapter on its own page

munotes.in661

Chapter Eighty-Seven

GSM and Its Mobile Services

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: GSM: Mobile services"

In one line

GSM was built to replace a continent of incompatible analogue systems with one digital standard, and what it offers a subscriber is defined in three layers: bearer services, which carry bits between access points; teleservices, which are complete applications such as telephony, emergency calls and the short message service; and supplementary services, which modify a basic service, such as forwarding, barring, waiting and calling line identification.

In the wording a student can write in an examination: GSM (originally Groupe Speciale Mobile, later the Global System for Mobile Communications) is a digital cellular standard developed in Europe in the 1980s to replace incompatible national analogue systems, offering international roaming, better spectrum efficiency, security and digital services. Its services are of three kinds:

  1. Bearer services carry data between two access points, providing only the transmission capability. They are described by attributes in four groups: "Information transfer attributes, which characterize the network capabilities for transferring information from a user access point in a PLMN to a user access point in another network"; "Access attributes, which describe the means for accessing network functions or facilities as seen at the access point in the PLMN"; "Interworking attributes, which describe properties of the terminating network and its access point"; and "General attributes, which deal with the service in general." GSM's circuit bearer services ran at 300 bit/s up to 9.6 kbit/s, later 14.4 kbit/s.
  2. Teleservices provide complete communication including the terminal's functions. The standard lists telephony, emergency calls, the short message service (mobile terminated, mobile originated and cell broadcast), alternate speech and facsimile group 3, automatic facsimile group 3, the voice group call service and the voice broadcast service.
  3. Supplementary services modify or supplement a basic service and cannot be offered alone. The standard's table lists, among others, CLIP and CLIR (calling line identification presentation and restriction), CoLP and CoLR, call deflection (CD), call forwarding unconditional (CFU), on busy (CFB), on no reply (CFNRy) and on not reachable (CFNRc), call waiting (CW) and hold (HOLD), multi party (MPTY), closed user group (CUG), advice of charge (AoCI, AoCC), the barring services (BAOC, BOIC, BOIC-exHC, BAIC, BAIC-Roam, ACR), explicit call transfer (ECT), completion of calls to busy subscribers (CCBS) and calling name presentation (CNAP).

Where GSM came from

Before GSM, each European country ran its own analogue system, and a phone bought in one did not work in the next. The standard was built to fix that and more: to be digital, so that speech could be compressed, protected by coding and encrypted; to be spectrally efficient, so that more subscribers fitted into the same band; to allow international roaming on one subscription; to separate the subscription from the handset, which is the SIM; and to define services precisely enough that any manufacturer's equipment would interoperate.

munotes.in662

GSM and Its Mobile Services

It succeeded so completely that its architecture, its vocabulary and its service model are still the shape of mobile networks, and its specifications are still maintained by 3GPP alongside the later generations. What the following chapters describe as GSM is therefore both a historical system and the skeleton of what came after.

Bearer services

A bearer service offers only transmission: bits in at one access point, bits out at another, with stated characteristics and no opinion about what they mean. TS 22.002 defines them by attributes, "which are intended to be independent", grouped into the four categories quoted above: information transfer, access, interworking and general.

That structure is worth understanding for its own sake. Information transfer says what the network does with the bits: the rate, whether the transfer is transparent (a fixed delay and rate, with errors passed through) or non-transparent (a protocol that retransmits, giving fewer errors and a variable delay), and whether it is circuit or packet. Access says how the user reaches it. Interworking says what the far network is, which for GSM was usually the fixed telephone network or an ISDN. General covers the rest.

In practice GSM's circuit bearer services offered 300 bit/s to 9.6 kbit/s on one slot, later 14.4 kbit/s, and the program shows what that meant for anyone trying to move a file.

Teleservices

A teleservice is a complete service, including what the terminal does: not merely a pipe but a telephone call. TS 22.003 lists them, and three matter most.

Telephony is the reason the system exists: speech, digitised and compressed by the codec, protected by channel coding, carried in a traffic channel. The program follows a speech block from the codec to the air.

Emergency calls are a teleservice in their own right, not a variant of telephony, because they must work when telephony would not: dialled by a number the mobile itself recognises, permitted where a normal call would be barred, and routed by the network to the appropriate centre.

The short message service carries up to 140 octets of payload, which is 160 characters in GSM's 7-bit alphabet, and rides on the signalling channels rather than on a traffic channel, so it works while a call is in progress and needs no call setup at all. The standard distinguishes mobile terminated and mobile originated point-to-point messages and the cell broadcast service, which sends the same message to every mobile in an area. SMS was an afterthought in the design and became the system's second business.

munotes.in663

GSM and Its Mobile Services

The list closes with alternate speech and facsimile, automatic facsimile, and the voice group call and voice broadcast services, which serve dispatch-style users, the ground that [TETRA] occupies properly.

Supplementary services

A supplementary service modifies a basic service and cannot stand alone: forwarding modifies a call, it is not a call. TS 22.004 tabulates them with the operations each supports, and the operations are themselves the vocabulary of the subject:

  • Provision and withdrawal, by the operator;
  • Registration and erasure, by the subscriber, of a parameter such as a forwarding number;
  • Activation and deactivation;
  • Invocation, either "automatic invocation by the network as a result of a particular condition" or "user invocation, by means of a control procedure";
  • Interrogation, either a "status check" or a "data request".

Reading the standard's table with those columns explains the everyday behaviour of a phone. Call forwarding on busy is registered and erased by the subscriber, activated by registration, and invoked by the network when the condition occurs. Call waiting has nothing to register, is activated by the subscriber, and is invoked by the network. Multi party has no registration, no activation and no interrogation: it is invoked by the user during a call and that is all.

The families are: number identification (CLIP, CLIR, CoLP, CoLR, CNAP, call deflection), call offering (the four forwardings, explicit call transfer), call completion (call waiting, hold, completion of calls to busy subscribers), multi party, community of interest (closed user group), charging (advice of charge, information or charging), call restriction (the barring services) and user-to-user signalling.

GSM's services, computed

The program follows a full-rate speech block from the codec through channel coding to the bursts and asks how much of a slot it fills; times a 100 kB file over each circuit bearer service; counts the characters that fit in a short message's 140 octets; and lists what makes an emergency call different.

# GSM's services, in numbers: what a speech channel costs, what the data
# bearer services carried, and what SMS fits into.
import math

# 1. From speech to a burst. GSM's full-rate speech codec makes 260 bits every
#    20 ms; channel coding takes that to 456 bits, carried as 4 half-bursts of
#    114 bits, interleaved over 8 bursts.
print("A full-rate GSM speech channel, block by block:")
speech_bits, block_ms = 260, 20.0
coded = 456
print("  the codec produces %d bits every %.0f ms: %.2f kbit/s" % (speech_bits, block_ms, speech_bits / block_ms))
print("  channel coding takes each block to %d bits: %.2f kbit/s on the air" % (coded, coded / block_ms))
print("  that is %d bursts of 114 payload bits, one per TDMA frame of 4.615 ms, spanning %.2f ms"
      % (coded / 114, coded / 114 * 4.615))
slot_kbps = 114 / 4.615                       # payload bits a burst, one burst a frame
print("  one slot carries %.2f kbit/s of payload; coded speech needs %.2f, so it fills %.0f%% of the slot"
      % (slot_kbps, coded / block_ms, 100 * (coded / block_ms) / slot_kbps))
print("  of that, the speech itself is %.0f%%: the rest is the protection that makes it survive the air"
      % (100 * speech_bits / coded))

# 2. The bearer services: data rates GSM offered on one slot, and how long a
#    file takes at each.
print("\nCircuit bearer services on one slot, and a 100 kB file:")
for name, kbps in (("300 bit/s", 0.3), ("1.2 kbit/s", 1.2), ("2.4 kbit/s", 2.4),
                   ("4.8 kbit/s", 4.8), ("9.6 kbit/s", 9.6), ("14.4 kbit/s", 14.4)):
    seconds = 100 * 1024 * 8 / (kbps * 1000)
    print("  %-12s %8.1f seconds, that is %5.1f minutes" % (name, seconds, seconds / 60))
print("  HSCSD and GPRS answer by bundling slots; EDGE by changing the modulation.")

# 3. Short messages. A short message is 140 octets of payload; in the 7-bit
#    alphabet that is 160 characters.
print("\nA short message:")
octets = 140
for bits, name in ((7, "the 7-bit GSM alphabet"), (8, "8-bit data"), (16, "UCS-2, for other scripts")):
    print("  %-26s %3d characters in %d octets" % (name, octets * 8 // bits, octets))
print("  a message of 160 characters sent on a signalling channel of 782 bit/s takes %.2f s of air time"
      % (octets * 8 / 782))

# 4. Emergency calls: why they are a teleservice of their own.
print("\nWhy emergency calls are a separate teleservice:")
for rule in ("dialled with a single number, recognised by the mobile itself",
             "placed without a SIM where regulation requires it",
             "not barred by call barring or by a full cell's access control",
             "routed to the nearest centre by the network, not by the subscriber"):
    print("  - %s" % rule)
munotes.in664

GSM and Its Mobile Services

A full-rate GSM speech channel, block by block:
  the codec produces 260 bits every 20 ms: 13.00 kbit/s
  channel coding takes each block to 456 bits: 22.80 kbit/s on the air
  that is 4 bursts of 114 payload bits, one per TDMA frame of 4.615 ms, spanning 18.46 ms
  one slot carries 24.70 kbit/s of payload; coded speech needs 22.80, so it fills 92% of the slot
  of that, the speech itself is 57%: the rest is the protection that makes it survive the air

Circuit bearer services on one slot, and a 100 kB file:
  300 bit/s      2730.7 seconds, that is  45.5 minutes
  1.2 kbit/s      682.7 seconds, that is  11.4 minutes
  2.4 kbit/s      341.3 seconds, that is   5.7 minutes
  4.8 kbit/s      170.7 seconds, that is   2.8 minutes
  9.6 kbit/s       85.3 seconds, that is   1.4 minutes
  14.4 kbit/s      56.9 seconds, that is   0.9 minutes
  HSCSD and GPRS answer by bundling slots; EDGE by changing the modulation.

A short message:
  the 7-bit GSM alphabet     160 characters in 140 octets
  8-bit data                 140 characters in 140 octets
  UCS-2, for other scripts    70 characters in 140 octets
  a message of 160 characters sent on a signalling channel of 782 bit/s takes 1.43 s of air time

Why emergency calls are a separate teleservice:
  - dialled with a single number, recognised by the mobile itself
  - placed without a SIM where regulation requires it
  - not barred by call barring or by a full cell's access control
  - routed to the nearest centre by the network, not by the subscriber
munotes.in665

GSM and Its Mobile Services

Speech. The full-rate codec makes 260 bits every 20 ms, which is 13.00 kbit/s, and channel coding takes each block to 456 bits, 22.80 kbit/s. Those 456 bits are carried as four bursts of 114 payload bits, one per TDMA frame, spanning 18.46 ms, which is why the frame is 4.615 ms long ([The GSM Radio Interface: Carriers, the TDMA Frame and Bursts]). One slot carries 24.70 kbit/s of payload, so coded speech fills 92 per cent of it, and of what is transmitted only 57 per cent is speech: the rest is protection. That ratio is the price of a channel that must survive fading, interference and handover.

Data. A 100 kB file takes 45.5 minutes at 300 bit/s, 1.4 minutes at 9.6 kbit/s and 0.9 minutes at 14.4 kbit/s. For 1991 that was reasonable; by the time the web existed it was not, and the answers, bundling slots (HSCSD), packet switching (GPRS) and better modulation (EDGE), are [New Data Services: HSCSD, GPRS and EDGE].

Short messages. 140 octets hold 160 characters of the 7-bit alphabet, 140 of 8-bit data, or 70 in UCS-2, which is why a message in Devanagari or any non-Latin script costs twice as much of the allowance. At the SDCCH's rate a 160-character message is 1.43 seconds of air time, which is why an operator could give it away and then charge for it.

Distinctions

Bearer serviceTeleserviceSupplementary service
ProvidesTransmission between access pointsA complete service including the terminalA modification of a basic service
Stands aloneYesYesNo
Defined inTS 22.002TS 22.003TS 22.004
Examples9.6 kbit/s circuit data, transparent or notTelephony, emergency calls, SMS, fax, voice group callCLIP, CLIR, CFU, CFB, CW, MPTY, CUG, AoC, barring, ECT
Described byInformation transfer, access, interworking and general attributesThe service and its terminal behaviourProvision, registration, activation, invocation, interrogation
munotes.in666

GSM and Its Mobile Services

OperationWho does itExample
Provision / withdrawalThe operatorAdding forwarding to a subscription
Registration / erasureThe subscriberSetting the forwarding number
Activation / deactivationSubscriber, operator, or as a result of registrationTurning forwarding on
InvocationThe network (on a condition) or the userForwarding a call because the mobile is busy
InterrogationThe subscriberAsking whether forwarding is on, and to where
TelephonyEmergency callSMS
Carried onA traffic channelA traffic channelSignalling channels
Needs a call setupYesYes, but privilegedNo
Works during a callNot another oneYesYes
PayloadCoded speech, 22.8 kbit/sThe same140 octets, 160 characters
Barred by call barringYesNoSeparately

What it does not mean

A bearer service is not a teleservice with less. It is a different layer: the bearer carries bits, the teleservice defines what the terminals do with them.

Supplementary services are not features of the handset. They are network services; the handset only invokes and interrogates them.

SMS is not a small call. It travels on signalling channels, which is why it needs no call setup and works while a call is in progress.

GSM is not only the radio. The service definitions, the architecture and the signalling are as much a part of the standard as the air interface.

13 kbit/s of speech is not what the air carries. Channel coding takes it to 22.8 kbit/s, and only 57 per cent of what is transmitted is the speech itself.

Quick revision

  • GSM: European digital standard for roaming, efficiency, security and digital services; the SIM separates subscription from handset.
  • Bearer services (TS 22.002): transmission only; attributes in four groups, information transfer, access, interworking, general; transparent against non-transparent; 300 bit/s to 9.6 kbit/s, later 14.4.
  • Teleservices (TS 22.003): telephony, emergency calls, SMS (MT/PP, MO/PP, cell broadcast), alternate speech/fax G3, automatic fax G3, voice group call, voice broadcast.
  • Supplementary services (TS 22.004): CLIP, CLIR, CoLP, CoLR, CNAP, CD; CFU, CFB, CFNRy, CFNRc, ECT; CW, HOLD, CCBS; MPTY; CUG; AoCI, AoCC; BAOC, BOIC, BOIC-exHC, BAIC, BAIC-Roam, ACR; UUS. Operations: provision/withdrawal, registration/erasure, activation/deactivation, invocation (network or user), interrogation (status or data).
  • Program: speech 260 bits / 20 ms = 13.00 kbit/s, coded 456 bits = 22.80 kbit/s, 4 bursts of 114 bits, 92 per cent of a slot, only 57 per cent of it speech; a 100 kB file takes 45.5 minutes at 300 bit/s and 0.9 at 14.4 kbit/s; SMS 140 octets = 160 / 140 / 70 characters at 7 / 8 / 16 bits.
munotes.in667

GSM and Its Mobile Services

Test yourself

1. What were GSM's goals? To replace the incompatible analogue national systems of Europe with a single digital standard, giving international roaming on one subscription; to use the spectrum more efficiently and so serve more subscribers; to be digital, which allows speech compression, error protection, encryption and authentication; to separate the subscription from the handset through the SIM; and to specify services and interfaces precisely enough for equipment from different manufacturers to interoperate.

2. Distinguish bearer services, teleservices and supplementary services with examples. A bearer service provides only the transmission of information between access points, with stated attributes and no interpretation of the content, for example a 9.6 kbit/s transparent circuit data service. A teleservice provides complete communication including the functions of the terminals, for example telephony, emergency calls, the short message service or facsimile. A supplementary service modifies or supplements a basic service and cannot be offered alone, for example call forwarding, call barring, call waiting, multi party or calling line identification presentation.

3. What attributes describe a bearer service? Four groups, as TS 22.002 states: information transfer attributes, which characterise the network's capability to transfer information from a user access point in one network to a user access point in another; access attributes, which describe how network functions are reached at the access point; interworking attributes, which describe the terminating network and its access point; and general attributes, which deal with the service in general. They include the rate, whether the transfer is transparent or non-transparent, and whether it is circuit or packet switched.

4. List GSM's teleservices and say why emergency calls are separate from telephony. Telephony, emergency calls, the short message service in its mobile terminated, mobile originated and cell broadcast forms, alternate speech and facsimile group 3, automatic facsimile group 3, the voice group call service and the voice broadcast service. Emergency calls are a separate teleservice because they must behave differently from ordinary telephony: they are dialled by a number the mobile itself recognises, they are permitted where an ordinary call would be barred or the cell would refuse access, in many countries they can be made without a valid subscription, and the network routes them to the appropriate emergency centre rather than to a subscriber-chosen destination.

5. Name the operations that apply to a supplementary service, and illustrate with call forwarding on busy. Provision and withdrawal by the operator; registration and erasure by the subscriber; activation and deactivation; invocation, either automatically by the network when a condition occurs or by the user through a control procedure; and interrogation, either a status check or a data request. Call forwarding on busy is provided by the operator, registered by the subscriber together with the number to forward to, activated as a result of that registration or by the subscriber, invoked automatically by the network when a call arrives while the mobile is busy, and interrogated by the subscriber as a data request to find out where calls are being sent.

munotes.in668

GSM and Its Mobile Services

6. How much of a GSM traffic channel is actual speech? The full-rate codec produces 260 bits every 20 ms, which is 13 kbit/s, and channel coding expands each block to 456 bits, which is 22.8 kbit/s on the air. Those 456 bits are sent as four bursts of 114 payload bits, one in each of four TDMA frames. One slot's payload capacity is about 24.7 kbit/s, so the coded speech fills about 92 per cent of it, and of the bits actually transmitted only 57 per cent are the speech itself; the remainder is the error protection that lets it survive fading and interference.

Contents This chapter on its own page

munotes.in669

Chapter Eighty-Eight

The GSM System Architecture

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: GSM: System architecture"

In one line

A GSM network is three subsystems: the radio subsystem, where mobile stations talk to base transceiver stations grouped under base station controllers; the network and switching subsystem, where mobile switching centres connect calls with the help of the home and visitor location registers; and the operation subsystem, where the authentication centre, the equipment identity register and the operation and maintenance centre keep the network honest and running.

In the wording a student can write in an examination: a GSM network has three subsystems.

The radio subsystem (RSS) contains the mobile station (MS), which is the handset plus the subscriber identity module (SIM) that holds the subscription, and the base station subsystem (BSS): "A Base Transceiver Station (BTS) is a network component which serves one cell", and "A Base Station Controller (BSC) is a network component in the PLMN with the functions for control of one or more BTS." The BSC manages radio channels, handovers within its area and the power of the mobiles.

The network and switching subsystem (NSS) contains the mobile services switching centre (MSC), which "constitutes the interface between the radio system and the fixed networks. The MSC performs all necessary functions in order to handle the circuit switched services to and from the mobile stations"; the gateway MSC (GMSC), the MSC "which performs the routing function to the actual location of the MS"; the home location register (HLR), the permanent database of every subscriber; and the visitor location register (VLR), which holds a working copy for the subscribers currently in an MSC area.

The operation subsystem (OSS) contains the authentication centre (AuC), which "stores an identity key for each mobile subscriber registered with the associated HLR"; the equipment identity register (EIR), "the logical entity which is responsible for storing in the network the International Mobile Equipment Identities (IMEIs)", where equipment "is classified as 'allowed', 'tracked', 'prohibited' or it may be unknown"; and the operation and maintenance centre (OMC).

The interfaces are Um (MS to BTS, the air interface), Abis (BTS to BSC), A (BSC to MSC), and the lettered interfaces of the core: B (MSC to VLR), C (GMSC to HLR), D (VLR to HLR), E (MSC to MSC), F (MSC to EIR), G (VLR to VLR) and H (HLR to AuC).

The three subsystems

A diagram in three dashed panels. Left, the radio subsystem: a mobile station with SIM connected over Um to two base transceiver stations, which connect over Abis to a base station controller. Centre, the network and switching subsystem: the BSC connects over A to an MSC, which links to a VLR below it and to a GMSC beside it; the GMSC links to the HLR below and to the PSTN or ISDN; the VLR links to the HLR. Right, the operation subsystem: the authentication centre, the equipment identity register and the operation and maintenance centre, joined by dashed lines to the HLR, the MSC and the BSC respectively

Figure 88.1 The GSM system architecture and its interfaces

The radio subsystem is everything that touches the air. The MS is a handset (mobile equipment) plus a SIM: the equipment has an identity, the IMEI, and the subscription has another, the IMSI, held on the SIM, which is why a subscription can be moved between handsets and a handset can be barred without touching the subscription. A BTS serves one cell: transceivers, antennas and the modulation, coding and encryption of the air interface. A BSC controls one or more BTSs, and it is where radio resource management lives: allocating and releasing channels, ordering handovers between its own cells, and controlling the power of mobiles and BTSs. Together a BSC and its BTSs are a base station subsystem (BSS).

munotes.in670

The GSM System Architecture

The network and switching subsystem is a telephone exchange that has learned about mobility. The MSC switches calls, but with the extra duties mobility brings: paging, handover between BSSs, and the signalling that goes with location updates. The GMSC is the door from other networks: a call arriving for a subscriber cannot be routed without knowing where that subscriber is, so it goes to a GMSC, which asks the HLR. The HLR holds each subscriber's permanent record; the VLR holds a copy of the record for each subscriber currently in its area, so that the MSC need not ask the HLR for every decision.

The operation subsystem is the part the subscriber never sees. The AuC holds the secret key of every SIM and produces the triplets that authenticate it ([GSM Security]); it is kept with the HLR and is deliberately hard to reach. The EIR holds the IMEIs of handsets and classifies them, so that a stolen handset can be refused service. The OMC is where the network is configured, watched and repaired.

The areas

GSM divides its coverage into a hierarchy, and every identity in [Localization and Calling in GSM] refers to one of its levels:

  • A cell is what one BTS serves, identified by a cell identity.
  • A location area is a group of cells, identified by a location area identity (LAI). A mobile reports when it crosses a location area boundary, not when it crosses a cell boundary, and an incoming call pages the whole location area. The size of a location area is therefore a trade between paging traffic and location updates.
  • An MSC area is everything one MSC serves, which is several location areas, and it matches a VLR area.
  • A PLMN (public land mobile network) is a whole operator's network in a country.

What a call asks of each entity

The architecture is best learned by following an incoming call, which the program does. The essential exchange is that a call arrives at a GMSC with only the subscriber's number, and the network must convert that into a route to a switch: "If a network delivering a call to the PLMN cannot interrogate the HLR, the call is routed to an MSC. This MSC will interrogate the appropriate HLR and then route the call to the MSC where the mobile station is located."

munotes.in671

The GSM System Architecture

Behind that lie the registers. The VLR exists because the alternative is unbearable: "A mobile station roaming in an MSC area or within a GERAN/UTRAN pool-area is controlled by a Visitor Location Register. When a Mobile Station (MS) enters a new location area it starts a registration procedure. An MSC in charge of that area notices this registration and transfers to a Visitor Location Register the identity of the location area where the MS is situated. If this MS is not yet registered in the VLR, the VLR and the HLR exchange information to allow the proper handling of CS calls involving the MS." One exchange with the home register on arrival, and then local decisions until the subscriber leaves.

The architecture, computed

The program walks an incoming call to a roaming subscriber through every entity and says what each is asked; lists the orders of magnitude of each entity in a national network; sets the HLR's record beside the VLR's; and tabulates the interfaces.

# The GSM architecture put to work: what each entity is asked for during one
# call, how many of each a network needs, and what the registers hold.
print("Setting up a call to a roaming subscriber: who is asked, and for what")
steps = [
    ("the caller's network", "GMSC", "where is this MSISDN?"),
    ("GMSC", "HLR", "interrogate: give me a roaming number"),
    ("HLR", "VLR", "the subscriber is in your area: allocate an MSRN"),
    ("VLR", "HLR", "here is the MSRN"),
    ("HLR", "GMSC", "route the call to this MSRN"),
    ("GMSC", "MSC", "call for the MSRN you allocated"),
    ("MSC", "VLR", "which location area, and is the subscriber allowed this service?"),
    ("MSC", "BSC", "page the mobile in that location area"),
    ("BSC", "BTS", "send the paging message on the air"),
    ("BTS", "MS", "paging on the air interface"),
    ("MS", "BTS", "channel request, then the call is set up"),
]
for i, (frm, to, what) in enumerate(steps, 1):
    print("  %2d. %-22s -> %-6s %s" % (i, frm, to, what))

print("\nHow many of each a national network has, in rough orders of magnitude:")
for entity, count, why in (("BTS", "tens of thousands", "one per cell"),
                           ("BSC", "hundreds", "each controls tens to hundreds of BTSs"),
                           ("MSC", "tens", "each serves a region"),
                           ("VLR", "one per MSC", "usually built into the MSC"),
                           ("GMSC", "a few", "the doors from other networks"),
                           ("HLR", "one, or a few", "the master record of every subscriber"),
                           ("AuC", "with the HLR", "holds the key of every SIM"),
                           ("EIR", "one", "the list of allowed and barred handsets"),
                           ("OMC", "one or a few", "operation and maintenance")):
    print("  %-6s %-20s %s" % (entity, count, why))

print("\nWhat the two registers hold for one subscriber:")
hlr = ["IMSI", "MSISDN", "the subscribed services and their parameters",
       "the current VLR address", "authentication data from the AuC"]
vlr = ["IMSI and TMSI", "MSISDN", "the services the subscriber may use here",
       "the current location area", "the MSRN when a call is being routed"]
for a, b in zip(hlr + [""] * 5, vlr + [""] * 5):
    if a or b:
        print("  HLR: %-44s VLR: %s" % (a, b))
print("  the HLR is permanent and one per subscriber; the VLR is a working copy, thrown away on leaving")

print("\nThe interfaces a call crosses:")
for name, between, carries in (("Um", "MS and BTS", "the air interface: traffic and signalling"),
                               ("Abis", "BTS and BSC", "traffic and BTS control"),
                               ("A", "BSC and MSC", "traffic and call signalling"),
                               ("B", "MSC and VLR", "the MSC's queries about a visitor"),
                               ("C", "GMSC and HLR", "interrogation for routing"),
                               ("D", "VLR and HLR", "location update and subscriber data"),
                               ("E", "MSC and MSC", "inter-MSC handover"),
                               ("F", "MSC and EIR", "is this handset allowed?"),
                               ("G", "VLR and VLR", "fetching data when a subscriber moves"),
                               ("H", "HLR and AuC", "authentication data")):
    print("  %-5s %-14s %s" % (name, between, carries))
munotes.in672

The GSM System Architecture

Setting up a call to a roaming subscriber: who is asked, and for what
   1. the caller's network   -> GMSC   where is this MSISDN?
   2. GMSC                   -> HLR    interrogate: give me a roaming number
   3. HLR                    -> VLR    the subscriber is in your area: allocate an MSRN
   4. VLR                    -> HLR    here is the MSRN
   5. HLR                    -> GMSC   route the call to this MSRN
   6. GMSC                   -> MSC    call for the MSRN you allocated
   7. MSC                    -> VLR    which location area, and is the subscriber allowed this service?
   8. MSC                    -> BSC    page the mobile in that location area
   9. BSC                    -> BTS    send the paging message on the air
  10. BTS                    -> MS     paging on the air interface
  11. MS                     -> BTS    channel request, then the call is set up

How many of each a national network has, in rough orders of magnitude:
  BTS    tens of thousands    one per cell
  BSC    hundreds             each controls tens to hundreds of BTSs
  MSC    tens                 each serves a region
  VLR    one per MSC          usually built into the MSC
  GMSC   a few                the doors from other networks
  HLR    one, or a few        the master record of every subscriber
  AuC    with the HLR         holds the key of every SIM
  EIR    one                  the list of allowed and barred handsets
  OMC    one or a few         operation and maintenance

What the two registers hold for one subscriber:
  HLR: IMSI                                         VLR: IMSI and TMSI
  HLR: MSISDN                                       VLR: MSISDN
  HLR: the subscribed services and their parameters VLR: the services the subscriber may use here
  HLR: the current VLR address                      VLR: the current location area
  HLR: authentication data from the AuC             VLR: the MSRN when a call is being routed
  the HLR is permanent and one per subscriber; the VLR is a working copy, thrown away on leaving

The interfaces a call crosses:
  Um    MS and BTS     the air interface: traffic and signalling
  Abis  BTS and BSC    traffic and BTS control
  A     BSC and MSC    traffic and call signalling
  B     MSC and VLR    the MSC's queries about a visitor
  C     GMSC and HLR   interrogation for routing
  D     VLR and HLR    location update and subscriber data
  E     MSC and MSC    inter-MSC handover
  F     MSC and EIR    is this handset allowed?
  G     VLR and VLR    fetching data when a subscriber moves
  H     HLR and AuC    authentication data
munotes.in673

The GSM System Architecture

One call, eleven steps. The call reaches a GMSC, which interrogates the HLR; the HLR asks the VLR for a roaming number; the VLR returns one; the HLR passes it back; the GMSC routes the call to the MSC; the MSC asks the VLR who this is and what they may have; the MSC tells the BSC to page the location area; the BSC tells the BTSs; the mobile answers, and the call is built. Every entity in the architecture earns its place in that list, which is why the diagram is worth memorising as a sequence rather than as a picture.

How many of each. Tens of thousands of BTSs, one per cell; hundreds of BSCs; tens of MSCs; a VLR with each MSC; a few GMSCs; one HLR, logically, for the whole network, with the AuC beside it. The shape of the network is a wide base and a narrow top, which is also the shape of its cost and of its failure consequences: losing a BTS darkens a cell, losing the HLR stops the network.

The two registers. The HLR holds the IMSI, the MSISDN, the subscribed services and the address of the VLR the subscriber is currently at. The VLR holds the IMSI and the temporary TMSI, the services the subscriber may use here, the current location area, and, while a call is being routed, the MSRN. The HLR entry is permanent and unique; the VLR entry is a working copy, created on arrival and discarded on departure.

The interfaces. Um, Abis and A across the radio side, then B, C, D, E, F, G and H across the core, each carrying one kind of question: C asks where a call should go, D reports that a subscriber has moved, F asks whether a handset is barred, and H fetches authentication data.

Distinctions

BTSBSCMSC
ServesOne cellOne or more BTSsSeveral BSSs, an MSC area
DoesTransceiver, coding, encryption on UmChannel allocation, handover within its area, power controlCall switching, paging, handover between BSSs, signalling to the core
CountTens of thousandsHundredsTens
InterfacesUm, AbisAbis, AA, B, C, E, F
munotes.in674

The GSM System Architecture

HLRVLR
HoldsThe permanent record of every subscriber of this networkA working copy for subscribers currently in this area
LifetimePermanentFrom arrival to departure
NumberOne (logically)One per MSC
KnowsWhich VLR the subscriber is atWhich location area the subscriber is in
Asked byThe GMSC (routing), the VLR (on arrival)The MSC (on every call)
AuCEIR
Keyed onThe subscription (IMSI)The equipment (IMEI)
HoldsThe secret key of each SIMLists of handsets: allowed, tracked, prohibited
Used forAuthentication and ciphering keysRefusing stolen or faulty handsets
Sits withThe HLROn its own
InterfaceBetweenCarries
UmMS and BTSThe air interface
AbisBTS and BSCTraffic and BTS control
ABSC and MSCTraffic and call signalling
B, C, DMSC-VLR, GMSC-HLR, VLR-HLRMobility and routing queries
E, F, G, HMSC-MSC, MSC-EIR, VLR-VLR, HLR-AuCHandover, equipment checks, data transfer, keys

What it does not mean

The MS is not the handset alone. It is the mobile equipment plus the SIM, and the two have separate identities and separate fates.

The VLR is not a cache of the HLR that can be stale. The HLR knows which VLR holds each subscriber and updates or cancels it when they move.

The GMSC is not a different kind of switch. It is an MSC doing the job of interrogating the HLR for an incoming call.

A location area is not a cell. It is a group of cells; mobiles report at its boundary and are paged across all of it.

The OSS is not optional. Without the AuC no subscriber can be authenticated, and without the OMC the network cannot be operated.

Quick revision

  • RSS: MS (mobile equipment + SIM), BTS (serves one cell), BSC (controls one or more BTSs; channel allocation, handover, power control); BTS + BSC = BSS.
  • NSS: MSC (interface between the radio system and the fixed networks; switching, paging, handover), GMSC (interrogates the HLR and routes an incoming call), HLR (permanent subscriber record), VLR (working copy for visitors).
  • OSS: AuC (a key per SIM, with the HLR), EIR (IMEIs: allowed, tracked, prohibited), OMC.
  • Interfaces: Um (MS-BTS), Abis (BTS-BSC), A (BSC-MSC), B (MSC-VLR), C (GMSC-HLR), D (VLR-HLR), E (MSC-MSC), F (MSC-EIR), G (VLR-VLR), H (HLR-AuC).
  • Areas: cell (one BTS), location area (paging and location update unit), MSC area = VLR area, PLMN.
  • Program: an incoming call touches GMSC, HLR, VLR, MSC, BSC, BTS, MS in that order; a national network has tens of thousands of BTSs and one HLR.
munotes.in675

The GSM System Architecture

Test yourself

1. Draw and explain the GSM system architecture. Three subsystems. The radio subsystem holds the mobile station, which is the mobile equipment plus the SIM, and the base station subsystem: base transceiver stations, each serving one cell, connected over the Abis interface to a base station controller that controls one or more of them. The network and switching subsystem holds the mobile services switching centre, which switches calls and manages mobility, the gateway MSC through which calls from other networks enter, the home location register with each subscriber's permanent record and the visitor location register with a copy for subscribers currently present. The operation subsystem holds the authentication centre, the equipment identity register and the operation and maintenance centre. The interfaces are Um between mobile and BTS, Abis between BTS and BSC, A between BSC and MSC, and the lettered interfaces B to H within the core.

2. What does a BSC do that a BTS does not? A BTS serves a single cell: it contains the transceivers and antennas and performs the modulation, coding, encryption and timing of the air interface. A BSC controls one or more BTSs and takes the decisions: it allocates and releases radio channels, decides and executes handovers between the cells it controls, controls the transmit power of the mobiles and the BTSs, and concentrates the traffic of its BTSs onto the A interface toward the MSC.

3. Distinguish the HLR and the VLR. The HLR is the permanent database of the operator's subscribers, one logical instance for the network, holding for each subscriber the IMSI, the MSISDN, the subscribed services and their parameters, and the address of the VLR where the subscriber currently is. The VLR is a temporary database, usually one per MSC, holding a working copy of the data for every subscriber currently in that MSC area, together with the current location area and the temporary identity; it is created when the subscriber arrives and discarded when they leave. The arrangement means that routine decisions are taken locally and the home register is consulted only on arrival and when a call must be routed.

4. What is the GMSC, and why is it needed? The gateway MSC is the MSC that receives a call arriving from another network for one of this network's subscribers and routes it to the MSC where the subscriber actually is. It is needed because the caller dials only a number, which says nothing about where the subscriber is; the GMSC interrogates the HLR, which asks the serving VLR for a mobile station roaming number, and the call is then routed to that number. Any MSC can act as a GMSC; it is a role, not a different kind of switch.

munotes.in676

The GSM System Architecture

5. What do the AuC and the EIR hold, and why are they separate? The authentication centre holds a secret key for each subscriber, associated with the HLR, and generates the data used to authenticate the SIM and derive the ciphering key. The equipment identity register holds the international mobile equipment identities of handsets and classifies them as allowed, tracked or prohibited. They are separate because they answer different questions about different things: the AuC is about the subscription and its secrecy, the EIR is about the physical handset, so a stolen handset can be barred without affecting its owner's subscription, and a subscription can move to another handset without affecting its keys.

6. What are the areas of a GSM network, and why does the location area matter? A cell is the area served by one BTS; a location area is a group of cells sharing a location area identity; an MSC area, which coincides with a VLR area, is all the location areas served by one MSC; and a PLMN is an operator's whole network. The location area matters because it is the unit of both location updating and paging: a mobile reports its position only when it crosses a location area boundary, and an incoming call is paged in every cell of the location area. Large location areas mean few updates and much paging; small ones mean the reverse.

Contents This chapter on its own page

munotes.in677

Chapter Eighty-Nine

The GSM Radio Interface: Carriers, the TDMA Frame and Bursts

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: GSM: Radio interface"

In one line

GSM divides its band into 200 kHz carriers and each carrier into a repeating frame of eight time slots, so a physical channel is one slot on one carrier; in that slot a mobile sends a burst of 156.25 bit periods, of which 116 carry data, 26 carry a training sequence for the equaliser and 8.25 are a guard time, and because the uplink runs three slots behind the downlink the mobile never transmits and receives at once and has spare time to measure its neighbours.

In the wording a student can write in an examination: GSM combines FDMA and TDMA. The band is divided into carriers 200 kHz apart; P-GSM 900 uses "890 MHz to 915 MHz: mobile transmit, base receive" and "935 MHz to 960 MHz: base transmit, mobile receive", a duplex spacing of 45 MHz and 124 usable carriers; E-GSM 900 extends it to "880 MHz to 915 MHz" and "925 MHz to 960 MHz", and DCS 1800 and GSM 850 are further bands. Each carrier is divided in time into a TDMA frame of 8 slots. At the standard's symbol rate of 1625/6 ksymbol/s (about 270.833 kbit/s), one bit period is 3.692 microseconds, a slot is 156.25 bit periods = 576.9 microseconds, and a frame is 4.615 ms. A physical channel is one slot in every frame on one carrier.

The normal burst carries, in order, 3 tail bits, 58 encrypted bits, a 26-bit training sequence, 58 more encrypted bits, 3 tail bits and a guard period of 8.25 bit periods. The training sequence in the middle lets the receiver's equaliser measure the channel where it matters most; the guard period, 30.5 microseconds, covers the ramping of the transmitter and the spread of propagation delays. The other bursts are the frequency correction burst, the synchronisation burst, the dummy burst and the access burst, which has a much longer guard period because a mobile's distance is not yet known. The uplink is delayed by three slots from the downlink, so a mobile need not transmit and receive simultaneously and has an idle period in which to measure neighbouring cells. Timing advance lets the network tell a distant mobile to transmit early, and power control keeps every mobile's signal at the level the base station needs and no more.

Carriers: the FDMA part

GSM's band is cut into carriers 200 kHz apart, and each carrier is used in both directions on different frequencies, which is frequency division duplex. The standard fixes the bands; the program counts what fits.

The duplex spacing of 45 MHz at 900 MHz is not arbitrary: it must be wide enough for a handset's duplex filter to keep its own transmitter out of its own receiver, which matters because the transmitter is a fraction of a metre from the receiver and may be 10^13 times stronger than the signal being received.

munotes.in678

The GSM Radio Interface: Carriers, the TDMA Frame and Bursts

A carrier gives 8 slots, so P-GSM 900's 124 usable carriers give 992 physical channels for a whole country's operator: hopelessly few, which is why the cellular reuse of [Cellular Systems: Cells, Clusters and Frequency Reuse] is not an optimisation but the entire basis of the system.

The frame: the TDMA part

Each carrier's time is divided into slots, and the slots are grouped into TDMA frames of eight. A mobile in a full-rate traffic channel uses one slot in every frame, so it transmits for one eighth of the time, which:

  • lets eight calls share one carrier;
  • lets the handset's transmitter be off seven eighths of the time, which saves battery and lets a small handset radiate at up to 2 W in its slot with a mean of a quarter of that;
  • leaves the handset free time to retune and measure other cells, which is what makes handover decisions possible ([Handover in GSM]).

Every duration follows from the symbol rate. The program computes them: a bit period of 3.692 microseconds, a slot of 576.9 microseconds, a frame of 4.615 ms.

The normal burst

A burst is what is actually transmitted in a slot. The standard's Table 5.2.3-1 gives the normal burst, and the program prints it with each field's duration.

Every field is there for a reason:

  • Tail bits (3 at each end, all zeros) bring the modulator into a known state and let the equaliser start and finish cleanly.
  • Encrypted bits, 58 and 58, are the payload: 116 bits of coded, interleaved, ciphered data.
  • The training sequence (26 bits) is a pattern the receiver already knows. By comparing what it received with what it expected, the receiver measures the channel's impulse response and sets its equaliser to undo the smearing that [Multipath, Fading and the Doppler Effect] measured. It sits in the middle so that the estimate is equally fresh for both halves of the burst.
  • The guard period (8.25 bit periods) lets the transmitter ramp up and down without splashing into neighbouring slots, and absorbs the difference in propagation delay between mobiles at different distances. The standard's own reason: "The guard period is provided because it is required for the MSs that transmission be attenuated for the period between bursts with the necessary ramp up and down occurring during the guard periods".

116 bits of payload in 156.25 bit periods is 74.2 per cent: a quarter of the air time goes on making the rest usable.

munotes.in679

The GSM Radio Interface: Carriers, the TDMA Frame and Bursts

Timing advance, and why the guard is not enough

A burst from a mobile 9.1 km away arrives one whole guard period late, as the program computes. Beyond that it would land in the next slot. GSM's answer is timing advance: the base station measures how late each mobile's burst arrives and tells it, in steps of one bit period, to transmit that much earlier. With the timing advance range the system defines, cells can reach about 35 km, and beyond that a special "extended range" arrangement is needed.

Power control works the same way and for the same reason: the base station tells each mobile to use the least power that still delivers an acceptable quality, which saves the handset's battery and, more importantly, lowers the interference every other cell sees.

The other bursts

  • The frequency correction burst is all zeros, which after GMSK modulation is a pure tone. A mobile searching for a network finds this tone and locks its own oscillator to it.
  • The synchronisation burst carries a long training sequence, easy to find, and the frame number, which tells the mobile where it is in the frame hierarchy of [GSM Logical Channels and the Frame Hierarchy] and is needed for ciphering and hopping.
  • The dummy burst is sent on the carrier that carries the broadcast channel whenever there is nothing else to send, so that the carrier's power stays constant and mobiles can measure it.
  • The access burst is a mobile's first transmission, before the network knows its distance. It carries less data and a much longer guard period, so that it cannot collide with the next slot however far away the mobile is.

The radio interface, computed

The program lists the bands with their carriers and duplex spacing, computes the frame's durations from the symbol rate, prints the normal burst field by field with each field's duration, computes the distance the guard period allows, explains the three-slot offset in milliseconds, and lists the burst types.

# The GSM radio interface, computed: the carriers in a band, the TDMA frame
# and its slots, the normal burst field by field, and what a slot carries.
C = 299792458.0
RATE = 1625e3 / 6                                   # symbols (and bits) a second, TS 45.004

# 1. The bands and their carriers, 200 kHz apart, from TS 45.005.
print("GSM bands (TS 45.005) and the carriers that fit, 200 kHz apart:")
for name, up_lo, up_hi, dn_lo, dn_hi in (("P-GSM 900", 890.0, 915.0, 935.0, 960.0),
                                         ("E-GSM 900", 880.0, 915.0, 925.0, 960.0),
                                         ("R-GSM 900", 876.0, 915.0, 921.0, 960.0),
                                         ("GSM 850", 824.0, 849.0, 869.0, 894.0)):
    n = round((up_hi - up_lo) / 0.2)
    print("  %-10s uplink %6.1f to %6.1f MHz, downlink %6.1f to %6.1f MHz: %3d carriers, duplex %4.0f MHz"
          % (name, up_lo, up_hi, dn_lo, dn_hi, n, dn_lo - up_lo))
print("  each carrier carries 8 slots, so P-GSM 900 has %d physical channels in all" % (125 * 8))

# 2. The frame. A bit period is 1/(1625/6 kbit/s); a slot is 156.25 bit
#    periods; a frame is 8 slots.
bit = 1 / RATE
slot = 156.25 * bit
frame = 8 * slot
print("\nThe TDMA frame:")
print("  bit period      %.6f ms   (%.3f microseconds, at %.3f kbit/s)" % (bit * 1e3, bit * 1e6, RATE / 1e3))
print("  time slot       %.6f ms   (156.25 bit periods)" % (slot * 1e3))
print("  TDMA frame      %.6f ms   (8 slots)" % (frame * 1e3))
print("  a mobile transmits in 1 slot of 8, so its duty cycle is %.1f%% and it may listen in the rest"
      % (100 / 8))

# 3. The normal burst, field by field (TS 45.002 Table 5.2.3-1).
print("\nThe normal burst (GMSK), 156.25 bit periods:")
fields = [("tail bits", 3), ("encrypted bits e0 to e57", 58), ("training sequence", 26),
          ("encrypted bits e58 to e115", 58), ("tail bits", 3), ("guard period", 8.25)]
total = 0
for name, bits in fields:
    total += bits
    print("  %-28s %6.2f bits %8.2f microseconds" % (name, bits, bits * bit * 1e6))
print("  %-28s %6.2f bits %8.2f microseconds" % ("total", total, total * bit * 1e6))
print("  payload: %d of %.2f bits, %.1f%%; the guard alone is %.1f microseconds, %.1f km of propagation"
      % (116, total, 100 * 116 / total, 8.25 * bit * 1e6, 8.25 * bit * C / 1000))

# 4. Uplink and downlink are offset by 3 slots, so a mobile never transmits and
#    receives at once and has time to measure its neighbours.
print("\nWhy the uplink is delayed by 3 slots:")
print("  the mobile receives in its slot, then transmits 3 slots later: a gap of %.2f ms" % (3 * slot * 1e3))
print("  and it has %.2f ms free before its next reception, which it spends measuring neighbour cells"
      % ((8 - 3 - 1) * slot * 1e3))

# 5. The other bursts, by what they are for.
print("\nThe five burst types:")
for name, what in (("normal (NB)", "traffic and most signalling: two blocks of 58 encrypted bits, a training sequence between"),
                   ("frequency correction (FB)", "all zeros: an unmodulated tone the mobile tunes to"),
                   ("synchronisation (SB)", "a long training sequence and the frame number"),
                   ("dummy (DB)", "sent on the BCCH carrier when there is nothing else, to keep it constant"),
                   ("access (AB)", "the mobile's first transmission, with a long guard period for unknown distance")):
    print("  %-26s %s" % (name, what))
munotes.in680

The GSM Radio Interface: Carriers, the TDMA Frame and Bursts

GSM bands (TS 45.005) and the carriers that fit, 200 kHz apart:
  P-GSM 900  uplink  890.0 to  915.0 MHz, downlink  935.0 to  960.0 MHz: 125 carriers, duplex   45 MHz
  E-GSM 900  uplink  880.0 to  915.0 MHz, downlink  925.0 to  960.0 MHz: 175 carriers, duplex   45 MHz
  R-GSM 900  uplink  876.0 to  915.0 MHz, downlink  921.0 to  960.0 MHz: 195 carriers, duplex   45 MHz
  GSM 850    uplink  824.0 to  849.0 MHz, downlink  869.0 to  894.0 MHz: 125 carriers, duplex   45 MHz
  each carrier carries 8 slots, so P-GSM 900 has 1000 physical channels in all

The TDMA frame:
  bit period      0.003692 ms   (3.692 microseconds, at 270.833 kbit/s)
  time slot       0.576923 ms   (156.25 bit periods)
  TDMA frame      4.615385 ms   (8 slots)
  a mobile transmits in 1 slot of 8, so its duty cycle is 12.5% and it may listen in the rest

The normal burst (GMSK), 156.25 bit periods:
  tail bits                      3.00 bits    11.08 microseconds
  encrypted bits e0 to e57      58.00 bits   214.15 microseconds
  training sequence             26.00 bits    96.00 microseconds
  encrypted bits e58 to e115    58.00 bits   214.15 microseconds
  tail bits                      3.00 bits    11.08 microseconds
  guard period                   8.25 bits    30.46 microseconds
  total                        156.25 bits   576.92 microseconds
  payload: 116 of 156.25 bits, 74.2%; the guard alone is 30.5 microseconds, 9.1 km of propagation

Why the uplink is delayed by 3 slots:
  the mobile receives in its slot, then transmits 3 slots later: a gap of 1.73 ms
  and it has 2.31 ms free before its next reception, which it spends measuring neighbour cells

The five burst types:
  normal (NB)                traffic and most signalling: two blocks of 58 encrypted bits, a training sequence between
  frequency correction (FB)  all zeros: an unmodulated tone the mobile tunes to
  synchronisation (SB)       a long training sequence and the frame number
  dummy (DB)                 sent on the BCCH carrier when there is nothing else, to keep it constant
  access (AB)                the mobile's first transmission, with a long guard period for unknown distance
munotes.in681

The GSM Radio Interface: Carriers, the TDMA Frame and Bursts

The bands. P-GSM 900 has 125 carrier positions in 25 MHz, of which 124 are used, with a 45 MHz duplex spacing; E-GSM adds 50 more and R-GSM, for railways, 20 beyond that. Eight slots on each gives about a thousand physical channels for an operator's whole country.

The frame. A bit period is 3.692 microseconds, a slot 576.9 microseconds, a frame 4.615 ms. Those three numbers, and 156.25, are worth memorising: every other timing in GSM is built from them, including the multiframes of the next chapter.

The burst. Of 156.25 bit periods, 116 are payload (74.2 per cent), 26 are training sequence, 6 are tail bits and 8.25 are guard. The guard alone is 30.5 microseconds, in which a radio wave travels 9.1 km: that distance is the reason timing advance exists.

munotes.in682

The GSM Radio Interface: Carriers, the TDMA Frame and Bursts

The offset. Because the uplink slot comes three slots after the downlink slot, the mobile has 1.73 ms between receiving and transmitting, and 2.31 ms after transmitting before it must receive again. It spends that idle time retuning to neighbouring carriers and measuring their broadcast channels, which is how it can report six neighbours to the network without ever missing its own slot.

Distinctions

FDMA in GSMTDMA in GSM
DividesThe band into carriersEach carrier into slots
UnitA carrier, 200 kHzA slot, 576.9 microseconds
GuardGuard bands inside the 200 kHz8.25 bit periods per burst
Gives124 carriers in P-GSM 9008 channels per carrier
Field of the normal burstBitsPurpose
Tail bits3 + 3Bring the modulator and equaliser to a known state
Encrypted bits58 + 58The payload, 116 bits
Training sequence26Lets the equaliser measure the channel, placed centrally
Guard period8.25Ramping, and differences in propagation delay
NormalFrequency correctionSynchronisationDummyAccess
Carries116 data bitsAll zeros (a tone)Frame number, long trainingNothing usefulA short request
Guard8.25 bits8.258.258.25Much longer
Used forTraffic and signallingFinding and tuning to a cellTime and frame alignmentKeeping the BCCH carrier constantThe mobile's first transmission

What it does not mean

A physical channel is not a carrier. It is one slot on one carrier, repeated every frame.

The training sequence is not overhead to be removed. Without it the equaliser cannot work, and without the equaliser the urban delay spread destroys the burst.

The guard period does not set the cell size. Timing advance extends the reach to about 35 km; the guard sets what can be tolerated without it.

The mobile is not idle for seven eighths of a frame. It uses the gaps to measure neighbouring cells, which is what makes handover work.

The uplink and downlink are not simultaneous. The three-slot offset is deliberate, so that a handset needs no duplexer for simultaneous transmission and reception.

Quick revision

  • Carriers 200 kHz apart, FDD. P-GSM 900: up 890 to 915, down 935 to 960, duplex 45 MHz, 124 usable carriers; E-GSM: 880 to 915 and 925 to 960; also R-GSM, GSM 850, DCS 1800.
  • Symbol rate 1625/6 ksymbol/s (270.833 kbit/s): bit 3.692 us, slot 156.25 bits = 576.9 us, frame 8 slots = 4.615 ms. A physical channel = one slot per frame on one carrier.
  • Normal burst: 3 tail, 58 data, 26 training, 58 data, 3 tail, 8.25 guard = 156.25. Payload 116 bits, 74.2 per cent; guard 30.5 us = 9.1 km.
  • Uplink offset by 3 slots: 1.73 ms between receive and transmit, 2.31 ms free for neighbour measurements.
  • Timing advance (measured by the BTS, in bit periods) extends the reach to about 35 km; power control cuts interference and saves battery.
  • Bursts: normal, frequency correction (all zeros, a tone), synchronisation (frame number, long training), dummy (keeps the BCCH carrier constant), access (long guard, unknown distance).
munotes.in683

The GSM Radio Interface: Carriers, the TDMA Frame and Bursts

Test yourself

1. How does GSM combine FDMA and TDMA? The allocated band is divided by frequency into carriers 200 kHz apart, which is FDMA, and each carrier's time is divided into a repeating frame of eight slots, which is TDMA. A physical channel is therefore one slot in every frame on one carrier, and eight such channels share a carrier. Uplink and downlink use different carriers, separated in GSM 900 by 45 MHz, which is frequency division duplex.

2. Compute the duration of a bit, a slot and a TDMA frame. The modulating symbol rate is 1625/6 ksymbol/s, about 270.833 kbit/s, so a bit period is its reciprocal, about 3.692 microseconds. A slot is 156.25 bit periods, which is 576.9 microseconds. A frame is eight slots, which is 4.615 milliseconds.

3. Describe the normal burst field by field and say why each field is there. Three tail bits, 58 encrypted bits, a 26-bit training sequence, 58 more encrypted bits, three tail bits and a guard period of 8.25 bit periods, 156.25 in all. The tail bits put the modulator and the receiver's equaliser into a known state at each end. The two blocks of encrypted bits are the 116 bits of payload. The training sequence is a known pattern from which the receiver measures the channel and configures its equaliser; it is placed in the middle so that the estimate is equally valid for both halves. The guard period allows the transmitter to ramp up and down without disturbing the neighbouring slots and absorbs differences in propagation delay.

4. Why is the guard period 8.25 bit periods, and what happens for distant mobiles? It is enough time, about 30.5 microseconds, for the transmitter's power to ramp up and down and for modest differences in propagation delay between mobiles. A wave travels about 9.1 km in that time, so a mobile further away than that would have its burst arrive more than a guard period late and overlap the next slot. The network therefore measures how late each mobile's burst arrives and orders a timing advance, so the mobile transmits correspondingly early; with the range of timing advance GSM provides, cells can reach about 35 km.

5. Why is the uplink delayed by three slots relative to the downlink? So that the mobile never has to transmit and receive at the same time, which would require an expensive duplexer and a much better filter, and so that it has idle periods. After receiving in its downlink slot the mobile has about 1.73 ms before it must transmit, and after transmitting about 2.31 ms before its next reception; it uses those gaps to retune to neighbouring carriers and measure their broadcast channels, which is what allows it to report neighbours for handover decisions.

munotes.in684

The GSM Radio Interface: Carriers, the TDMA Frame and Bursts

6. Name the five burst types and state the purpose of each. The normal burst carries traffic and most signalling, with 116 encrypted bits and a training sequence. The frequency correction burst consists of all zeros, which after GMSK modulation is an unmodulated tone that a mobile uses to find a cell and correct its oscillator. The synchronisation burst carries a long, easily detected training sequence and the frame number, letting the mobile align itself with the frame structure. The dummy burst is transmitted on the broadcast carrier when there is nothing else to send, keeping that carrier's power constant so mobiles can measure it. The access burst is the mobile's first transmission, carrying little data and a long guard period because its distance from the base station is not yet known.

Contents This chapter on its own page

munotes.in685

Chapter Ninety

GSM Logical Channels and the Frame Hierarchy

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: GSM: Logical channels and frame hierarchy"

In one line

A physical channel is a slot on a carrier; a logical channel is a job, and GSM maps many jobs onto the same slots by giving each one its turn in a repeating multiframe: 26 frames for traffic, where 24 carry speech, one carries the slow control channel and one is idle so the mobile can look around, and 51 frames for control, where the mobile finds the tone to tune to, the frame number, the cell's broadcast data and its own pages.

In the wording a student can write in an examination: a physical channel is one time slot on one carrier, recurring every TDMA frame. A logical channel is a particular kind of information flow, mapped onto physical channels by the frame structure. The logical channels are:

  • Traffic channels (TCH): TCH/F (full rate) and TCH/H (half rate, two calls sharing one slot).
  • Broadcast channels, downlink only, point to multipoint: BCCH (broadcast control channel: cell identity, neighbours, access parameters), FCCH (frequency correction channel: an unmodulated tone), SCH (synchronisation channel: the frame number and the base station identity code).
  • Common control channels (CCCH), for mobiles not yet on a dedicated channel: PCH (paging channel, downlink), RACH (random access channel, uplink, using slotted ALOHA), AGCH (access grant channel, downlink).
  • Dedicated control channels: SDCCH (stand-alone dedicated control channel, for signalling before a traffic channel is assigned), SACCH (slow associated control channel, always accompanying a traffic or SDCCH channel, carrying measurements and power control), FACCH (fast associated control channel, which steals traffic frames when urgent signalling such as handover is needed).

The frame hierarchy: "The complete cycle of TDMA frame numbers from 0 to FN_MAX is defined as a hyperframe." A hyperframe holds "2048 superframes", and a superframe is 26 multiplied by 51 TDMA frames. "A 26-multiframe, comprising 26 TDMA frames, is used to support traffic and associated control channels and a 51-multiframe, comprising 51 TDMA frames, is used to support broadcast, common control and stand alone dedicated control (and their associated control) channels." So a superframe is "51 traffic/associated control multiframes or 26 broadcast/common control multiframes", and the standard gives FN_MAX as 2715647. The hyperframe is long because "The need for a hyperframe of a substantially longer period than a superframe arises from the requirements of the encryption process which uses FN as an input parameter."

Physical and logical

One slot on one carrier is a physical channel: a fixed amount of capacity, 114 payload bits every 4.615 ms. What that capacity is used for changes from frame to frame, and each use is a logical channel.

The mapping is a timetable. In a traffic channel's 26-frame cycle, 24 frames carry speech, one carries the slow control channel and one is idle. In a control channel's 51-frame cycle, particular frames carry the tone, the frame number, the broadcast data and the paging messages. A mobile that knows the frame number therefore knows exactly what every slot contains, which is why the frame number is broadcast and why it matters so much.

munotes.in686

GSM Logical Channels and the Frame Hierarchy

The four families

Traffic channels carry what the user is paying for. TCH/F uses one slot in every frame of the 26-multiframe except two, giving the 22.8 kbit/s that [GSM and Its Mobile Services] followed from the codec. TCH/H gives each of two calls alternate frames, halving the rate and doubling the capacity of a cell, at some cost in speech quality.

Broadcast channels are how a mobile finds and joins a cell, and they are downlink only:

  • FCCH is a burst of all zeros, which after GMSK is a pure tone. A mobile scanning for a network looks for this first; finding it corrects the mobile's own oscillator and identifies the carrier as a BCCH carrier.
  • SCH follows, carrying the frame number and the base station identity code, so the mobile learns where it is in the hierarchy.
  • BCCH then tells it everything about the cell: the identity, the location area, which carriers the neighbours use, whether access is allowed, the parameters for random access, and more.

Common control channels serve mobiles that have no dedicated channel yet:

  • PCH: the network pages a mobile when a call arrives, in every cell of its location area.
  • RACH: the mobile answers, or asks for a channel, in a slotted ALOHA contention ([Contention: ALOHA and CSMA] gives the throughput of that scheme, and the access burst of [The GSM Radio Interface: Carriers, the TDMA Frame and Bursts] is what it sends).
  • AGCH: the network replies with an assignment to a dedicated channel.

Dedicated control channels serve one mobile:

  • SDCCH carries the signalling of location updates, authentication, SMS and call setup, before any traffic channel is committed. Giving a whole traffic channel to a location update would be wasteful, so GSM gives eight SDCCHs the capacity of one traffic channel.
  • SACCH is always present alongside a traffic channel or an SDCCH, in the 26-multiframe's frame 12, and carries the measurements the mobile makes of its neighbours and the power and timing commands coming back. It is the channel on which handover decisions are fed ([Handover in GSM]).
  • FACCH has no frames of its own: when something urgent must be said, such as a handover command, it steals a traffic frame, flagging the theft so the receiver does not mistake it for speech. The speech gap is short enough to be masked by the codec.
munotes.in687

GSM Logical Channels and the Frame Hierarchy

The frame hierarchy

The standard's own account, quoted above, gives the whole structure:

  • TDMA frame: 8 slots, 4.615 ms.
  • 26-multiframe: 26 frames, for traffic and associated control.
  • 51-multiframe: 51 frames, for broadcast, common and dedicated control.
  • Superframe: 1,326 frames, that is 26 × 51, which is 51 traffic multiframes or 26 control multiframes: the point at which the two cycles line up again.
  • Hyperframe: 2,048 superframes, 2,715,648 frames.

The hyperframe exists for ciphering. The frame number is an input to the A5 algorithm ([GSM Security]), so the keystream repeats when the frame number does; a hyperframe of nearly three and a half hours makes that irrelevant for any call.

The location area trade

Paging happens on the PCH, in every cell of the location area, so a large location area costs paging traffic on every incoming call. Location updating happens whenever a mobile crosses a location area boundary, so a small location area costs signalling from every moving mobile. The operator chooses the size where the two balance, and the program shows the shape of the trade.

The hierarchy, computed

The program computes the duration of every level of the hierarchy from the symbol rate; lays out the 26-multiframe and the 51-multiframe; lists the logical channels by family; and prices the location area trade.

# GSM's frame hierarchy and its logical channels, computed from the standard's
# own numbers (TS 45.002 clause 4.3.3).
FRAME = 8 * 156.25 / (1625e3 / 6) * 1000           # 4.615 ms, from the symbol rate

print("The frame hierarchy (TS 45.002, 4.3.3): FN_MAX = (26 x 51 x 2048) - 1 = %d" % (26 * 51 * 2048 - 1))
levels = [("TDMA frame", 1), ("26-multiframe (traffic)", 26), ("51-multiframe (control)", 51),
          ("superframe = 26 x 51 frames", 26 * 51), ("hyperframe = 2048 superframes", 26 * 51 * 2048)]
print("   level                              frames        duration")
for name, frames in levels:
    s = frames * FRAME / 1000        # in seconds
    if s < 1:
        d = "%8.2f ms" % (s * 1000)
    elif s < 600:
        d = "%8.3f s " % s
    else:
        d = "%8.2f hours" % (s / 3600)
    print("  %-34s %8d %s" % (name, frames, d))
print("  the hyperframe is long because the frame number is an input to the ciphering: it must")
print("  not repeat while a key is in use.")

# A superframe two ways, as the standard puts it.
print("\n  a superframe is %d traffic multiframes of 26, or %d control multiframes of 51"
      % (51, 26))

# 2. The 26-multiframe: 24 traffic frames, one for the slow associated control
#    channel, one idle. What that leaves for speech.
print("\nThe 26-multiframe (traffic), frame by frame:")
print("  frames  0 to 11 and 13 to 24: traffic (TCH), %d frames" % 24)
print("  frame  12: slow associated control channel (SACCH)")
print("  frame  25: idle, which the mobile uses to measure neighbours")
print("  so speech gets %d of 26 frames, %.1f%% of the channel, and the idle frame comes every %.0f ms"
      % (24, 100 * 24 / 26, 26 * FRAME))

# 3. The 51-multiframe on the downlink of a BCCH carrier, in outline.
print("\nThe 51-multiframe (control), what it carries on the downlink:")
for name, count, what in (("FCCH", 5, "frequency correction, so a mobile can tune"),
                          ("SCH", 5, "synchronisation, carrying the frame number"),
                          ("BCCH", 4, "the cell's broadcast information"),
                          ("CCCH", 36, "paging and access grant"),
                          ("idle", 1, "")):
    print("  %-6s %3d frames  %s" % (name, count, what))
print("  total %d frames, repeating every %.1f ms" % (5 + 5 + 4 + 36 + 1, 51 * FRAME))

# 4. The logical channels, by group.
print("\nThe logical channels:")
groups = [("Traffic", [("TCH/F", "full-rate traffic, 24 of every 26 frames"),
                       ("TCH/H", "half-rate traffic, two calls share one slot")]),
          ("Broadcast", [("BCCH", "cell identity, the neighbours, access parameters"),
                         ("FCCH", "a tone to tune to"),
                         ("SCH", "frame number and base station identity code")]),
          ("Common control", [("PCH", "paging: the network calls a mobile"),
                              ("RACH", "random access: the mobile's first uplink message"),
                              ("AGCH", "access grant: here is your dedicated channel")]),
          ("Dedicated control", [("SDCCH", "signalling before a traffic channel is given"),
                                 ("SACCH", "slow: measurements and power control, always present"),
                                 ("FACCH", "fast: steals traffic frames when something urgent happens")])]
for group, chans in groups:
    print("  %s:" % group)
    for name, what in chans:
        print("    %-6s %s" % (name, what))

# 5. What paging costs. A location area of n cells, paged for every incoming
#    call, against the location updates a smaller area would cause.
print("\nThe location area trade: paging against updating (a model)")
print("   cells in the area   paging messages per call   updates per user per hour (crossing at 6/hour)")
for cells in (1, 10, 50, 200):
    crossings = 6 / cells ** 0.5
    print("  %18d %26d %36.2f" % (cells, cells, crossings))
print("  a big area pages many cells for each call; a small one is crossed often. The operator")
print("  picks the size where the two traffics balance.")
munotes.in688

GSM Logical Channels and the Frame Hierarchy

The frame hierarchy (TS 45.002, 4.3.3): FN_MAX = (26 x 51 x 2048) - 1 = 2715647
   level                              frames        duration
  TDMA frame                                1     4.62 ms
  26-multiframe (traffic)                  26   120.00 ms
  51-multiframe (control)                  51   235.38 ms
  superframe = 26 x 51 frames            1326    6.120 s
  hyperframe = 2048 superframes       2715648     3.48 hours
  the hyperframe is long because the frame number is an input to the ciphering: it must
  not repeat while a key is in use.

  a superframe is 51 traffic multiframes of 26, or 26 control multiframes of 51

The 26-multiframe (traffic), frame by frame:
  frames  0 to 11 and 13 to 24: traffic (TCH), 24 frames
  frame  12: slow associated control channel (SACCH)
  frame  25: idle, which the mobile uses to measure neighbours
  so speech gets 24 of 26 frames, 92.3% of the channel, and the idle frame comes every 120 ms

The 51-multiframe (control), what it carries on the downlink:
  FCCH     5 frames  frequency correction, so a mobile can tune
  SCH      5 frames  synchronisation, carrying the frame number
  BCCH     4 frames  the cell's broadcast information
  CCCH    36 frames  paging and access grant
  idle     1 frames
  total 51 frames, repeating every 235.4 ms

The logical channels:
  Traffic:
    TCH/F  full-rate traffic, 24 of every 26 frames
    TCH/H  half-rate traffic, two calls share one slot
  Broadcast:
    BCCH   cell identity, the neighbours, access parameters
    FCCH   a tone to tune to
    SCH    frame number and base station identity code
  Common control:
    PCH    paging: the network calls a mobile
    RACH   random access: the mobile's first uplink message
    AGCH   access grant: here is your dedicated channel
  Dedicated control:
    SDCCH  signalling before a traffic channel is given
    SACCH  slow: measurements and power control, always present
    FACCH  fast: steals traffic frames when something urgent happens

The location area trade: paging against updating (a model)
   cells in the area   paging messages per call   updates per user per hour (crossing at 6/hour)
                   1                          1                                 6.00
                  10                         10                                 1.90
                  50                         50                                 0.85
                 200                        200                                 0.42
  a big area pages many cells for each call; a small one is crossed often. The operator
  picks the size where the two traffics balance.
munotes.in689

GSM Logical Channels and the Frame Hierarchy

The hierarchy. A frame is 4.615 ms; a 26-multiframe 120.00 ms; a 51-multiframe 235.38 ms; a superframe 6.120 s; a hyperframe 3.48 hours. The 120 ms of the traffic multiframe is why SACCH measurements arrive about eight times a second, which is the rhythm of handover decisions. The 3.48 hours is the ciphering requirement, and it is the reason GSM's frame number is 22 bits rather than something convenient.

The traffic multiframe. 24 of 26 frames carry speech, 92.3 per cent; frame 12 is the SACCH and frame 25 is idle. The idle frame is not waste: it is when the mobile retunes to a neighbouring BCCH carrier and reads its synchronisation channel, which is how it can report six neighbours by name.

The control multiframe. On a BCCH carrier's downlink, 5 frames carry the FCCH, 5 the SCH, 4 the BCCH and 36 the common control channels, with one idle, repeating every 235.4 ms. A mobile switching on finds the tone within a fraction of a second, reads the frame number from the next SCH, and has the cell's whole description from the BCCH within a second or two.

munotes.in690

GSM Logical Channels and the Frame Hierarchy

The location area. Pageing a 200-cell location area costs 200 paging messages per incoming call; a one-cell area costs one. But a mobile crossing cells six times an hour crosses a 200-cell area only about 0.42 times an hour, against 6 times for a one-cell area, and each crossing is a location update with authentication and database traffic. The two curves cross somewhere in the middle, and that is where operators set it.

Distinctions

Physical channelLogical channel
IsOne slot on one carrier, every frameA kind of information flow
FixedYes, by the frame structureNo: mapped onto physical channels by a timetable
ExampleSlot 3 of carrier 62BCCH, SDCCH, TCH/F
FamilyDirectionChannelsFor
TrafficBothTCH/F, TCH/HSpeech and data
BroadcastDownlinkBCCH, FCCH, SCHFinding, tuning to and learning about a cell
Common controlBothPCH (down), RACH (up), AGCH (down)Reaching a mobile with no channel yet
Dedicated controlBothSDCCH, SACCH, FACCHSignalling for one mobile
SDCCHSACCHFACCH
WhenBefore a traffic channel is assignedAlways, beside a TCH or SDCCHOnly when needed
CapacitySmall, eight in one traffic channel's spaceSmall, one frame in 26A stolen traffic frame
CarriesLocation update, authentication, SMS, call setupMeasurements, power and timing commandsUrgent signalling, chiefly handover
LevelFramesDurationWhy
TDMA frame14.615 ms8 slots
26-multiframe26120.00 msTraffic: 24 speech, 1 SACCH, 1 idle
51-multiframe51235.38 msBroadcast, common and dedicated control
Superframe1,3266.120 sWhere 26 and 51 line up
Hyperframe2,715,6483.48 hoursCiphering needs a long, non-repeating frame number

What it does not mean

A logical channel is not a frequency. It is a use of slots, defined by the frame number.

The idle frame is not spare capacity. It is when the mobile measures its neighbours; without it there would be no handover.

FACCH does not have a channel. It takes traffic frames when it must, and the codec covers the gap.

The hyperframe is not a timing requirement. Nothing in the radio needs 3.48 hours; the ciphering does.

A bigger location area is not simply better. It cuts updates and multiplies paging; the choice is a balance.

Quick revision

  • Physical channel = one slot on one carrier. Logical channel = a job, mapped by the frame structure.
  • Traffic: TCH/F, TCH/H. Broadcast (downlink): FCCH (tone), SCH (frame number, BSIC), BCCH (cell data). Common control: PCH (down), RACH (up, slotted ALOHA), AGCH (down). Dedicated: SDCCH, SACCH (always, frame 12), FACCH (steals traffic frames).
  • Hierarchy: frame 4.615 ms; 26-multiframe 120.00 ms (24 traffic, 1 SACCH, 1 idle, 92.3 per cent traffic); 51-multiframe 235.38 ms (5 FCCH, 5 SCH, 4 BCCH, 36 CCCH, 1 idle); superframe 1,326 frames = 6.120 s; hyperframe 2,048 superframes = 2,715,648 frames = 3.48 hours, needed because FN is an input to the ciphering.
  • FN_MAX = 2715647: one less than the 2,715,648 frames of a hyperframe.
  • Location area: big means much paging, small means many updates.
munotes.in691

GSM Logical Channels and the Frame Hierarchy

Test yourself

1. Distinguish a physical channel from a logical channel in GSM. A physical channel is one time slot on one carrier, recurring in every TDMA frame; it is a fixed quantity of capacity. A logical channel is a kind of information flow, such as a traffic channel or a paging channel, and it is realised by using particular frames of a physical channel according to the multiframe timetable. Many logical channels therefore share the same physical channel at different moments, and a mobile that knows the frame number knows which logical channel each slot carries.

2. List the GSM logical channels by family. Traffic channels: TCH/F full rate and TCH/H half rate. Broadcast channels, downlink only: the frequency correction channel, the synchronisation channel and the broadcast control channel. Common control channels: the paging channel downlink, the random access channel uplink and the access grant channel downlink. Dedicated control channels: the stand-alone dedicated control channel, the slow associated control channel and the fast associated control channel.

3. Describe the frame hierarchy and give the duration of each level. Eight slots make a TDMA frame of 4.615 ms. Twenty-six frames make a traffic multiframe of 120 ms, and fifty-one frames make a control multiframe of 235.38 ms. A superframe is 26 times 51 frames, 1,326 frames or 6.12 seconds, which is 51 traffic multiframes or 26 control multiframes, the point at which both cycles realign. A hyperframe is 2,048 superframes, 2,715,648 frames, about 3.48 hours, and the frame number runs from 0 to 2,715,647.

4. Why is the hyperframe so long? Because the frame number is an input to the ciphering algorithm, so the keystream repeats when the frame number repeats. A cycle of about three and a half hours ensures that no call can ever see the same frame number twice with the same key, which would let an attacker combine two ciphertexts.

5. How is the 26-multiframe divided, and what is the idle frame for? Twenty-four of its frames carry the traffic channel, one, frame 12, carries the slow associated control channel with the measurement reports and the power and timing commands, and one, frame 25, is idle. The idle frame gives the mobile time to retune to a neighbouring cell's broadcast carrier and read its synchronisation channel, which is how it identifies and measures the neighbours it reports; without it, handover decisions could not be made.

munotes.in692

GSM Logical Channels and the Frame Hierarchy

6. What is the FACCH, and why does it exist? The fast associated control channel is a dedicated control channel with no frames of its own: when urgent signalling is needed, typically a handover command, it takes over, or steals, frames that would have carried traffic, marking them so the receiver does not decode them as speech. It exists because the slow associated control channel, with one frame in 26, is far too slow for a handover command, and because reserving fast capacity permanently would waste it; the interruption to speech is short enough for the codec to conceal.

7. What trade does the size of a location area represent? Every incoming call is paged in every cell of the location area, so a large area multiplies paging traffic by the number of cells. Every crossing of a location area boundary causes a location update, with signalling and database transactions, so a small area multiplies updating traffic, since mobiles cross boundaries more often. The operator sizes the location area where the two costs balance, which depends on how many calls arrive and how fast the subscribers move.

Contents This chapter on its own page

munotes.in693

Chapter Ninety-One

GSM Protocols

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: GSM: Protocols"

In one line

GSM's signalling is a stack of five layers whose interest is not what each does but where each ends: the physical layer and LAPDm stop at the base transceiver station, radio resource management stops at the base station controller, and mobility management and connection management run all the way from the handset to the mobile switching centre, relayed untouched by everything in between.

In the wording a student can write in an examination: on the Um interface the protocol stack is, from the bottom: the physical layer (bursts, channel coding, interleaving, ciphering, the material of [The GSM Radio Interface: Carriers, the TDMA Frame and Bursts]); LAPDm, a data link protocol derived from LAPD, which frames layer-3 messages, in acknowledged or unacknowledged mode, with no CRC of its own because the physical layer already protects; and layer 3, which has three sublayers:

  • Radio resource management (RR): assigning and releasing channels, changing power and timing, measurement reporting, handover, ciphering start, paging. RR runs between the MS and the BSC, with part of it in the BTS.
  • Mobility management (MM): location updating, attach and detach, authentication, identity requests and the allocation of the TMSI. MM runs between the MS and the MSC/VLR.
  • Connection management (CM): call control (CC), the short message service (SMS) and supplementary services (SS). CM runs between the MS and the MSC.

On Abis (BTS to BSC) the layers are the physical layer, LAPD and the BTS management part of RR. On A (BSC to MSC) they are MTP levels 1 to 3 and SCCP carrying BSSAP, which splits into BSSMAP, the part the BSC and MSC exchange with each other, and DTAP, the part that belongs to MM and CM and is relayed transparently through the BSC. Between the core entities, MAP over SCCP and MTP carries the queries between MSC, VLR, HLR, AuC and EIR.

The stack, and where each layer ends

A diagram of the GSM protocol stack across four columns, MS, BTS, BSC and MSC, with the interfaces Um, Abis and A between them. At the bottom of each column is a physical layer. On Um the MS and BTS share LAPDm above the physical layer; on Abis the BTS and BSC share LAPD; on A the BSC and MSC share MTP, SCCP and BSSAP. Above these, a radio resource layer spans the MS, the BTS in part, and the BSC; above that, mobility management and connection management layers span the MS and the MSC alone, drawn as passing straight through the BTS and BSC

Figure 91.1 The GSM protocol stack, and where each layer terminates

Reading the diagram from the top is the way to remember it:

  • CM and MM are a conversation between the handset and the MSC. The BTS and the BSC carry those messages and do not interpret them, which is what DTAP means on the A interface: TS 23.009 says of an inter-MSC handover that "The DTAP signalling is relayed transparently by MSC-B between MSC-A and the MS", and the same transparency applies at the BSC.
  • RR is a conversation between the handset and the BSC, because the BSC is what owns the radio resources: it allocates the channels, orders the handovers and controls the power. The BTS executes parts of it.
  • LAPDm and the physical layer are a conversation between the handset and the BTS, because they are about the air itself.
munotes.in694

GSM Protocols

That layering is why a BSC can be replaced without the MSC noticing, and why a handover between cells of one BSC never involves the MSC at all ([Handover in GSM]).

Layer by layer

The physical layer on Um turns bits into bursts: channel coding (convolutional coding and a CRC on the important bits), interleaving across eight bursts so that a fade damages a little of many blocks rather than all of one, ciphering, and modulation. Its job is to deliver 456-bit blocks with as few errors as the channel allows.

LAPDm is the data link layer on Um, a stripped-down LAPD. It frames layer-3 messages, distinguishes unacknowledged operation (used for broadcast and for messages where speed matters) from acknowledged operation (used where reliability matters), and segments long messages across frames. It carries no frame check sequence of its own, because the physical layer's coding already detects errors, and its address field is short, because there are only a few logical links on one channel.

RR is the busiest layer during a call's life and is invisible to the user. It assigns an SDCCH after a random access, assigns a traffic channel when a call is set up, starts ciphering, collects the measurement reports on the SACCH, decides and executes handovers, and releases everything afterwards.

MM deals with who and where: location updating when a mobile enters a new location area, IMSI attach and detach, authentication ([GSM Security]), the allocation of a TMSI so that the permanent identity is not sent in clear, and identity requests.

CM is what the user asked for: call control (setup, alerting, connect, disconnect), SMS, and the supplementary services of [GSM and Its Mobile Services].

On the fixed side

Abis carries LAPD frames over 64 kbit/s links, with traffic channels at 16 kbit/s (four to a 64 kbit/s timeslot, since GSM speech is 13 kbit/s). Part of RR lives here as BTS management.

A carries the real signalling load. MTP levels 1 to 3 are SS7's transport: the physical links, error control and routing. SCCP adds connections and the ability to address by global title, which is how a query can be addressed to whichever HLR serves an IMSI, without knowing which machine that is. BSSAP rides on SCCP and splits in two: BSSMAP is the BSC and the MSC talking about resources, handovers and paging; DTAP is the MS and the MSC talking, with the BSC as a postman.

MAP is the protocol of the core registers: the VLR asking the HLR for a subscriber's data, the HLR asking the VLR for a roaming number, the MSC asking the EIR about a handset, the HLR fetching triplets from the AuC. It too runs over SCCP and MTP, which is why a GSM core network is, underneath, a telephone signalling network.

munotes.in695

GSM Protocols

The protocols, computed

The program tabulates which layer terminates in which entity; follows one location updating request from the handset to the MSC, adding each envelope; states the question each layer answers; and lists the core signalling stack.

# GSM's protocol stack: which layer terminates where, what one message crosses,
# and what the layers cost in bits.
print("Where each protocol layer ends, across the three interfaces:")
print("  %-28s %5s %6s %6s %6s  %s" % ("layer", "MS", "BTS", "BSC", "MSC", "what it does"))
rows = [("CM (call control, SMS, SS)", "yes", ".", ".", "yes", "sets up and clears calls, carries SMS"),
        ("MM (mobility management)", "yes", ".", ".", "yes", "location updating, authentication, identity"),
        ("RR (radio resources)", "yes", "part", "yes", ".", "channels, handover, measurements"),
        ("LAPDm (data link, Um)", "yes", "yes", ".", ".", "frames, acknowledged or not, on the air"),
        ("physical (Um)", "yes", "yes", ".", ".", "bursts, coding, interleaving, ciphering")]
for name, ms, bts, bsc, msc, what in rows:
    print("  %-28s %5s %6s %6s %6s  %s" % (name, ms, bts, bsc, msc, what))
print("  CM and MM messages pass through the BTS and BSC untouched: they are relayed, not read.")

# 2. One message, and the headers it gathers on the way.
print("\nA LOCATION UPDATING REQUEST, and what wraps it at each step:")
layers = [("the message itself (CM/MM, layer 3)", 20),
          ("LAPDm header and length (Um)", 3),
          ("channel coding and interleaving (Um)", 0),
          ("on the air: 4 bursts of 114 payload bits", 0),
          ("Abis: LAPD frame to the BSC", 6),
          ("A: BSSAP over SCCP over MTP to the MSC", 14)]
total = 0
for name, overhead in layers:
    total += overhead
    print("  %-42s adds %2d octets, running total %2d" % (name, overhead, total))
print("  the same 20 octets of meaning arrive at the MSC inside %d octets of envelopes" % total)

# 3. What the layers are for, as a list of the questions each answers.
print("\nEach layer answers one question:")
for layer, question in (("physical", "how do these bits survive the air?"),
                        ("LAPDm", "did this frame arrive, and in order?"),
                        ("RR", "which channel should this mobile use, and where should it go next?"),
                        ("MM", "who is this subscriber, where are they, and are they genuine?"),
                        ("CM", "what does the user want: a call, a message, a service?")):
    print("  %-9s %s" % (layer, question))

# 4. On the network side: the signalling that carries MM and CM beyond the MSC.
print("\nBeyond the MSC, the core signalling stack:")
for layer, what in (("MTP levels 1 to 3", "the SS7 transport: links, error control, routing"),
                    ("SCCP", "connections and global title addressing over MTP"),
                    ("BSSAP", "split into BSSMAP (BSC and MSC talking) and DTAP (relayed to the MS)"),
                    ("MAP", "MSC, VLR, HLR, AuC and EIR talking to each other")):
    print("  %-18s %s" % (layer, what))
print("  DTAP is the part of BSSAP that the BSC does not read: it belongs to MM and CM.")
munotes.in696

GSM Protocols

Where each protocol layer ends, across the three interfaces:
  layer                           MS    BTS    BSC    MSC  what it does
  CM (call control, SMS, SS)     yes      .      .    yes  sets up and clears calls, carries SMS
  MM (mobility management)       yes      .      .    yes  location updating, authentication, identity
  RR (radio resources)           yes   part    yes      .  channels, handover, measurements
  LAPDm (data link, Um)          yes    yes      .      .  frames, acknowledged or not, on the air
  physical (Um)                  yes    yes      .      .  bursts, coding, interleaving, ciphering
  CM and MM messages pass through the BTS and BSC untouched: they are relayed, not read.

A LOCATION UPDATING REQUEST, and what wraps it at each step:
  the message itself (CM/MM, layer 3)        adds 20 octets, running total 20
  LAPDm header and length (Um)               adds  3 octets, running total 23
  channel coding and interleaving (Um)       adds  0 octets, running total 23
  on the air: 4 bursts of 114 payload bits   adds  0 octets, running total 23
  Abis: LAPD frame to the BSC                adds  6 octets, running total 29
  A: BSSAP over SCCP over MTP to the MSC     adds 14 octets, running total 43
  the same 20 octets of meaning arrive at the MSC inside 43 octets of envelopes

Each layer answers one question:
  physical  how do these bits survive the air?
  LAPDm     did this frame arrive, and in order?
  RR        which channel should this mobile use, and where should it go next?
  MM        who is this subscriber, where are they, and are they genuine?
  CM        what does the user want: a call, a message, a service?

Beyond the MSC, the core signalling stack:
  MTP levels 1 to 3  the SS7 transport: links, error control, routing
  SCCP               connections and global title addressing over MTP
  BSSAP              split into BSSMAP (BSC and MSC talking) and DTAP (relayed to the MS)
  MAP                MSC, VLR, HLR, AuC and EIR talking to each other
  DTAP is the part of BSSAP that the BSC does not read: it belongs to MM and CM.

Where things end. The table is the examination answer in one picture: CM and MM at the MS and the MSC, RR at the MS and the BSC with part in the BTS, LAPDm and the physical layer at the MS and the BTS.

What a message costs. Twenty octets of meaning, a location updating request, arrive at the MSC inside about 43 octets once LAPDm, the Abis frame and the BSSAP, SCCP and MTP headers are counted, and that is before the channel coding that doubles the bits on the air. Signalling is not cheap, which is why GSM sends it on the small SDCCH rather than on a traffic channel, and why the location area size of [GSM Logical Channels and the Frame Hierarchy] matters.

munotes.in697

GSM Protocols

One question each. Physical: how do these bits survive the air? LAPDm: did this frame arrive, in order? RR: which channel, and where next? MM: who is this, where are they, are they genuine? CM: what does the user want? A protocol architecture is well designed when each layer has exactly one such question, and GSM's does.

Distinctions

LayerBetweenTerminates atConcerned with
Physical (Um)MS and BTSBTSBursts, coding, interleaving, ciphering
LAPDmMS and BTSBTSFraming, acknowledged and unacknowledged transfer
RRMS and BSCBSC (part in the BTS)Channels, power, measurements, handover
MMMS and MSCMSCLocation updating, authentication, identities
CMMS and MSCMSCCall control, SMS, supplementary services
BSSMAPDTAP
BetweenBSC and MSCMS and MSC
Read by the BSCYesNo: relayed transparently
CarriesResource assignment, handover, pagingMM and CM messages
Part ofBSSAP, over SCCP and MTPThe same
UmAbisA
JoinsMS and BTSBTS and BSCBSC and MSC
Data linkLAPDmLAPDMTP and SCCP
ApplicationRR, and MM and CM passing throughBTS management, plus relayingBSSAP: BSSMAP and DTAP
Traffic22.8 kbit/s on the air16 kbit/s per channel64 kbit/s per channel

What it does not mean

The BTS is not a mere repeater. It terminates the physical layer and LAPDm and executes parts of RR; what it does not do is read MM or CM.

LAPDm is not LAPD. It is a simplified version for the air: no frame check sequence, short addresses, and segmentation suited to the small frames of a GSM channel.

DTAP is not a protocol of the BSC. It is the MS-to-MSC conversation that the BSC relays; BSSMAP is the BSC's own.

RR does not end at the BTS. The BTS holds part of it; the decisions belong to the BSC.

MAP is not a GSM radio protocol. It is the signalling between core databases, carried by the same SS7 machinery that carries telephone signalling everywhere.

Quick revision

  • Um stack: physical (bursts, coding, interleaving, ciphering), LAPDm (framing, acknowledged and unacknowledged), layer 3: RR, MM, CM.
  • Termination: physical and LAPDm at the BTS; RR at the BSC (part in the BTS); MM and CM at the MSC, relayed transparently by the BTS and BSC.
  • RR: channels, power and timing, measurement reports, handover, ciphering start, paging. MM: location updating, attach and detach, authentication, TMSI. CM: call control, SMS, supplementary services.
  • Abis: physical, LAPD, BTS management. A: MTP 1 to 3, SCCP, BSSAP = BSSMAP (BSC and MSC) + DTAP (relayed, MS to MSC).
  • MAP over SCCP and MTP: MSC, VLR, HLR, AuC, EIR.
  • Program: 20 octets of message reach the MSC inside about 43 octets of envelopes; each layer answers exactly one question.
munotes.in698

GSM Protocols

Test yourself

1. Draw the GSM protocol stack on the Um interface and say what each layer does. From the bottom: the physical layer, which performs channel coding, interleaving, ciphering and modulation and delivers blocks in bursts; LAPDm, the data link layer, which frames layer-3 messages and offers acknowledged and unacknowledged transfer with segmentation, carrying no frame check sequence of its own because the physical layer's coding detects errors; and layer 3, divided into radio resource management, which handles channels, power, measurements and handover, mobility management, which handles location updating, authentication and identities, and connection management, which handles call control, the short message service and supplementary services.

2. Where does each layer terminate, and why does it matter? The physical layer and LAPDm terminate at the BTS, because they concern the air interface itself. Radio resource management terminates at the BSC, with part of it executed in the BTS, because the BSC owns the radio resources and takes the decisions about them. Mobility management and connection management terminate at the MSC, and the BTS and BSC relay them without reading them. It matters because it fixes who must be involved in what: a handover between two cells of one BSC needs no MSC, while an authentication involves the MSC and the VLR but no radio decision.

3. What is the difference between BSSMAP and DTAP? Both are parts of BSSAP, carried over SCCP and MTP on the A interface. BSSMAP is the conversation between the BSC and the MSC themselves, about assigning and releasing resources, paging and handovers. DTAP is the conversation between the mobile station and the MSC, that is the mobility management and connection management messages, which the BSC relays transparently without interpreting; 3GPP TS 23.009 makes the same point about an inter-MSC handover, where the DTAP signalling is relayed transparently by the second MSC between the controlling MSC and the mobile.

4. What protocols run on the A interface, and why is SS7 used? Message transfer part levels 1 to 3 provide the links, error control and routing; the signalling connection control part adds connections and global title addressing; BSSAP rides on top and splits into BSSMAP and DTAP. SS7 is used because the GSM core is a telephone network: its switches already spoke SS7 to each other and to the fixed network, and global title addressing lets a query be routed to the register responsible for a subscriber without knowing which machine holds it.

munotes.in699

GSM Protocols

5. What is MAP and what does it carry? The mobile application part is the protocol by which the core entities talk: the VLR asking the HLR for a subscriber's data when the subscriber arrives, the HLR asking the VLR for a roaming number when a call must be routed, the MSC asking the EIR whether a handset is barred, and the HLR obtaining authentication data from the AuC. It runs over SCCP and MTP, the same signalling transport as the A interface uses.

6. Why does a 20-octet signalling message occupy far more than 20 octets by the time it reaches the MSC? Because each layer adds its own envelope: LAPDm frames it for the air, the physical layer adds channel coding and interleaving that roughly double the bits transmitted, the Abis interface wraps it in a LAPD frame, and the A interface wraps it in BSSAP, SCCP and MTP headers. In the chapter's illustration the 20 octets of meaning arrived inside about 43 octets of headers, before the coding on the air was counted, which is why GSM sends signalling on the small SDCCH rather than on a traffic channel.

Contents This chapter on its own page

munotes.in700

Chapter Ninety-Two

Localization and Calling in GSM

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: GSM: Localization and calling"

In one line

A GSM network never knows exactly where a subscriber is, and does not try to: the mobile reports only when it crosses a location area, its home register records only which visitor register is looking after it, and when a call comes the home register asks that visitor register for a temporary number to route to, after which the whole location area is paged.

In the wording a student can write in an examination: the identities are:

  • IMSI, the international mobile subscriber identity, stored on the SIM and at most 15 digits: "IMSI is composed of three parts: 1) Mobile Country Code (MCC) consisting of three digits. The MCC identifies uniquely the country of domicile of the mobile subscription; 2) Mobile Network Code (MNC) consisting of two or three digits ... 3) Mobile Subscriber Identification Number (MSIN) identifying the mobile subscription within a PLMN".
  • MSISDN, the number a caller dials, which says nothing about where the subscriber is.
  • TMSI, a temporary identity allocated by the VLR and meaningful only there, used instead of the IMSI on the air so that a subscriber cannot be tracked.
  • MSRN, the mobile station roaming number, a temporary number from the visited network's range, used to route one call.
  • LAI, the location area identification: "MCC, MNC, LAC".
  • IMEI, which identifies the equipment, not the subscription.

Location updating keeps the registers current: normal updating when the mobile enters a new location area, periodic updating on a timer so that the network can tell a switched-off mobile from a silent one, and IMSI attach and detach when the mobile is switched on and off. A mobile terminated call is routed by interrogation: the GMSC asks the HLR for the subscriber; the HLR asks the serving VLR for an MSRN; the call is routed to that MSC, which pages the mobile in its location area. A mobile originated call begins with a RACH request, an SDCCH, authentication and ciphering, and then a traffic channel.

The identities, and why there are so many

Each identity exists because a different party needs to name the subscriber for a different purpose, and the separations are deliberate.

  • The IMSI is the true name, known to the SIM and the home network. It is sent on the air as rarely as possible.
  • The MSISDN is the public name. Its digits are a telephone number in the ordinary numbering plan, which is what makes the subscriber reachable from any telephone in the world, and it deliberately carries no location.
  • The TMSI is a local alias. The VLR allocates one, sends it ciphered, and changes it from time to time, so that an eavesdropper who hears a page cannot link it to a person.
  • The MSRN is a routing alias. It looks like an ordinary number in the visited network so that the fixed network can route to it, and it lives only for the duration of one call setup.
  • The LAI names the location area, and is broadcast so that a mobile can tell when it has moved into a new one.
  • The IMEI names the handset, and belongs to the EIR ([The GSM System Architecture]).
munotes.in701

Localization and Calling in GSM

Location updating

The network's knowledge of where a subscriber is has exactly two levels: the HLR knows which VLR, and the VLR knows which location area. Nothing knows the cell until the mobile answers a page.

  • Normal location updating happens when a mobile, listening to the BCCH, sees a location area identity different from the one it has stored. It requests a channel, identifies itself (by TMSI if it can), is authenticated, and the VLR records the new location area. If the new location area belongs to a different VLR, that VLR fetches the subscriber's data from the HLR, the HLR updates its record and tells the old VLR to delete its copy.
  • Periodic location updating happens on a timer broadcast by the network. Without it, a mobile that has gone out of coverage or lost its battery would still appear present, and every call to it would cost a fruitless paging of the whole location area.
  • IMSI attach and detach mark switching on and off. Detach is a courtesy that saves the network from paging a phone that is off.

The cost of all this is signalling, and the size of the location area sets the balance, as [GSM Logical Channels and the Frame Hierarchy] computed.

A mobile terminated call

This is the sequence the examination asks for, and every step exists because of what the previous step did not know.

  1. The caller dials the MSISDN. The fixed network routes the call to the subscriber's home network, because that is all the number says.
  2. A GMSC receives it and interrogates the HLR, as [The GSM System Architecture] quoted: an MSC that cannot be reached directly "will interrogate the appropriate HLR and then route the call to the MSC where the mobile station is located."
  3. The HLR finds the IMSI and the VLR currently serving it, and asks that VLR for a roaming number.
  4. The VLR allocates an MSRN and returns it.
  5. The HLR passes the MSRN to the GMSC.
  6. The GMSC routes the call to the MSC that owns that MSRN, using ordinary telephone routing.
  7. That MSC asks its VLR for the subscriber's location area and whether the service is allowed.
  8. The MSC pages the mobile, by TMSI, in every cell of the location area.
  9. The mobile answers on the RACH and is assigned a signalling channel on the AGCH.
  10. Authentication and ciphering follow ([GSM Security]), a traffic channel is assigned, and the phone rings.
munotes.in702

Localization and Calling in GSM

A mobile originated call

Shorter, because the mobile knows where it is.

  1. The mobile sends a channel request on the RACH and is given an SDCCH.
  2. It sends a service request with its TMSI; the VLR authenticates it.
  3. Ciphering starts, and the mobile sends the dialled number.
  4. The MSC checks with the VLR that the subscriber may make this call (barring, credit, service subscription).
  5. The MSC routes the call toward the destination and assigns a traffic channel.
  6. The destination rings and, on answer, the call is connected.

Localization and calling, computed

The program takes the identities apart, states who holds each, and walks the two call flows and the four kinds of location update.

# The identities GSM uses to find a subscriber, and the two calls, step by step.
print("The identities (TS 23.003), taken apart:")
imsi = "404451234567890"                      # an Indian example: MCC 404, MNC 45
print("  IMSI  %s  MCC %s (country) MNC %s (network) MSIN %s (the subscriber)"
      % (imsi, imsi[:3], imsi[3:5], imsi[5:]))
print("        %d digits, at most 15; stored on the SIM and never sent in clear when it can be helped"
      % len(imsi))
msisdn = "919876543210"
print("  MSISDN +%s  country code %s, then the national number: what a caller dials"
      % (msisdn, msisdn[:2]))
lai = ("404", "45", "0x1A2B")
print("  LAI   MCC %s MNC %s LAC %s: which location area the mobile is in" % lai)
print("  TMSI  4 octets, meaningful only inside one VLR, changed often so that the IMSI is not tracked")
print("  MSRN  a temporary number in the visited network's range, used to route one call")
print("  IMEI  identifies the handset, not the subscription: the EIR's business")

print("\nWho holds which identity:")
for who, holds in (("SIM", "IMSI, the key Ki, the TMSI in use"),
                   ("handset", "IMEI"),
                   ("HLR", "IMSI, MSISDN, services, the address of the current VLR"),
                   ("VLR", "IMSI, TMSI, MSISDN, the current LAI, and an MSRN while a call is routed"),
                   ("the caller", "only the MSISDN")):
    print("  %-9s %s" % (who, holds))

print("\nA mobile terminated call (someone rings the handset):")
for i, step in enumerate([
        "the caller dials the MSISDN; the fixed network routes it to the subscriber's home network",
        "a GMSC receives the call and interrogates the HLR with the MSISDN",
        "the HLR looks up the IMSI and the VLR the subscriber is at, and asks that VLR for a roaming number",
        "the VLR allocates an MSRN and returns it",
        "the HLR gives the MSRN to the GMSC",
        "the GMSC routes the call to the MSC that owns the MSRN",
        "that MSC asks its VLR: which location area, and is this service allowed?",
        "the MSC pages the mobile in every cell of the location area, using the TMSI",
        "the mobile answers on the RACH; it is given a channel on the AGCH",
        "authentication and ciphering; then a traffic channel is assigned and the phone rings"], 1):
    print("  %2d. %s" % (i, step))

print("\nA mobile originated call (the user dials):")
for i, step in enumerate([
        "the mobile requests a channel on the RACH and is given an SDCCH",
        "it sends a service request with its TMSI; the VLR authenticates it",
        "ciphering starts; the mobile sends the dialled number",
        "the MSC checks with the VLR that the subscriber may make this call",
        "the MSC routes the call toward the destination and assigns a traffic channel",
        "the destination rings; when it answers, the call is connected"], 1):
    print("  %2d. %s" % (i, step))

print("\nLocation updating, and what each kind costs:")
for kind, when, touches in (("normal", "the mobile enters a new location area", "VLR, and the HLR if the VLR changed"),
                            ("periodic", "a timer expires, to prove the mobile is still there", "VLR"),
                            ("IMSI attach", "the mobile is switched on", "VLR, HLR"),
                            ("IMSI detach", "the mobile is switched off", "VLR")):
    print("  %-12s %-46s touches %s" % (kind, when, touches))
print("  without IMSI detach the network would page a switched-off phone and wait for the timeout each time")
munotes.in703

Localization and Calling in GSM

The identities (TS 23.003), taken apart:
  IMSI  404451234567890  MCC 404 (country) MNC 45 (network) MSIN 1234567890 (the subscriber)
        15 digits, at most 15; stored on the SIM and never sent in clear when it can be helped
  MSISDN +919876543210  country code 91, then the national number: what a caller dials
  LAI   MCC 404 MNC 45 LAC 0x1A2B: which location area the mobile is in
  TMSI  4 octets, meaningful only inside one VLR, changed often so that the IMSI is not tracked
  MSRN  a temporary number in the visited network's range, used to route one call
  IMEI  identifies the handset, not the subscription: the EIR's business

Who holds which identity:
  SIM       IMSI, the key Ki, the TMSI in use
  handset   IMEI
  HLR       IMSI, MSISDN, services, the address of the current VLR
  VLR       IMSI, TMSI, MSISDN, the current LAI, and an MSRN while a call is routed
  the caller only the MSISDN

A mobile terminated call (someone rings the handset):
   1. the caller dials the MSISDN; the fixed network routes it to the subscriber's home network
   2. a GMSC receives the call and interrogates the HLR with the MSISDN
   3. the HLR looks up the IMSI and the VLR the subscriber is at, and asks that VLR for a roaming number
   4. the VLR allocates an MSRN and returns it
   5. the HLR gives the MSRN to the GMSC
   6. the GMSC routes the call to the MSC that owns the MSRN
   7. that MSC asks its VLR: which location area, and is this service allowed?
   8. the MSC pages the mobile in every cell of the location area, using the TMSI
   9. the mobile answers on the RACH; it is given a channel on the AGCH
  10. authentication and ciphering; then a traffic channel is assigned and the phone rings

A mobile originated call (the user dials):
   1. the mobile requests a channel on the RACH and is given an SDCCH
   2. it sends a service request with its TMSI; the VLR authenticates it
   3. ciphering starts; the mobile sends the dialled number
   4. the MSC checks with the VLR that the subscriber may make this call
   5. the MSC routes the call toward the destination and assigns a traffic channel
   6. the destination rings; when it answers, the call is connected

Location updating, and what each kind costs:
  normal       the mobile enters a new location area          touches VLR, and the HLR if the VLR changed
  periodic     a timer expires, to prove the mobile is still there touches VLR
  IMSI attach  the mobile is switched on                      touches VLR, HLR
  IMSI detach  the mobile is switched off                     touches VLR
  without IMSI detach the network would page a switched-off phone and wait for the timeout each time
munotes.in704

Localization and Calling in GSM

The identities. The example IMSI splits into MCC 404, MNC 45 and a ten-digit MSIN: a country, a network and a subscriber, in that order, which is exactly the order in which a query is routed. The MSISDN, by contrast, is a telephone number and nothing more; it is the pair of them, joined only in the HLR, that makes a mobile reachable.

Who holds what. The SIM holds the IMSI and the key; the handset holds the IMEI; the HLR holds the subscription and which VLR has it; the VLR holds the working copy and the location area; the caller holds only the MSISDN. Every step of a call is a consequence of that distribution: the call must reach the HLR because only the HLR knows the VLR, and it must reach the VLR because only the VLR knows the location area.

The two calls. A mobile terminated call takes ten steps and two database interrogations before the phone rings; a mobile originated call takes six and none. That asymmetry is why incoming calls take noticeably longer to connect, and why the MSRN, a number allocated and released for one call, exists at all.

munotes.in705

Localization and Calling in GSM

Location updating. Normal updating touches the VLR, and the HLR too if the VLR has changed; periodic updating touches only the VLR but happens whether or not anything has moved; attach and detach bracket the mobile's day. Without detach, every call to a switched-off phone would page a whole location area and wait for a timeout.

Distinctions

IdentityNamesLivesKnown to
IMSIThe subscriptionPermanentlySIM, HLR, VLR (not sent on air when avoidable)
MSISDNThe subscriber, publiclyPermanentlyEveryone; the HLR maps it to the IMSI
TMSIThe subscription, locallyUntil the VLR changes itThe mobile and one VLR
MSRNA call's routeOne call setupHLR, GMSC, the visited MSC
LAIA location areaAs long as the area existsBroadcast on the BCCH
IMEIThe handsetPermanentlyThe handset and the EIR
The HLR knowsThe VLR knowsThe network learns on paging
GranularityWhich VLRWhich location areaWhich cell
Updated byLocation updating that changes VLREvery location updateThe mobile's answer
Cost of finer granularityMore HLR trafficMore location updatesNothing: it is per call
Mobile terminatedMobile originated
Starts atThe caller, with an MSISDNThe mobile, on the RACH
Database queriesHLR and VLR, for an MSRNVLR only, for authentication and rights
PagingYes, the whole location areaNo
Steps in the program106

What it does not mean

The network does not know which cell a mobile is in. It knows the location area, and finds the cell by paging.

The MSISDN does not identify the SIM. The HLR maps between them, which is why a number can be moved to a new SIM.

The MSRN is not the subscriber's number. It is a temporary routing number in the visited network, released after the call is set up.

A TMSI is not secret. It is temporary and local; its purpose is to stop an eavesdropper linking transmissions to a person, not to authenticate anyone.

Periodic updating is not wasted traffic. It is what lets the network distinguish a mobile that is switched off from one that is merely silent.

Quick revision

  • IMSI = MCC (3) + MNC (2 or 3) + MSIN, at most 15 digits, on the SIM. MSISDN: the dialled number. TMSI: temporary, per VLR. MSRN: temporary routing number for one call. LAI = MCC + MNC + LAC. IMEI: the handset.
  • The HLR knows the VLR; the VLR knows the location area; paging finds the cell.
  • Location updating: normal (new location area), periodic (timer), IMSI attach and detach.
  • Mobile terminated call: MSISDN to the home network, GMSC interrogates the HLR, HLR asks the VLR for an MSRN, GMSC routes to that MSC, MSC asks the VLR, pages the location area by TMSI, mobile answers on RACH, channel on AGCH, authentication and ciphering, traffic channel, ring.
  • Mobile originated call: RACH, SDCCH, service request with TMSI, authentication, ciphering, dialled number, checks, traffic channel, connect.
  • Program: 10 steps and two interrogations for an incoming call against 6 and none for an outgoing one.
munotes.in706

Localization and Calling in GSM

Test yourself

1. Name the identities used in GSM and say what each is for. The IMSI identifies the subscription, is stored on the SIM and consists of a three-digit mobile country code, a two or three digit mobile network code and a mobile subscriber identification number, at most fifteen digits in all. The MSISDN is the number a caller dials and carries no location information. The TMSI is a temporary identity allocated by a VLR, used on the air instead of the IMSI so that a subscriber cannot be tracked. The MSRN is a temporary number from the visited network used to route one incoming call. The LAI, made of the mobile country code, the mobile network code and a location area code, identifies a location area and is broadcast. The IMEI identifies the handset rather than the subscription and is checked against the equipment identity register.

2. Describe location updating and its kinds. A mobile listens to the broadcast channel and compares the location area identity with the one it has stored; when they differ, it requests a channel and performs a normal location update, in which it is authenticated and the VLR records the new location area, fetching the subscriber's data from the HLR and causing the old VLR's copy to be deleted if the VLR has changed. Periodic location updating repeats on a timer so that the network can tell a mobile that has lost coverage or power from one that is merely idle. IMSI attach and IMSI detach are performed when the mobile is switched on and off, so that the network does not page a mobile that is known to be off.

3. Explain step by step how an incoming call reaches a roaming subscriber. The caller dials the MSISDN and the fixed network routes the call to the subscriber's home network, where a gateway MSC receives it. The GMSC interrogates the HLR with the MSISDN; the HLR finds the IMSI and the VLR currently serving the subscriber and asks that VLR for a mobile station roaming number. The VLR allocates an MSRN and returns it through the HLR to the GMSC, which routes the call by ordinary telephone routing to the MSC that owns that number. That MSC consults its VLR for the subscriber's location area and rights, then pages the mobile, using its TMSI, in every cell of the location area. The mobile answers on the random access channel, is granted a signalling channel, is authenticated and ciphering is started, a traffic channel is assigned, and the phone rings.

munotes.in707

Localization and Calling in GSM

4. Why is an MSRN needed at all? Because the fixed network can only route to a telephone number, and the subscriber's own number, the MSISDN, points to the home network, not to wherever the subscriber currently is. The MSRN is a temporary number belonging to the visited network's numbering range, so ordinary telephone routing carries the call to the correct MSC; it is allocated by the VLR when the HLR asks, used for that one call setup, and then released, so that a small pool of numbers serves many visitors.

5. Why does GSM use a TMSI instead of sending the IMSI? Because the IMSI is the permanent identity of the subscription, and anyone who could hear it on the air could follow that subscriber from place to place and link their transmissions. The VLR therefore allocates a temporary identity, sends it under ciphering, and changes it from time to time; only that VLR can map it back to the IMSI. The IMSI is sent in clear only when no TMSI can be used, for example the very first time a mobile registers in a network or after a VLR failure.

6. Compare a mobile terminated and a mobile originated call. A mobile terminated call begins outside the network with only an MSISDN, so it needs two database interrogations, the HLR for the serving VLR and the VLR for a roaming number, then rerouting to the serving MSC and paging of the whole location area before the mobile can even answer: about ten steps. A mobile originated call begins at the mobile, which already knows its cell, so it needs only a random access, a signalling channel, authentication, ciphering and the dialled number before the traffic channel is assigned: about six steps and no paging. That is why incoming calls take longer to connect.

Contents This chapter on its own page

munotes.in708

Chapter Ninety-Three

Handover in GSM

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: GSM: Handover"

In one line

Handover moves a call in progress from one channel to another without the user noticing, and almost all of the difficulty is in the decision rather than the mechanism: the mobile measures its neighbours and reports twice a second, and the network must decide when a neighbour is genuinely better rather than momentarily luckier, which is what hysteresis and a dwell requirement are for.

In the wording a student can write in an examination: handover (or handoff) transfers an ongoing call from one radio channel to another. It is needed because the mobile moves out of a cell, because the signal quality falls through interference or fading even without movement, and because the network wants to balance load between cells. GSM's handover is hard (the old channel is released before the new one is used, break before make) and mobile-assisted (the mobile measures neighbouring cells and reports, but the network decides).

The four types are: intra-cell (a different channel in the same cell, usually because of interference), inter-cell within one BSC, inter-BSC within one MSC, and inter-MSC. The higher the type, the more entities take part; the BSC handles the first two alone.

The measurements come from both ends: the mobile measures the received level and quality of its serving cell and the level of up to six neighbours, using the idle frame of the 26-multiframe, and reports on the SACCH about twice a second; the BTS measures the uplink. The BSC decides, using hysteresis (the neighbour must be better by a margin) and a requirement that the condition persist for several reports, so that a momentary fade does not cause a handover and the mobile does not ping-pong between two cells.

In an inter-MSC handover, MSC-A remains the anchor: it "controls the call and the mobility management of the Mobile during the call, before, during and after a basic or subsequent handover", and the connection from the caller is never rerouted. A later move is a subsequent handover, either back to MSC-A or on to a third MSC.

Why handover, and when

Three reasons trigger it:

  • Movement. The subscriber leaves the cell. This is the obvious one, and in small cells it is frequent: [Cellular Systems: Cells, Clusters and Frequency Reuse] counted 50 handovers an hour for a car in 500 m cells.
  • Quality. The signal may fail without any movement: a fade, an interferer, a lorry parked between the mobile and the base station. An intra-cell handover to another channel of the same cell answers interference; an inter-cell handover answers a failing path.
  • Load. A cell that is full may hand a mobile to a neighbour that can hear it, even though nothing is wrong with the radio link. This is traffic handover, and it is how an operator squeezes a busy cell.
munotes.in709

Handover in GSM

Measurements

GSM is mobile-assisted: the network could never measure what the mobile hears, so the mobile measures for it.

In each 26-multiframe the mobile has an idle frame ([GSM Logical Channels and the Frame Hierarchy]) in which it retunes to a neighbouring BCCH carrier and measures it. Over several multiframes it builds a picture of up to six neighbours, identified by their base station identity codes, and reports their levels, along with the level and quality of its own serving cell, on the SACCH, whose frame comes round every 120 ms, giving a report about twice a second.

The BTS measures the uplink at the same time: level and quality of what it receives, and the timing advance, which is a crude measure of distance.

All of it goes to the BSC, which is where the decision is taken, because the BSC owns the radio resources ([GSM Protocols]).

The decision, and why it is hard

The naive rule, hand over whenever a neighbour is stronger, fails badly, because fading makes the comparison flicker: two cells whose signals differ by a decibel will swap places many times a minute. Each swap costs signalling, a short interruption, and a risk of failure. The symptom has a name: ping-pong.

Two mechanisms fix it:

  • Hysteresis. The neighbour must be better by a margin, typically a few decibels, before a handover is considered, and the margin means the reverse handover needs the same advantage in the other direction.
  • A dwell requirement. The condition must hold for several consecutive reports, so that a single deep fade does not trigger anything.

Both delay the handover, and delay is dangerous: a mobile that clings to a failing cell loses the call. The program measures exactly that trade.

The message sequence

TS 23.009's Figure 7 gives the basic external intra-MSC handover, and the program lists it. The shape is worth learning:

  1. BSS-A decides and asks: "When the BSS (BSS-A), currently supporting the MS, determines that the MS requires to be handed over it will send an A-HANDOVER-REQUIRED message to the MSC (MSC-A)." That message "shall contain a list of cells, or a single cell, to which the MS can be handed over ... in order of preference".
  2. MSC-A asks the target to allocate a channel, and the target acknowledges with the details.
  3. MSC-A tells BSS-A, which tells the mobile to go to the new channel.
  4. The mobile accesses the new cell; the target detects it and reports, and when the mobile completes, the target tells MSC-A.
  5. MSC-A clears the old resources.
munotes.in710

Handover in GSM

The mobile is off the air only between leaving the old channel and being detected on the new one, which is why GSM's handover is heard, at most, as a click.

Inter-MSC handover and the anchor

When the target cell belongs to another MSC, the call could in principle be rerouted from the caller, but that would involve the whole fixed network in a local radio event. GSM instead keeps MSC-A as the anchor: "In the Inter-MSC handover case, MSC-A is the MSC which controls the call and the mobility management of the Mobile during the call, before, during and after a basic or subsequent handover."

A circuit is established from MSC-A to MSC-B, and the call now travels to MSC-A and onward to MSC-B. If the mobile moves again, that is a subsequent handover, and it may go back to MSC-A, in which case the extra leg is released, or on to a third MSC, MSC-B', in which case MSC-A sets up a new leg and releases the old one. The anchor never changes during the call, so a mobile that crosses many MSC areas can end up with a long path, which is the price of not disturbing the caller.

Handover, computed

The program walks a mobile at 15 m/s between two cells 1 km apart, 200 times, with correlated shadowing and a measurement report every 480 ms, and counts the handovers and the dropped calls under five decision rules; lists the four handover types with the entities each involves; and prints the message sequence of an intra-MSC handover.

# Handover: the decision, and why hysteresis and a dwell timer are needed.
import math
import random

# A mobile walks from cell A to cell B. Each cell's signal falls with distance
# (path-loss exponent 3.5) and fades; the mobile reports every 480 ms on the
# SACCH. Three decision rules are compared.
rnd = random.Random(93)
STEP_S, SPEED, SEP = 0.48, 15.0, 1000.0            # a report every 480 ms, 15 m/s, cells 1 km apart

def level(d, shadow):
    return -15 - 35 * math.log10(max(d, 1.0)) + shadow      # dBm, 35 dB a decade

def walk(rule, hysteresis=0.0, dwell=1, trials=200):
    swaps, dropped = 0, 0
    for _ in range(trials):
        serving, count, x, s_a, s_b = 'A', 0, 0.0, 0.0, 0.0
        for _ in range(int(SEP / (SPEED * STEP_S))):
            x += SPEED * STEP_S
            s_a = 0.7 * s_a + rnd.gauss(0, 4)      # slow fading, correlated
            s_b = 0.7 * s_b + rnd.gauss(0, 4)
            a, b = level(x, s_a), level(SEP - x, s_b)
            here, there = (a, b) if serving == 'A' else (b, a)
            if rule == 'none':
                better = there > here
            else:
                better = there > here + hysteresis
            count = count + 1 if better else 0
            if count >= dwell:
                serving = 'B' if serving == 'A' else 'A'
                swaps += 1
                count = 0
            if here < -125:                        # too weak: the call drops
                dropped += 1
                break
    return swaps / trials, 100 * dropped / trials

print("A mobile crossing between two cells 1 km apart at 15 m/s, 200 walks:")
print("  rule                                  handovers per crossing   calls dropped")
for label, rule, hyst, dwell in (("no hysteresis, act at once", 'none', 0.0, 1),
                                 ("hysteresis 3 dB", 'h', 3.0, 1),
                                 ("hysteresis 6 dB", 'h', 6.0, 1),
                                 ("hysteresis 6 dB, 3 reports", 'h', 6.0, 3),
                                 ("hysteresis 12 dB, 3 reports", 'h', 12.0, 3)):
    swaps, drops = walk(rule, hyst, dwell)
    print("  %-38s %18.2f %15.1f%%" % (label, swaps, drops))
print("  one handover is wanted. Without hysteresis the mobile ping-pongs; too much, and it")
print("  clings to a cell until the call drops.")

# 2. The four types of handover, and who is involved in each.
print("\nThe four handover types, and who must take part:")
for name, what, involved in (("intra-cell", "another channel in the same cell (interference)", "BSC"),
                             ("inter-cell, intra-BSC", "another cell of the same BSC", "BSC"),
                             ("inter-BSC, intra-MSC", "a cell of another BSC under the same MSC", "BSC, MSC"),
                             ("inter-MSC", "a cell under another MSC", "BSC, MSC-A, MSC-B")):
    print("  %-22s %-46s %s" % (name, what, involved))
print("  MSC-A anchors the call through an inter-MSC handover: the route from the caller never moves")

# 3. The message sequence of an intra-MSC handover (TS 23.009, figure 7).
print("\nAn external intra-MSC handover, message by message:")
for i, (frm, to, msg) in enumerate([
        ("BSS-A", "MSC-A", "A-HANDOVER-REQUIRED (a list of target cells, in order of preference)"),
        ("MSC-A", "BSS-B", "A-HANDOVER-REQUEST (allocate a channel)"),
        ("BSS-B", "MSC-A", "A-HANDOVER-REQUEST-ACK (here it is)"),
        ("MSC-A", "BSS-A", "A-HANDOVER-COMMAND"),
        ("BSS-A", "MS", "RI-HO-COMMAND (go to that channel)"),
        ("MS", "BSS-B", "RI-HO-ACCESS (access bursts on the new channel)"),
        ("BSS-B", "MSC-A", "A-HANDOVER-DETECT"),
        ("MS", "BSS-B", "RI-HO-COMPLETE"),
        ("BSS-B", "MSC-A", "A-HANDOVER-COMPLETE"),
        ("MSC-A", "BSS-A", "A-CLEAR-COMMAND, then A-CLEAR-COMPLETE")], 1):
    print("  %2d. %-6s -> %-6s %s" % (i, frm, to, msg))
munotes.in711

Handover in GSM

A mobile crossing between two cells 1 km apart at 15 m/s, 200 walks:
  rule                                  handovers per crossing   calls dropped
  no hysteresis, act at once                          10.24             1.5%
  hysteresis 3 dB                                      5.32             0.5%
  hysteresis 6 dB                                      3.00             3.5%
  hysteresis 6 dB, 3 reports                           1.18             5.5%
  hysteresis 12 dB, 3 reports                          0.93            14.5%
  one handover is wanted. Without hysteresis the mobile ping-pongs; too much, and it
  clings to a cell until the call drops.

The four handover types, and who must take part:
  intra-cell             another channel in the same cell (interference) BSC
  inter-cell, intra-BSC  another cell of the same BSC                   BSC
  inter-BSC, intra-MSC   a cell of another BSC under the same MSC       BSC, MSC
  inter-MSC              a cell under another MSC                       BSC, MSC-A, MSC-B
  MSC-A anchors the call through an inter-MSC handover: the route from the caller never moves

An external intra-MSC handover, message by message:
   1. BSS-A  -> MSC-A  A-HANDOVER-REQUIRED (a list of target cells, in order of preference)
   2. MSC-A  -> BSS-B  A-HANDOVER-REQUEST (allocate a channel)
   3. BSS-B  -> MSC-A  A-HANDOVER-REQUEST-ACK (here it is)
   4. MSC-A  -> BSS-A  A-HANDOVER-COMMAND
   5. BSS-A  -> MS     RI-HO-COMMAND (go to that channel)
   6. MS     -> BSS-B  RI-HO-ACCESS (access bursts on the new channel)
   7. BSS-B  -> MSC-A  A-HANDOVER-DETECT
   8. MS     -> BSS-B  RI-HO-COMPLETE
   9. BSS-B  -> MSC-A  A-HANDOVER-COMPLETE
  10. MSC-A  -> BSS-A  A-CLEAR-COMMAND, then A-CLEAR-COMPLETE
munotes.in712

Handover in GSM

Ping-pong. With no hysteresis, one crossing caused 10.24 handovers on average: the mobile swapped cells every time fading moved the balance, which is ten times the signalling and ten chances to fail, for one crossing that needed one handover.

Hysteresis. 3 dB halves it to 5.32; 6 dB gives 3.00; 6 dB with a requirement of three consecutive reports gives 1.18, which is about right: one handover per crossing, with the occasional extra where the shadowing genuinely reversed.

Too much. At 12 dB with three reports, handovers fall to 0.93, but 14.5 per cent of calls drop, against 5.5 per cent at 6 dB: the mobile holds on to a cell that has become too weak. That is the whole design problem in one table: too little hysteresis wastes signalling, too much loses calls, and the operator tunes the margin for each cell.

The types. Intra-cell and inter-cell within a BSC are the BSC's business alone, which is why they are fast and common. An inter-BSC handover adds the MSC; an inter-MSC handover adds a second MSC and a circuit between them, with MSC-A anchoring the call.

The sequence. Ten messages, of which only two cross the air: the command telling the mobile to move, and the access and completion on the new channel. Everything else is the network preparing and cleaning up.

Distinctions

TypeMoves the call toInvolvesTypical cause
Intra-cellAnother channel in the same cellBSCInterference on the current channel
Inter-cell, intra-BSCA cell of the same BSCBSCMovement
Inter-BSC, intra-MSCA cell of another BSCBSC and MSCMovement across a BSC boundary
Inter-MSCA cell of another MSCBSC, MSC-A (anchor), MSC-BMovement across an MSC boundary
Hard handover (GSM)Soft handover (CDMA)
TimingBreak before makeMake before break: both cells serve at once
Possible becauseDifferent frequencies or slotsEvery cell uses the same frequency
InterruptionA brief gapNone
CostA risk of failure at the moment of changeTwo cells' resources during the overlap
munotes.in713

Handover in GSM

Too little hysteresisToo much hysteresis
EffectPing-pong: many needless handoversThe mobile clings to a failing cell
Program10.24 handovers a crossing at 0 dB14.5 per cent of calls dropped at 12 dB
CostsSignalling, and a failure risk each timeDropped calls, and interference to others

What it does not mean

The mobile does not decide. It measures and reports; the BSC decides. That is what mobile-assisted means.

A handover is not always caused by movement. Interference and load cause handovers in a stationary mobile, and an intra-cell handover does not change cell at all.

The anchor MSC is not a detour that can be avoided. Keeping MSC-A in the path is what spares the caller's network any knowledge of the mobile's movement.

Hysteresis is not a delay setting. It is a margin in decibels; the dwell requirement is the delay, and the two are tuned separately.

A dropped call at a boundary is not always a coverage fault. It can be a handover decision that waited too long, which is a parameter problem, not a radio one.

Quick revision

  • Handover: moves an ongoing call to another channel. Causes: movement, quality, load. GSM's is hard (break before make) and mobile-assisted (mobile measures, network decides).
  • Four types: intra-cell (BSC), inter-cell intra-BSC (BSC), inter-BSC intra-MSC (BSC, MSC), inter-MSC (MSC-A anchors, MSC-B serves).
  • Measurements: the mobile uses the idle frame to measure up to six neighbours and reports on the SACCH about twice a second; the BTS measures the uplink and the timing advance; the BSC decides.
  • Decision: hysteresis (a margin in dB) plus a dwell requirement (several consecutive reports). Program: 10.24 handovers a crossing with none; 1.18 at 6 dB and three reports; 14.5 per cent of calls dropped at 12 dB.
  • Sequence (TS 23.009 Fig. 7): HANDOVER-REQUIRED (with a preference-ordered cell list), HANDOVER-REQUEST, ACK, HANDOVER-COMMAND, the mobile's HO-ACCESS, HANDOVER-DETECT, HO-COMPLETE, HANDOVER-COMPLETE, CLEAR-COMMAND, CLEAR-COMPLETE.
  • Inter-MSC: MSC-A "controls the call ... before, during and after a basic or subsequent handover"; later moves are subsequent handovers, back to MSC-A or on to MSC-B'.

Test yourself

1. Why is handover needed, and what are its causes? Because a call in progress must survive changes in the radio link. The usual cause is movement out of the serving cell, but a handover may also be triggered by falling quality without movement, for example interference or a fade, in which case even a change of channel within the same cell may help, and by load, when a busy cell hands a mobile to a neighbour that can serve it.

2. Describe the four types of handover in GSM and say who takes part. An intra-cell handover moves the call to another channel of the same cell and is handled by the BSC alone, usually because of interference. An inter-cell handover within one BSC also involves only that BSC. An inter-BSC handover within one MSC involves both BSCs and the MSC, which switches the connection. An inter-MSC handover involves the original MSC, which remains the anchor and keeps control of the call, a second MSC which takes over the radio side, and a circuit between them.

munotes.in714

Handover in GSM

3. How does GSM collect the information a handover decision needs? The handover is mobile-assisted. Each mobile uses the idle frame of its 26-multiframe to retune to neighbouring broadcast carriers, measuring the received level of up to six neighbours and identifying them by their base station identity codes, and it reports these together with the level and quality of its serving cell on the slow associated control channel, about twice a second. The BTS measures the uplink level, quality and timing advance. All these reports go to the BSC, which takes the decision.

4. What are hysteresis and the dwell requirement, and why are both needed? Hysteresis is a margin: a neighbouring cell must be better than the serving cell by several decibels before a handover is considered, so that the reverse handover would require the same advantage in the other direction. The dwell requirement is that the condition must hold for several consecutive measurement reports. Without them, fading makes the comparison flicker and the mobile ping-pongs between cells: in the chapter's simulation, a single crossing produced 10.24 handovers with no hysteresis and 1.18 with 6 dB and three reports. Too much of either is also harmful: at 12 dB with three reports, 14.5 per cent of calls dropped because the mobile clung to a failing cell.

5. Give the message sequence of an intra-MSC handover. BSS-A, which is serving the mobile, decides and sends A-HANDOVER-REQUIRED to MSC-A with a list of candidate cells in order of preference. MSC-A sends A-HANDOVER-REQUEST to BSS-B, which allocates a channel and replies with A-HANDOVER-REQUEST-ACK. MSC-A sends A-HANDOVER-COMMAND to BSS-A, which passes the handover command to the mobile over the air. The mobile accesses the new channel, BSS-B reports A-HANDOVER-DETECT, the mobile completes, and BSS-B sends A-HANDOVER-COMPLETE. MSC-A then clears the old resources with A-CLEAR-COMMAND and receives A-CLEAR-COMPLETE.

6. What is an anchor MSC, and what is a subsequent handover? In an inter-MSC handover the original MSC, MSC-A, keeps control of the call and of mobility management before, during and after the handover, and the connection from the caller continues to reach MSC-A, which extends it over a circuit to the new MSC, MSC-B. MSC-A is therefore the anchor. If the mobile moves again, the result is a subsequent handover: either back to MSC-A, in which case the extra leg is released, or onward to a third MSC, in which case MSC-A establishes a new leg and releases the old one. The anchor does not change during the call, which spares the caller's network any involvement at the cost of a longer path.

Contents This chapter on its own page

munotes.in715

Chapter Ninety-Four

GSM Security

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: GSM: Security"

In one line

GSM protects the air interface and nothing else: the SIM and the authentication centre share a secret key that never moves, a challenge and response proves the subscriber to the network, the same challenge derives a ciphering key for the radio link, and a temporary identity hides who is talking, but the network never proves itself to the subscriber, the ciphering stops at the base station, nothing protects the integrity of a message, and the original ciphers have been broken.

In the wording a student can write in an examination: GSM's security rests on a secret key Ki, 128 bits, held only in the SIM and in the authentication centre (AuC). Three algorithms use it:

  • A3, authentication: "the purpose of Algorithm A3 is to allow authentication of a mobile subscriber's identity." To do that, "Algorithm A3 must compute an expected response SRES from a random challenge RAND sent by the network. For this computation, Algorithm A3 makes use of the secret authentication key Ki."
  • A8, key generation: "Algorithm A8 must compute the ciphering key Kc from the random challenge RAND sent during the authentication procedure, using the authentication key Ki." On the mobile side, "Algorithm A8 is contained in the SIM"; on the network side it "is co-located with Algorithm A3".
  • A5, encryption: "Algorithm A5 realizes the protection of both user data and signalling information elements at the physical layer on the dedicated channels (TCH or DCCH)", and "Algorithm A5 is implemented into both the MS and the BSS".

The authentication procedure: the AuC computes triplets (RAND, SRES, Kc) from Ki and sends them to the VLR; the VLR sends RAND to the mobile; the SIM computes SRES with A3 and Kc with A8 and returns SRES; the VLR compares. Ki never leaves the SIM or the AuC, and Kc is never transmitted. Ciphering with A5 then uses Kc and the frame number, so the keystream differs in every frame. Anonymity comes from the TMSI, a temporary identity allocated by the VLR in place of the IMSI.

Its weaknesses are: the network is not authenticated to the mobile, so a false base station can impersonate it and can order A5/0, no ciphering; encryption covers only MS to BTS, so the Abis link and the core network carry the call in clear unless separately protected; there is no integrity protection of signalling; and A5/1 and A5/2 have been broken by published cryptanalysis, with A5/3 and A5/4 added later. UMTS answers the first three with mutual authentication and integrity protection ([UMTS and IMT-2000]).

What GSM set out to protect

Analogue systems could be listened to with a scanner and cloned by copying an identity off the air. GSM's designers therefore set three goals, and it is worth noticing which three:

munotes.in716

GSM Security

  1. Authenticate the subscriber, so that calls are billed to whoever really made them and a cloned identity does not work.
  2. Encrypt the radio path, so that casual interception fails.
  3. Hide the subscriber's identity on the air, so that a listener cannot track a person.

The goals are all about the operator's interests on the radio link: fraud, casual eavesdropping, and traffic analysis. Protecting the subscriber from the network, or the call once it is inside the network, was not among them, and that omission explains every weakness below.

The triplet, and why it exists

The clever part of the design is that a visited network authenticates a subscriber it knows nothing about, without ever learning the subscriber's secret.

The AuC, sitting with the HLR, holds Ki. On request it computes a batch of triplets, each one a random challenge, the response that the SIM will give, and the ciphering key that both will derive, and sends them to the VLR. The VLR can then authenticate the subscriber repeatedly, and start ciphering, without another word to the home network, and without ever holding Ki.

The exchange on the air is two messages: RAND down, SRES up. Both travel in clear, which is safe because knowing RAND and SRES does not give Ki, and because the next challenge is different.

Ciphering

Once authenticated, the network orders ciphering, naming which A5 variant to use, and both ends load Kc. A5 is a stream cipher: it produces a keystream from Kc and the frame number, and the keystream is added to the bits. Using the frame number means the keystream changes every 4.615 ms and does not repeat until the hyperframe does, about three and a half hours later ([GSM Logical Channels and the Frame Hierarchy] computed it, and this is the reason it is so long).

Ciphering runs between the MS and the BTS, as the standard says. Beyond the BTS, the call travels over Abis and into the core in clear, unless the operator has separately protected those links.

Anonymity

The IMSI is the permanent identity, and sending it on the air would let anyone track a subscriber. GSM therefore allocates a TMSI, sends it ciphered, and changes it, as [Localization and Calling in GSM] described. It is a real improvement, and it is not a guarantee: a network can always demand the IMSI, and a mobile that has just entered a network with no usable TMSI must send it.

munotes.in717

GSM Security

Where the design falls short

One-way authentication. The subscriber proves itself to the network; the network proves nothing. A false base station can therefore broadcast a stronger signal claiming to be the operator, attract nearby mobiles, decline to authenticate them (nothing requires it to), order A5/0, no ciphering, and relay the calls onward while listening. The program lists the steps. This is the weakness UMTS was designed to close, by having the network prove knowledge of the key too.

Encryption stops at the BTS. Everything in the core is in clear to anyone with access to it, which includes microwave links between a BTS and its BSC.

No integrity protection. A GSM message is encrypted but not authenticated, so an attacker who can modify bits changes the message, and the receiver has no way to notice. Integrity protection of signalling is again a UMTS addition.

Weak algorithms. A5/1 was deliberately weakened for export, A5/2 more so, and both have been broken by published cryptanalysis; A5/0 is no encryption at all. A5/3 (based on KASUMI) and later A5/4 were added, and the standard's clauses on negotiation exist because a network and a mobile must agree which variant to use, which is itself an attack surface: an attacker who can force the choice downward gets a weaker cipher.

And one more, worth knowing. Some networks reduced the effective length of Kc, fixing part of it to zero, so that the 64-bit key carried fewer than 64 bits of entropy.

GSM security, computed

The program computes triplets with stand-in algorithms; tabulates what crosses the air and what never does; lists the properties the design provides and those it does not; explains the frame number's role in the keystream; and walks a false base station attack.

# GSM security: the triplet, where each secret lives, and what the design does
# and does not protect against.
import hashlib
import random

# 1. The triplet, with stand-in algorithms. The real A3 and A8 are operator
#    secrets; here SHA-256 stands in for both, which is enough to show the flow.
def a3a8(ki, rand):
    h = hashlib.sha256(ki + rand).digest()
    return h[:4], h[4:12]                      # SRES (32 bits), Kc (64 bits)

ki = bytes(range(16))                          # the SIM's secret key, 128 bits
print("Authentication, as the AuC and the SIM each compute it:")
rnd = random.Random(94)
for i in range(3):
    rand = bytes(rnd.randrange(256) for _ in range(16))
    sres, kc = a3a8(ki, rand)
    print("  RAND %s -> SRES %s, Kc %s" % (rand[:6].hex(), sres.hex(), kc.hex()))
print("  the AuC computes triplets in advance and sends them to the VLR; the SIM computes")
print("  the same values when challenged, and Ki never leaves either place.")

# 2. What is sent and what is not.
print("\nWhat crosses the air, and what does not:")
for item, crosses in (("RAND, the challenge", "yes, in clear"),
                      ("SRES, the response", "yes, in clear"),
                      ("Ki, the subscriber key", "never"),
                      ("Kc, the ciphering key", "never: both ends derive it"),
                      ("IMSI", "only when no TMSI can be used"),
                      ("TMSI", "yes, but it is temporary and local")):
    print("  %-26s %s" % (item, crosses))

# 3. Where GSM's security stops. Each row is a property and whether GSM has it.
print("\nWhat the design provides, and what it does not:")
for prop, verdict, why in (
        ("the network authenticates the subscriber", "yes", "A3 with a challenge and response"),
        ("the subscriber authenticates the network", "NO", "nothing stops a false base station"),
        ("confidentiality on the air", "yes", "A5 between the MS and the BTS"),
        ("confidentiality beyond the BTS", "NO", "Abis and the core are in clear unless separately protected"),
        ("subscriber anonymity", "partly", "the TMSI hides the IMSI, but the IMSI can be demanded"),
        ("integrity of signalling", "NO", "no message authentication: messages can be altered"),
        ("strong ciphers", "NO", "A5/1 and A5/2 are broken; A5/3 and A5/4 came later")):
    print("  %-42s %-7s %s" % (prop, verdict, why))

# 4. Why a 64-bit key and a 22-bit frame number matter: the keystream must not
#    repeat, and the hyperframe is what guarantees it.
FRAME_MS = 4.615384615
print("\nThe keystream and the frame number:")
print("  A5 takes Kc and the frame number, so the keystream repeats when the frame number does")
print("  the hyperframe is 2,715,648 frames, %.2f hours: longer than any call" % (2715648 * FRAME_MS / 3600000))
print("  Kc is 64 bits, but some networks fixed 10 of them to zero, leaving %d bits of real key" % 54)

# 5. A false base station, in one exchange.
print("\nWhy a false base station works against GSM:")
for step in ("the attacker broadcasts a stronger BCCH claiming to be the network",
             "the mobile camps on it and offers its identity; the attacker may demand the IMSI",
             "the attacker never asks for authentication, because nothing requires it to",
             "the attacker orders A5/0, that is no ciphering, and the mobile complies",
             "the attacker relays the call onward, listening to everything"):
    print("  - %s" % step)
print("  the fix is mutual authentication, which UMTS introduced.")
munotes.in718

GSM Security

Authentication, as the AuC and the SIM each compute it:
  RAND 5d3f8f9adc08 -> SRES 3d552e3b, Kc 6ae886362dafe511
  RAND e8b36d17d2ee -> SRES 4721b982, Kc 02454ad4311f3e2b
  RAND b214098d7314 -> SRES f434170e, Kc 121b14a7f9f20069
  the AuC computes triplets in advance and sends them to the VLR; the SIM computes
  the same values when challenged, and Ki never leaves either place.

What crosses the air, and what does not:
  RAND, the challenge        yes, in clear
  SRES, the response         yes, in clear
  Ki, the subscriber key     never
  Kc, the ciphering key      never: both ends derive it
  IMSI                       only when no TMSI can be used
  TMSI                       yes, but it is temporary and local

What the design provides, and what it does not:
  the network authenticates the subscriber   yes     A3 with a challenge and response
  the subscriber authenticates the network   NO      nothing stops a false base station
  confidentiality on the air                 yes     A5 between the MS and the BTS
  confidentiality beyond the BTS             NO      Abis and the core are in clear unless separately protected
  subscriber anonymity                       partly  the TMSI hides the IMSI, but the IMSI can be demanded
  integrity of signalling                    NO      no message authentication: messages can be altered
  strong ciphers                             NO      A5/1 and A5/2 are broken; A5/3 and A5/4 came later

The keystream and the frame number:
  A5 takes Kc and the frame number, so the keystream repeats when the frame number does
  the hyperframe is 2,715,648 frames, 3.48 hours: longer than any call
  Kc is 64 bits, but some networks fixed 10 of them to zero, leaving 54 bits of real key

Why a false base station works against GSM:
  - the attacker broadcasts a stronger BCCH claiming to be the network
  - the mobile camps on it and offers its identity; the attacker may demand the IMSI
  - the attacker never asks for authentication, because nothing requires it to
  - the attacker orders A5/0, that is no ciphering, and the mobile complies
  - the attacker relays the call onward, listening to everything
  the fix is mutual authentication, which UMTS introduced.
munotes.in719

GSM Security

The triplet. Three challenges give three different responses and three different ciphering keys, all computed from the same Ki. Two things follow: the VLR can authenticate as often as it likes from a batch, and an attacker who records one exchange learns nothing usable, because the next RAND will be different.

What crosses. RAND and SRES cross in clear; Ki never crosses at all, and Kc never crosses either, since both ends compute it. That is the strength of the design, and it is genuinely strong: cloning a SIM by listening to the air does not work.

What is missing. The table is the examination answer: subscriber authenticated yes, network authenticated no, air encrypted yes, core encrypted no, anonymity partly, integrity no, strong ciphers no in the original algorithms. Four of seven are absent, and every one of the four is a consequence of the same decision, that security was designed to protect the operator from the subscriber and from the casual listener, not the subscriber from anyone.

The false base station. Five steps, none of which requires breaking any cipher. That is the difference between a cryptographic weakness and a protocol weakness: the ciphers can be replaced, but a protocol that never authenticates the network cannot be patched from the handset.

munotes.in720

GSM Security

Distinctions

A3A8A5
PurposeAuthenticate the subscriberDerive the ciphering keyEncrypt on the dedicated channels
InputsKi, RANDKi, RANDKc, frame number
OutputSRESKcKeystream
WhereSIM and AuCSIM and AuC, co-located with A3MS and BSS
Chosen byThe operator (secret)The operator (secret)The standard (A5/0 to A5/4), negotiated
Held by the SIMHeld by the AuCSent over the air
KiYesYesNever
RANDReceivedGeneratedYes, in clear
SRESComputedPrecomputedYes, in clear
KcComputedPrecomputedNever
IMSIYesYes (with the HLR)Only when no TMSI can be used
PropertyGSMUMTS
Subscriber authenticatedYesYes
Network authenticatedNoYes (mutual)
Integrity of signallingNoYes
Encryption reachesThe BTSThe radio network controller
CipherA5/1, A5/2 broken; A5/3 laterKASUMI-based from the start

What it does not mean

Encrypted does not mean end to end. The call is in clear from the BTS onward.

Authentication is not mutual. The network never proves who it is, which is what a false base station exploits.

A TMSI is not anonymity. It hides the IMSI from a casual listener; the network can always ask for the IMSI.

Breaking A5 is not the only attack. The protocol weaknesses need no cryptanalysis at all.

A secret algorithm is not a strong one. A3 and A8 are operator choices, and the early common example, COMP128, was itself broken.

Quick revision

  • Ki: 128-bit key, only in the SIM and the AuC.
  • A3: computes SRES from RAND and Ki, to "allow authentication of a mobile subscriber's identity". A8: computes Kc from RAND and Ki, in the SIM, co-located with A3 in the network. A5: protects "both user data and signalling information elements at the physical layer on the dedicated channels (TCH or DCCH)", in the MS and the BSS, keyed by Kc and the frame number.
  • Triplet (RAND, SRES, Kc): computed by the AuC, used by the VLR, so the visited network never sees Ki. RAND and SRES cross in clear; Ki and Kc never cross.
  • Anonymity: the TMSI, temporary and local.
  • Weaknesses: no network authentication (false base station, can force A5/0), encryption only to the BTS, no integrity protection, A5/1 and A5/2 broken (A5/3, A5/4 later), and Kc sometimes effectively shortened.
  • UMTS answers with mutual authentication and integrity protection.

Test yourself

1. Describe GSM's authentication procedure. The authentication centre, which shares the subscriber key Ki with the SIM, computes triplets consisting of a random challenge RAND, the expected response SRES that the SIM will produce from RAND and Ki using algorithm A3, and the ciphering key Kc that both will derive from RAND and Ki using algorithm A8. These triplets are sent to the visitor location register. To authenticate, the network sends RAND to the mobile; the SIM computes SRES and returns it; the network compares it with the value in the triplet. The key Ki never leaves the SIM or the authentication centre, and Kc is never transmitted.

munotes.in721

GSM Security

2. What are A3, A5 and A8, and where does each run? A3 authenticates the subscriber's identity by computing an expected response SRES from a random challenge RAND using the secret key Ki; it runs in the SIM and in the authentication centre. A8 computes the ciphering key Kc from RAND and Ki; it is contained in the SIM and, on the network side, is co-located with A3. A5 protects both user data and signalling at the physical layer on the dedicated channels; it is implemented in the mobile station and in the base station subsystem, and is keyed by Kc together with the frame number.

3. Why does the authentication centre send triplets rather than the key? So that a visited network can authenticate a subscriber and encrypt the radio link without ever learning the subscriber's secret key. The triplets contain only a challenge, the expected answer and the derived session key, so a compromised or untrusted visited network cannot impersonate the subscriber in the future or derive keys for other challenges; and because triplets are sent in batches, the visited network can authenticate repeatedly without further traffic to the home network.

4. How does a false base station attack GSM, and why does it work? The attacker transmits a stronger broadcast channel claiming to be the operator's network, so nearby mobiles camp on it. Because GSM authenticates only the subscriber to the network, and never the network to the subscriber, the false base station is not obliged to authenticate the mobile at all, and it can order the null cipher A5/0 or a weak variant, which the mobile accepts. It can then relay the traffic to the real network while reading it. No cipher has to be broken; the weakness is in the protocol, and the answer is mutual authentication, which UMTS introduced.

5. What does GSM's encryption not cover? It covers only the radio path between the mobile station and the base transceiver station. The Abis link from the BTS to the BSC, the A interface, and everything in the core network carry the call unencrypted unless the operator protects them separately, so anyone with access to those links, including microwave hops between sites, can listen. GSM also provides no integrity protection at all, so a message can be altered without detection.

munotes.in722

GSM Security

6. Summarise GSM's security weaknesses and how UMTS answers them. GSM authenticates the subscriber but not the network, so false base stations work; it encrypts only to the base transceiver station; it provides no integrity protection of signalling; and its original ciphers, A5/1 and A5/2, have been broken, with A5/0 providing none at all. UMTS introduces mutual authentication, so the network must prove knowledge of the key; it adds integrity protection of signalling messages; it extends the encryption to the radio network controller; and it uses a published, stronger algorithm from the start.

Contents This chapter on its own page

munotes.in723

Chapter Ninety-Five

New Data Services: HSCSD, GPRS and EDGE

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: GSM: New data services"

In one line

GSM's 9.6 kbit/s came from giving a data call one speech slot, so the three answers are to take more slots (HSCSD), to stop reserving them at all and send packets instead (GPRS), and to put more bits in each symbol by changing the modulation (EDGE), of which only the second changes the architecture and only the third changes the radio.

In the wording a student can write in an examination: HSCSD (high speed circuit switched data) bundles several time slots for one circuit-switched call, giving up to about 57.6 kbit/s with four slots at 14.4 kbit/s each. It is simple, needs mostly software in the network, and keeps circuit switching's weakness: the slots are reserved for the whole call whether data flows or not.

GPRS (general packet radio service) adds packet switching. Slots are allocated per block, so many users share them and a user pays for volume rather than time; it adds two nodes to the core, the SGSN (serving GPRS support node), which serves the mobile with routing, mobility and ciphering, and the GGSN (gateway GPRS support node), the gateway to external packet networks; a packet control unit in the BSS; a packet subscription in the HLR; and the concepts of attach and the PDP context, the packet equivalent of a call. It defines four coding schemes, CS-1 to CS-4, at 9.05, 13.4, 15.6 and 21.4 kbit/s a slot, chosen according to the channel's quality, so that up to eight slots give a theoretical maximum of about 171 kbit/s.

EDGE (enhanced data rates for GSM evolution) changes the modulation from GMSK to 8PSK, three bits per symbol instead of one, in the same 200 kHz carrier and the same slots, and defines nine modulation and coding schemes, MCS-1 to MCS-9, with the highest giving about 59.2 kbit/s a slot. With GPRS it is EGPRS; with HSCSD, ECSD. Its higher schemes need a better signal, so the rate falls with distance from the base station.

HSCSD: more slots, same idea

The cheapest answer to a rate of 9.6 kbit/s being too slow is to give the user more than one slot. HSCSD does exactly that: a mobile is assigned several slots of the same carrier, and the network bundles them into one circuit.

What it gains is proportionality: four slots give four times the rate. What it does not change is the reservation. A circuit is held for the whole call, so a user reading a page occupies four slots while sending nothing, and the operator has to charge by time because that is what is being consumed. The program measures how bad that is for a browsing session.

munotes.in724

New Data Services: HSCSD, GPRS and EDGE

Its virtue is that it needed almost no new equipment: the same slots, the same channels, mostly software.

GPRS: packet switching added to GSM

GPRS is the real change, and it is a change to the core network as much as to the radio.

On the radio side, slots are no longer reserved. A packet control unit in the base station subsystem allocates radio blocks to whichever mobile has data, so several mobiles share the same slots and one mobile can use several slots. Uplink and downlink are allocated separately, which is why a mobile's capability is quoted as a class, so many slots down and so many up.

In the core, two nodes appear:

  • The SGSN serves the mobile: it keeps track of where the mobile is, authenticates and ciphers, routes packets to and from it, and counts the traffic for billing. It is to packets what the MSC is to calls.
  • The GGSN is the gateway to the outside: it holds the connection to the external packet network, gives the mobile an address there, and tunnels packets to the SGSN currently serving it.

In the subscription, the mobile attaches to the GPRS service, which is the packet equivalent of IMSI attach, and then activates a PDP context, which sets up the path to the GGSN and gives the mobile an address. A mobile can be attached, and reachable, while sending nothing, which is what "always on" means and what circuit switching could not offer.

The coding schemes are where the radio's quality enters. CS-1 protects heavily and carries little; CS-4 has no error protection at all, "1" code rate in the standard's table, and needs a clean channel. The network chooses per block. The standard notes that "CS-1 is the same coding scheme as specified for SACCH in 3GPP TS 45.003", which is a nice economy: the most protected packet scheme is the one GSM already used for its control channel.

EDGE: more bits per symbol

EDGE leaves the frame, the slots and the 200 kHz carrier exactly as they are, and changes what a symbol means. GSM's GMSK carries one bit per symbol ([Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM]); EDGE's 8PSK carries three. At the same symbol rate of 270.833 ksymbol/s, that is three times the raw rate.

The cost is robustness: eight points on a circle are much closer together than two, so the highest schemes need a strong, clean signal. EDGE therefore defines nine schemes, MCS-1 to MCS-9, the lower ones using GMSK and the higher ones 8PSK with progressively less coding, and switches between them as the channel changes. A mobile near the base station gets MCS-9; one at the edge falls back to a GMSK scheme and is no worse off than with GPRS.

munotes.in725

New Data Services: HSCSD, GPRS and EDGE

EDGE also improved the retransmission scheme (incremental redundancy, where a failed block is retransmitted with different coding and the two attempts are combined), which is why its throughput at a given quality beats what the raw rates suggest.

The new services, computed

The program tabulates the GPRS coding schemes and what one to eight slots give; computes HSCSD's bundled rates; models an hour of browsing and asks how much of a reserved circuit would be wasted; compares GMSK and 8PSK at eight slots; and lists what GPRS added to the core.

# From one slot of GSM data to HSCSD, GPRS and EDGE: what each step buys.
# The GPRS coding schemes are TS 43.064's own table 1.
CS = [("CS-1", 0.5, 181, 9.05, 8.0), ("CS-2", 2 / 3, 268, 13.4, 12.0),
      ("CS-3", 0.75, 312, 15.6, 14.4), ("CS-4", 1.0, 428, 21.4, 20.0)]

print("GPRS coding schemes (TS 43.064, Table 1), and what n slots give:")
print("  scheme  code rate  bits a block  rate a slot   2 slots   4 slots   8 slots")
for name, rate, bits, kbps, user in CS:
    print("  %-7s %9.2f %13d %8.2f kb/s %9.1f %9.1f %9.1f"
          % (name, rate, bits, kbps, 2 * kbps, 4 * kbps, 8 * kbps))
print("  CS-1 has the most protection and the least rate; CS-4 has no coding at all and needs a")
print("  good channel. The scheme is chosen per block from the quality the network sees.")

# 2. HSCSD: bundling circuit slots. The rate is a multiple of one slot's.
print("\nHSCSD: bundling circuit-switched slots (14.4 kb/s a slot):")
for slots in (1, 2, 3, 4):
    print("  %d slot(s): %5.1f kb/s downlink; a slot is held for the whole call whether or not it is used"
          % (slots, 14.4 * slots))

# 3. What packet switching changes. A user browsing: bursts of data with long
#    gaps. Compare a circuit held for the session with packets sharing slots.
print("\nOne hour of browsing: 40 pages, each 100 kB, read for 60 s, over 4 slots at 57.6 kb/s:")
pages, size_kb, read_s, rate_kbps = 40, 100, 60, 57.6
transfer = size_kb * 8 / rate_kbps
session = pages * (transfer + read_s)
print("  transferring one page takes %.1f s; a session lasts %.0f minutes" % (transfer, session / 60))
print("  circuit switched: the 4 slots are held for %.0f minutes, and used for %.1f minutes, %.0f%%"
      % (session / 60, pages * transfer / 60, 100 * pages * transfer / session))
print("  packet switched: the slots are used only while data flows, so the same 4 slots could")
print("  serve about %d such users at once before they queue" % int(session / (pages * transfer)))

# 4. EDGE: the same 200 kHz and the same slots, but 8PSK instead of GMSK.
print("\nEDGE: 3 bits a symbol instead of 1, in the same carrier:")
for name, bits_per_symbol, note in (("GMSK (GPRS)", 1, "CS-1 to CS-4, up to 21.4 kb/s a slot"),
                                    ("8PSK (EDGE)", 3, "MCS-1 to MCS-9, up to about 59.2 kb/s a slot")):
    print("  %-12s %d bit(s) a symbol: %s" % (name, bits_per_symbol, note))
print("  8 slots of the fastest scheme: GPRS %.1f kb/s, EDGE about %.0f kb/s"
      % (8 * 21.4, 8 * 59.2))
print("  EDGE needs a better signal for its highest schemes, so the rate falls with distance.")

# 5. The architecture GPRS added.
print("\nWhat GPRS added to the GSM core:")
for entity, what in (("SGSN", "serves the mobile: routing, mobility, ciphering, billing for packets"),
                     ("GGSN", "the gateway to external packet networks; holds the mobile's IP address"),
                     ("GPRS register in the HLR", "the packet subscription, beside the circuit one"),
                     ("PCU in the BSS", "shares radio blocks between packet users")):
    print("  %-26s %s" % (entity, what))
print("  a mobile 'attaches' and activates a 'PDP context', which is the packet equivalent of a call.")
munotes.in726

New Data Services: HSCSD, GPRS and EDGE

GPRS coding schemes (TS 43.064, Table 1), and what n slots give:
  scheme  code rate  bits a block  rate a slot   2 slots   4 slots   8 slots
  CS-1         0.50           181     9.05 kb/s      18.1      36.2      72.4
  CS-2         0.67           268    13.40 kb/s      26.8      53.6     107.2
  CS-3         0.75           312    15.60 kb/s      31.2      62.4     124.8
  CS-4         1.00           428    21.40 kb/s      42.8      85.6     171.2
  CS-1 has the most protection and the least rate; CS-4 has no coding at all and needs a
  good channel. The scheme is chosen per block from the quality the network sees.

HSCSD: bundling circuit-switched slots (14.4 kb/s a slot):
  1 slot(s):  14.4 kb/s downlink; a slot is held for the whole call whether or not it is used
  2 slot(s):  28.8 kb/s downlink; a slot is held for the whole call whether or not it is used
  3 slot(s):  43.2 kb/s downlink; a slot is held for the whole call whether or not it is used
  4 slot(s):  57.6 kb/s downlink; a slot is held for the whole call whether or not it is used

One hour of browsing: 40 pages, each 100 kB, read for 60 s, over 4 slots at 57.6 kb/s:
  transferring one page takes 13.9 s; a session lasts 49 minutes
  circuit switched: the 4 slots are held for 49 minutes, and used for 9.3 minutes, 19%
  packet switched: the slots are used only while data flows, so the same 4 slots could
  serve about 5 such users at once before they queue

EDGE: 3 bits a symbol instead of 1, in the same carrier:
  GMSK (GPRS)  1 bit(s) a symbol: CS-1 to CS-4, up to 21.4 kb/s a slot
  8PSK (EDGE)  3 bit(s) a symbol: MCS-1 to MCS-9, up to about 59.2 kb/s a slot
  8 slots of the fastest scheme: GPRS 171.2 kb/s, EDGE about 474 kb/s
  EDGE needs a better signal for its highest schemes, so the rate falls with distance.

What GPRS added to the GSM core:
  SGSN                       serves the mobile: routing, mobility, ciphering, billing for packets
  GGSN                       the gateway to external packet networks; holds the mobile's IP address
  GPRS register in the HLR   the packet subscription, beside the circuit one
  PCU in the BSS             shares radio blocks between packet users
  a mobile 'attaches' and activates a 'PDP context', which is the packet equivalent of a call.
munotes.in727

New Data Services: HSCSD, GPRS and EDGE

The coding schemes. CS-1 carries 9.05 kbit/s a slot with a half-rate code; CS-4 carries 21.4 with none. Eight slots of CS-4 give 171.2 kbit/s, the number GPRS was advertised with, and which no user ever saw, because eight slots are rarely free and CS-4 needs a channel that a moving mobile rarely has.

HSCSD. Four slots give 57.6 kbit/s, which is genuinely four times one slot; the arithmetic is the least interesting thing about it.

What a circuit wastes. Forty pages of 100 kB, each read for a minute, over four slots at 57.6 kbit/s: each page takes 13.9 seconds to transfer, the session lasts 49 minutes, and the slots are actually used for 9.3 minutes, 19 per cent of the time. A circuit-switched session reserves the other 81 per cent for nothing. Packet switching gives those slots to other users, so the same four slots serve about five such browsers at once. That single figure is why GPRS existed and why data tariffs changed from per minute to per megabyte.

EDGE. Eight slots of GPRS's best scheme give 171.2 kbit/s; eight of EDGE's give about 474 kbit/s, nearly three times, in the same spectrum, from the same sites, with new software and a new modulator. It was the cheapest capacity any operator ever bought.

What GPRS added. SGSN, GGSN, a packet subscription in the HLR, a packet control unit in the BSS, and the attach and PDP context procedures. The MSC and the circuit core are untouched, which is why GPRS could be deployed on a live network.

Distinctions

HSCSDGPRSEDGE
ChangesHow many slots a call getsThe switching: packets, not circuitsThe modulation: 8PSK, not GMSK
New core nodesNoneSGSN, GGSNNone
RadioUnchangedShared blocks, PCU, CS-1 to CS-4MCS-1 to MCS-9, link adaptation
Rate14.4 kbit/s a slot, up to 57.69.05 to 21.4 a slot, up to 171.2Up to about 59.2 a slot, about 474
Charged byTimeVolumeVolume
Idle userHolds its slotsHolds nothingHolds nothing
munotes.in728

New Data Services: HSCSD, GPRS and EDGE

Circuit switchedPacket switched
ResourceReserved for the sessionAllocated per block
SetupOnce, then guaranteedPer context, then contention
SuitsSpeech, steady streamsBursty traffic: web, email, messaging
Program19 per cent of the reserved slots usedThe same slots serve about 5 users
SGSNGGSN
Compare withThe MSC and VLRThe GMSC
JobServing the mobile: mobility, ciphering, routing, countingThe gateway: an address in the outside network, tunnels to the SGSN
KnowsWhere the mobile isWhich SGSN is serving it

What it does not mean

HSCSD is not a small GPRS. It is still circuit switched: the slots are held whether used or not.

GPRS's 171 kbit/s is not a user rate. It assumes eight free slots and CS-4's unprotected coding at once.

EDGE is not a new network. It is a modulation and a set of coding schemes on the same carriers, frames and slots.

A PDP context is not a call. It is a path to the packet network that can exist while nothing is being sent.

Packet switching does not remove contention. It replaces reservation with sharing, so a busy cell delays everyone a little instead of refusing a few.

Quick revision

  • HSCSD: bundles circuit slots, 14.4 kbit/s each, up to 57.6; no new nodes; slots reserved for the whole call; billed by time.
  • GPRS: packet switched; SGSN (serving: mobility, ciphering, routing, billing) and GGSN (gateway: address and tunnel); PCU in the BSS; attach and PDP context; coding schemes CS-1 9.05, CS-2 13.4, CS-3 15.6, CS-4 21.4 kbit/s a slot, so 171.2 with eight; billed by volume; CS-1 is the SACCH's coding scheme.
  • EDGE: 8PSK, three bits a symbol, same carrier and slots; MCS-1 to MCS-9, up to about 59.2 kbit/s a slot, about 474 with eight; link adaptation and incremental redundancy; with GPRS it is EGPRS.
  • Program: a browsing session uses 19 per cent of a reserved four-slot circuit, so packets let the same slots serve about 5 users.

Test yourself

1. What is HSCSD and what are its limitations? High speed circuit switched data gives one data call several time slots of the same carrier, bundling them into one circuit, so four slots at 14.4 kbit/s give about 57.6 kbit/s. It needed little new equipment, mostly software, because it changed nothing about switching or the radio. Its limitation is that it remains circuit switched: the slots are reserved for the whole call whether data is flowing or not, so a user reading a web page ties up capacity for nothing, and the operator must charge by time.

munotes.in729

New Data Services: HSCSD, GPRS and EDGE

2. What does GPRS add to a GSM network? Packet switching. On the radio side a packet control unit in the base station subsystem allocates radio blocks per block rather than reserving slots, so several mobiles share slots and one mobile may use several. In the core it adds the serving GPRS support node, which handles mobility, authentication, ciphering, routing and accounting for the mobile, and the gateway GPRS support node, which connects to external packet networks, gives the mobile an address and tunnels its packets to the serving node. In the subscription it adds GPRS attach and the PDP context, so a mobile can be reachable while sending nothing, and it defines four coding schemes for the radio blocks.

3. Name GPRS's coding schemes and their rates, and explain the choice between them. CS-1 at 9.05 kbit/s per slot uses a half-rate convolutional code, the same coding as the SACCH; CS-2 at 13.4 and CS-3 at 15.6 use progressively weaker coding; CS-4 at 21.4 has a code rate of 1, that is no forward error correction at all. The network chooses per block according to the quality it sees: a mobile with a strong, clean signal is given a fast, lightly protected scheme, while one at the cell edge is given CS-1, whose protection is what lets its blocks arrive at all.

4. How does EDGE increase the data rate? It changes the modulation from GMSK, which carries one bit per symbol, to 8PSK, which carries three, while keeping the same 200 kHz carrier, the same symbol rate of about 270.833 ksymbol/s, the same frame and the same slots. It defines nine modulation and coding schemes, the lower ones still GMSK and the higher ones 8PSK with progressively less coding, and adapts between them as the channel changes; it also uses incremental redundancy so that a failed block's retransmission is combined with the first attempt. The highest scheme gives about 59.2 kbit/s per slot, so eight slots give about 474 kbit/s against GPRS's 171.

5. Why is packet switching so much better than circuit switching for web browsing? Because browsing is bursty: a page is fetched in seconds and then read for a minute. In the chapter's model of forty 100 kB pages read for a minute each over four slots at 57.6 kbit/s, the transfers take 9.3 minutes out of a 49-minute session, so a reserved circuit would be idle 81 per cent of the time. Packet switching allocates the slots only while data is flowing, so the same four slots can serve about five such users at once, and the user can be charged for the volume transferred rather than for the time connected.

munotes.in730

New Data Services: HSCSD, GPRS and EDGE

6. Compare HSCSD, GPRS and EDGE as answers to the same problem. All three raise the data rate of a GSM network. HSCSD gives one call more slots and changes nothing else, so it is easy but wasteful and still billed by time. GPRS changes the switching, adding the SGSN and GGSN and sharing radio blocks between users, which suits bursty traffic, allows always-on connectivity and volume billing, and reaches up to 171.2 kbit/s in theory. EDGE changes only the modulation, from GMSK to 8PSK, tripling the bits per symbol in the same spectrum and reaching about 474 kbit/s over eight slots, at the cost of needing a better signal for its highest schemes.

Contents This chapter on its own page

munotes.in731

Chapter Ninety-Six

DECT: System Architecture and Protocol Architecture

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: DECT"

In one line

DECT is a cordless system rather than a cellular one: short range, very high density, and no frequency planning at all, because a handset measures every one of the 240 physical channels and takes a quiet one for itself, in a frame of 24 slots on ten carriers where the same slot pair serves both directions ten milliseconds apart.

In the wording a student can write in an examination: DECT (Digital Enhanced Cordless Telecommunications) is an ETSI standard for cordless communication: short range, high density, low power, chiefly for cordless telephones, wireless PABXs and local loops. Its reference model runs from the portable part (PP), the handset, over the DECT air interface to the fixed part (FP), "physical grouping that contains all of the elements in the DECT network between the local network and the DECT air interface", whose radio elements are radio fixed parts (RFPs); the fixed part connects to a local network (a PABX, a LAN, the telephone network) and thence to a global network.

Its physical layer "divides the radio spectrum into the physical channels. This division occurs in two fixed dimensions, frequency and time", using "Time Division Multiple Access (TDMA) operation on multiple RF carriers", with "ten carriers ... within the actual frequency band (1 880 MHz to 1 900 MHz ...)". A frame lasts 10 ms and holds 24 slots; slots 0 to 11 carry the fixed part to the portable part and 12 to 23 the reverse, so duplex is by time division (TDD) on one carrier. The carrier rate is 1152 kbit/s, so a slot carries 480 bits, of which the B-field's 320 give a 32 kbit/s ADPCM speech channel. Ten carriers by 24 slots give 240 physical channels, that is 120 duplex pairs at one site.

Its layers are the physical layer (PHL), the MAC layer, which "selects physical channels, and then establishes and releases connections on those channels" and "multiplexes ... control information, together with higher layer information and error control information, into slot-sized packets", the DLC layer, "concerned with the provision of very reliable data links to the NWK layer", and the network (NWK) layer, with a management entity beside them. Channels are chosen by dynamic channel selection: the handset measures and picks, so no planning is needed.

What DECT is for

GSM covers a country; DECT covers a building. The consequences run through every design decision:

  • Range is tens to a few hundred metres, so the power is milliwatts and the cells are rooms.
  • Density is what matters: a hundred handsets in an office, all wanting a channel at once.
  • Planning must not be needed, because the buyer of a cordless phone will not plan frequencies, and neither will the installer of an office system.
  • Interference comes mostly from other DECT systems in the same building, which cannot be coordinated.
munotes.in732

DECT: System Architecture and Protocol Architecture

DECT's answers are a small band used intensively, time division duplex so that no paired spectrum is needed, and dynamic channel selection so that the equipment plans itself.

The reference model

The standard's terms are worth using exactly:

  • The portable part (PP) is the handset, which may be one of several sharing a subscription.
  • The fixed part (FP) is everything between the local network and the air: "physical grouping that contains all of the elements in the DECT network between the local network and the DECT air interface". Within it, a radio fixed part (RFP) is the radio end, and a central control fixed part (CCFP) is the "physical grouping that contains the central elements of a FP".
  • The local network is what the fixed part attaches to: a PABX in an office, a telephone line at home, a LAN.
  • The global network is beyond that: the public telephone network or the Internet.

Alongside those physical groupings the standard names two logical ones, and the protocol architecture is defined between them rather than between the boxes:

  • A portable radio termination (PT) is a "logical group of functions that contains all of the DECT processes and procedures on the portable side of the DECT air interface".
  • A fixed radio termination (FT) is the same "on the fixed side of the DECT air interface".

The distinction is not pedantry. A portable part is a piece of plastic with a battery; a portable radio termination is the set of protocol functions inside it that the standard actually specifies, and every layer in the next section is a peer relationship between a PT and an FT. One physical fixed part with several radio fixed parts still presents one fixed radio termination to a handset, which is why moving between radio fixed parts is a matter of changing bearers rather than of talking to a different party.

Between these sit interworking units, and the same air interface serves all of them, which is why DECT appears as a cordless phone, an office system and a wireless local loop without changing the radio.

The physical layer

Ten carriers, 2 MHz apart, in 1880 to 1900 MHz. A frame of 10 ms with 24 slots, so a slot is 417 microseconds, carrying 480 bits at 1152 kbit/s. The first twelve slots are the downlink and the last twelve the uplink, so a duplex bearer is a slot and its partner twelve slots later on the same carrier: time division duplex, which needs only one band and no duplex filter.

munotes.in733

DECT: System Architecture and Protocol Architecture

Inside a full slot, the fields are the S-field (preamble and synchronisation), the A-field (control, with its own error check), the B-field (320 bits of user data, with its check) and a guard. The B-field's 320 bits every 10 ms are exactly 32 kbit/s, which is one ADPCM speech channel: DECT does not compress speech as hard as GSM because it does not have to, and the quality is correspondingly better.

Ten carriers by 24 slots is 240 physical channels, 120 duplex pairs at one site. GSM gives 8 channels per carrier; DECT gives 24 half-duplex slots per carrier and can use all ten carriers at one base station because they are not planned for reuse.

Dynamic channel selection

This is DECT's cleverest idea, and the reason it needs no planning. Every portable part continuously scans all 240 physical channels and keeps a list of their interference levels. When a call is set up, or when quality falls, the handset chooses the quietest channel itself and tells the fixed part, which accepts if it can. If the chosen channel deteriorates, the handset performs a connection handover: it sets up a second bearer on a better channel and drops the old one, with no interruption to the call, which is DECT's version of seamless handover.

Two consequences. An installer can put base stations anywhere without a frequency plan, and two DECT systems from different vendors in the same building will simply avoid each other. The program measures the difference against a fixed plan.

The protocol layers

The standard sets them out plainly:

  • Physical layer (PHL): "divides the radio spectrum into the physical channels ... in two fixed dimensions, frequency and time".
  • MAC layer: "performs two main functions. Firstly, it selects physical channels, and then establishes and releases connections on those channels. Secondly, it multiplexes (and demultiplexes) control information, together with higher layer information and error control information, into slot-sized packets." It offers three services: broadcast, connection oriented and connectionless.
  • DLC layer: "concerned with the provision of very reliable data links to the NWK layer ... designed to work closely with the MAC layer to provide higher levels of data integrity than can be provided by the MAC layer alone."
  • Network layer (NWK): call control, mobility management, and the services above.

Beside all four sits a management entity, which is where DECT puts the functions that cross layers, such as choosing channels and deciding on handovers.

DECT, computed

The program computes the carrier spacing, the frame, the slot and its contents; counts the physical channels and duplex pairs; measures dynamic channel selection against a fixed plan as a building fills; and lists the layers.

munotes.in734

DECT: System Architecture and Protocol Architecture

# DECT's radio: carriers, the frame of 24 slots, and what one duplex pair
# carries. From ETSI EN 300 175-2.
CARRIERS, LOW, HIGH = 10, 1880.0, 1900.0
FRAME_MS, SLOTS = 10.0, 24
BITRATE = 1152e3                                  # bits a second on a carrier

print("DECT's physical layer (EN 300 175-2):")
print("  band %.0f to %.0f MHz, %d carriers, so %.3f MHz apart"
      % (LOW, HIGH, CARRIERS, (HIGH - LOW) / CARRIERS))
print("  a frame is %.0f ms and holds %d full slots, so a slot lasts %.4f ms"
      % (FRAME_MS, SLOTS, FRAME_MS / SLOTS))
print("  the carrier runs at %.0f kbit/s, so a slot carries %.0f bits"
      % (BITRATE / 1000, BITRATE * FRAME_MS / 1000 / SLOTS))
print("  slots 0 to 11 carry the fixed part to the portable part, 12 to 23 the other way,")
print("  so a duplex bearer is one slot and the slot 12 later: time division duplex.")

# Physical channels, and how many calls a base station can hold.
print("\nHow much capacity that gives:")
pairs = CARRIERS * SLOTS // 2
print("  %d carriers, %d slots each: %d physical channels, that is %d duplex pairs"
      % (CARRIERS, SLOTS, CARRIERS * SLOTS, pairs))
print("  each pair carries a 32 kbit/s speech channel (ADPCM), so one site can hold %d calls at once"
      % pairs)
print("  against GSM's 8 per carrier: DECT trades range for density, which is what a building needs")

# 2. The slot's own budget: what a full slot carries.
print("\nInside a full slot (%d bits):" % int(BITRATE * FRAME_MS / 1000 / SLOTS))
for name, bits in (("S-field: preamble and sync word", 32), ("A-field: control (with its CRC)", 64),
                   ("B-field: user data (with its CRC)", 320), ("Z-field and guard", 64)):
    print("  %-38s %4d bits" % (name, bits))
print("  the B-field's 320 bits every 10 ms are the %.0f kbit/s of one speech channel"
      % (320 / (FRAME_MS / 1000) / 1000))

# 3. Dynamic channel selection: a portable part picks the quietest channel
#    itself, instead of being assigned one. What that buys in a building.
import random
rnd = random.Random(96)
print("\nDynamic channel selection against a fixed plan (a building, 200 trials):")
print("  handsets   fixed plan: blocked   DECT's own choice: blocked")
for handsets in (10, 30, 60, 100, 120):
    fixed_blocked = dynamic_blocked = 0
    for _ in range(200):
        plan = [0] * 40                            # a fixed plan giving 40 channels
        free = pairs
        for _ in range(handsets):
            k = rnd.randrange(40)                  # a fixed plan assigns by identity
            if plan[k]:
                fixed_blocked += 1
            else:
                plan[k] = 1
            if free:
                free -= 1
            else:
                dynamic_blocked += 1
    n = 200 * handsets
    print("  %8d %20.1f%% %26.1f%%" % (handsets, 100 * fixed_blocked / n, 100 * dynamic_blocked / n))
print("  a fixed plan collides long before the capacity is used; DECT's handsets measure every")
print("  channel and take a free one, which is why a base station needs no planning at all.")

# 4. The protocol layers, and what each does (EN 300 175-1, clause 7).
print("\nDECT's layers:")
for layer, what in (("physical (PHL)", "divides the spectrum into physical channels, in frequency and time"),
                    ("MAC", "selects physical channels, establishes and releases connections, multiplexes"),
                    ("DLC", "provides very reliable data links to the network layer"),
                    ("network (NWK)", "call control, mobility, and the services above")):
    print("  %-16s %s" % (layer, what))
print("  and a management entity beside them all, which is where DECT puts what does not fit a layer.")
munotes.in735

DECT: System Architecture and Protocol Architecture

DECT's physical layer (EN 300 175-2):
  band 1880 to 1900 MHz, 10 carriers, so 2.000 MHz apart
  a frame is 10 ms and holds 24 full slots, so a slot lasts 0.4167 ms
  the carrier runs at 1152 kbit/s, so a slot carries 480 bits
  slots 0 to 11 carry the fixed part to the portable part, 12 to 23 the other way,
  so a duplex bearer is one slot and the slot 12 later: time division duplex.

How much capacity that gives:
  10 carriers, 24 slots each: 240 physical channels, that is 120 duplex pairs
  each pair carries a 32 kbit/s speech channel (ADPCM), so one site can hold 120 calls at once
  against GSM's 8 per carrier: DECT trades range for density, which is what a building needs

Inside a full slot (480 bits):
  S-field: preamble and sync word          32 bits
  A-field: control (with its CRC)          64 bits
  B-field: user data (with its CRC)       320 bits
  Z-field and guard                        64 bits
  the B-field's 320 bits every 10 ms are the 32 kbit/s of one speech channel

Dynamic channel selection against a fixed plan (a building, 200 trials):
  handsets   fixed plan: blocked   DECT's own choice: blocked
        10                 10.2%                        0.0%
        30                 29.6%                        0.0%
        60                 47.9%                        0.0%
       100                 63.2%                        0.0%
       120                 68.3%                        0.0%
  a fixed plan collides long before the capacity is used; DECT's handsets measure every
  channel and take a free one, which is why a base station needs no planning at all.

DECT's layers:
  physical (PHL)   divides the spectrum into physical channels, in frequency and time
  MAC              selects physical channels, establishes and releases connections, multiplexes
  DLC              provides very reliable data links to the network layer
  network (NWK)    call control, mobility, and the services above
  and a management entity beside them all, which is where DECT puts what does not fit a layer.

The frame. Ten carriers 2 MHz apart; a slot of 417 microseconds carrying 480 bits; twelve slots each way. Of those 480 bits, 320 are user data, which is 32 kbit/s, and the rest is synchronisation, control and guard.

munotes.in736

DECT: System Architecture and Protocol Architecture

Capacity. 240 physical channels, 120 duplex pairs, at one site. That is a density GSM cannot approach, and it is possible only because the range is short: the same channels are reused in the next building without any coordination.

Dynamic selection. With a fixed plan of 40 channels assigned by the handset's identity, collisions begin immediately: 10.2 per cent of 10 handsets are blocked, 47.9 per cent of 60. With DECT's own scheme, where each handset takes a free channel from the 120 available, nothing is blocked until the capacity is genuinely exhausted. The comparison is deliberately crude, but the mechanism it shows is real: measuring beats planning when the environment cannot be planned.

The layers. Four, plus a management entity, and the division of labour is clear from the standard's own sentences: the physical layer makes channels, the MAC chooses and fills them, the DLC makes them reliable, and the network layer runs calls.

Distinctions

GSMDECT
CoversA country, in cells of kilometresA building, in cells of tens of metres
Band890 to 960 MHz and others, paired1880 to 1900 MHz, unpaired
DuplexFrequency division (45 MHz apart)Time division (12 slots apart)
Frame4.615 ms, 8 slots10 ms, 24 slots
Channels at a site8 per carrier, a few carriers240 physical, 120 duplex
Speech13 kbit/s, heavily coded32 kbit/s ADPCM
PlanningFrequency plan, cluster, reuse distanceNone: dynamic channel selection
HandoverNetwork decides, hardHandset decides, connection handover without a break
LayerDoes
PHLDivides the spectrum into physical channels, in frequency and time
MACSelects physical channels, establishes and releases connections, multiplexes into slot-sized packets
DLCProvides very reliable data links to the network layer
NWKCall control, mobility, services
Management entityWhat crosses layers: channel selection, handover decisions
Fixed part (FP)Portable part (PP)
IsEverything between the local network and the airThe handset
ContainsRadio fixed parts, a central control fixed partOne radio
Connects toThe local network, then the global networkThe fixed part it has rights on

What it does not mean

DECT is not a cellular system. It has no frequency plan, no location registers of the GSM kind, and no expectation of wide-area mobility.

Time division duplex is not half duplex. Both directions are carried, twelve slots apart, which the user cannot perceive.

Dynamic channel selection is not contention. The handset measures before choosing, so two handsets rarely pick the same channel, and if they do the MAC recovers.

DECT's 32 kbit/s is not better coding. It is less compression, which a short link can afford and which is why DECT speech sounds better than GSM speech.

munotes.in737

DECT: System Architecture and Protocol Architecture

A fixed part is not a base station. It is the whole fixed installation; the radio unit is a radio fixed part within it.

Quick revision

  • DECT: cordless, short range, high density, no planning. PP (handset), FP (all between the local network and the air; RFP radios, CCFP central control), local network, global network.
  • PHL: 1880 to 1900 MHz, 10 carriers 2 MHz apart, frame 10 ms, 24 slots (0 to 11 down, 12 to 23 up), TDD; 1152 kbit/s, slot 480 bits, B-field 320 bits = 32 kbit/s ADPCM; 240 physical channels, 120 duplex pairs at one site.
  • Dynamic channel selection: the handset scans all channels and picks the quietest; connection handover moves a call to a better channel without a break.
  • Layers: PHL (makes channels), MAC (selects, connects, multiplexes; broadcast, connection oriented and connectionless services), DLC (very reliable links), NWK (call control, mobility), plus a management entity.
  • Program: a fixed plan blocks 47.9 per cent of 60 handsets; dynamic selection blocks none until the 120 pairs are used.

Test yourself

1. What is DECT for, and how does it differ from a cellular system such as GSM? DECT is a standard for cordless communication over short ranges at high density: cordless telephones, wireless PABXs and wireless local loops. Unlike GSM it covers a building rather than a country, uses tens of milliwatts over tens of metres, works in a single unpaired band with time division duplex rather than paired bands, offers 240 physical channels at one site rather than eight per carrier, uses far less speech compression, and, most importantly, requires no frequency planning at all, because handsets select channels dynamically instead of being assigned planned ones.

2. Describe DECT's reference model. The portable part is the handset. The fixed part is the physical grouping containing all the elements of the DECT network between the local network and the air interface; within it the radio fixed parts provide the radio and the central control fixed part contains the central elements. The fixed part attaches to a local network, such as a PABX, a telephone line or a LAN, which in turn connects to a global network such as the public telephone network. The same air interface serves all these arrangements.

3. Give DECT's physical layer parameters and explain time division duplex. Ten carriers 2 MHz apart in the band 1880 to 1900 MHz; a TDMA frame of 10 ms divided into 24 slots, so a slot lasts about 417 microseconds; a carrier rate of 1152 kbit/s, so a slot carries 480 bits, of which the B-field's 320 bits every 10 ms give a 32 kbit/s ADPCM speech channel. Slots 0 to 11 carry the fixed part to the portable part and slots 12 to 23 the reverse, so a duplex connection uses one slot and the slot twelve later on the same carrier: both directions share one frequency, separated in time, which is time division duplex and needs no paired spectrum and no duplex filter.

munotes.in738

DECT: System Architecture and Protocol Architecture

4. What is dynamic channel selection, and why does DECT need it? Every portable part continuously scans all 240 physical channels and records their interference levels; when a connection is needed it chooses the quietest channel itself and proposes it to the fixed part. It is needed because DECT equipment is installed by people who will not plan frequencies, and because the main source of interference is other DECT systems in the same building, which no one can coordinate. It also allows connection handover: if the chosen channel deteriorates, the handset establishes a bearer on a better channel and releases the old one without interrupting the call.

5. Name DECT's protocol layers and say what each does. The physical layer divides the radio spectrum into physical channels in the two dimensions of frequency and time. The MAC layer selects physical channels and establishes and releases connections on them, and multiplexes control information, higher layer information and error control into slot-sized packets, offering broadcast, connection oriented and connectionless services. The data link control layer provides very reliable data links to the network layer, working closely with the MAC layer to achieve greater data integrity than the MAC alone can. The network layer provides call control, mobility management and the services above. A management entity beside these handles functions that cross layers, such as channel selection and handover.

6. How many simultaneous calls can one DECT base station support, and why so many more than a GSM cell? Ten carriers with 24 slots each give 240 physical channels, which is 120 duplex pairs, so in principle 120 simultaneous speech channels at one site. A GSM carrier gives eight channels and a cell has only the carriers its reuse plan allows. DECT can use all ten of its carriers at every site because its range is only tens of metres, so the same channels are reused in the next building with no coordination; it trades range for density, which is exactly the trade an office needs.

Contents This chapter on its own page

munotes.in739

Chapter Ninety-Seven

TETRA

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: TETRA"

In one line

TETRA is radio for people whose job depends on it: the channels are pooled so that a pressed transmit key almost always finds one, a whole talkgroup hears one transmission on one channel, the call sets up in less than half a second, an emergency can take a channel from an ordinary user, and when the network is gone the radios talk to each other directly.

In the wording a student can write in an examination: TETRA (Terrestrial Trunked Radio) is an ETSI standard for professional mobile radio, used by police, fire, ambulance, transport and utilities. Its access scheme is TDMA with 4 physical channels per carrier, and "For phase modulation the carrier bandwidth is 25 kHz", with a modulation rate of 36 kbit/s using pi/4-DQPSK. Its frame structure, as the standard draws it, is one TDMA frame of 4 timeslots, about 56.67 ms; one multiframe of 18 TDMA frames, 1.02 s, of which frame 18 is the control frame; and one hyperframe of 60 multiframes, 61.2 s.

Trunking means that channels are held in a common pool and assigned to a call when it is made, rather than each user group having its own channel; because demand averages out, a pool serves far more traffic than the same channels split up. The standard distinguishes message trunked, transmission trunked and quasi-transmission trunked systems, which differ in whether the channel is held for a whole conversation or released after each transmission.

Its distinguishing services are the group call (one transmission heard by a whole talkgroup, using one downlink channel however many listeners there are), the broadcast call, push to talk half-duplex working, call setup in well under a second, priority and pre-emption, late entry (a radio switched on during a call joins it), direct mode operation (DMO), in which radios talk to each other with no network, and end-to-end encryption in addition to the air interface's own. It carries voice plus data (the standard's "V+D"), including short data and packet data.

Why not just use a phone network

A public cellular network is optimised for one-to-one conversations between strangers, set up over several seconds, with best-effort service and a single grade of user. A fire crew needs none of those things and cannot use a system that offers only them:

  • The natural unit is the group, not the pair. Twelve firefighters need to hear one another, and a public network can only build that as a conference bridge, badly and slowly.
  • The natural act is pressing a key and speaking, immediately. A setup time of seconds is unusable when the message is "get out".
  • Priority must be real: an emergency call must be able to take a channel from somebody else.
  • The network must work when everything else has failed, which includes working with no network at all.
  • The user often must not trust the operator with the content, which means end-to-end encryption under the user's own keys.
munotes.in740

TETRA

TETRA is what those requirements produce.

The radio

Four channels per 25 kHz carrier with pi/4-DQPSK at 36 kbit/s, so a channel is about 9 kbit/s gross, of which the speech codec uses 4.567 kbit/s and the rest is protection. That is a quarter of GSM's per-channel rate, and it is deliberate: TETRA's priority is range and robustness, not rate. Its spectral efficiency, four channels in 25 kHz, is four times GSM's eight in 200 kHz, which matters because professional users are allocated narrow bands.

The frame structure is unusual in one respect: the control frame, frame 18 of every multiframe, is reserved for signalling, so there is always a known moment at which a radio can be reached or can ask for something. That is part of how call setup stays fast.

Trunking

A traditional professional radio system gave each user group its own channel: the fire service had one, the water board another. Half the channels were idle while another group queued.

Trunking pools them. A radio that wants to talk asks a control channel, and is assigned any free traffic channel for the duration. The gain is the trunking efficiency of [Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing], and the program measures it: five groups of four channels carry 5.46 erlangs in all at 2 per cent blocking; the same twenty channels pooled carry 13.18, 141 per cent more.

The standard distinguishes how long a channel is held:

  • Message trunked: the channel is held for the whole conversation, including the pauses between transmissions. Simple, and wasteful of the pauses.
  • Transmission trunked: the channel is released at the end of every transmission and requested again for the next. Efficient, at the cost of a small delay and the risk that the channel is gone when the reply comes.
  • Quasi-transmission trunked: a compromise, holding the channel briefly after each transmission in case the conversation continues.

The group call

This is the service that defines the system. A talkgroup is a set of radios that share an identity; a transmission to the group goes out once on one downlink channel, and every radio in the group in that cell hears it. The program counts what the alternatives would cost: twelve people talking to each other as individual calls would need 66 channels for everyone to hear everyone, a network conference would need twelve, and a TETRA group call needs one downlink and one uplink.

munotes.in741

TETRA

Around it sit the features that make it usable:

  • Push to talk: the channel is half duplex and one radio at a time transmits, which is not a limitation but the discipline the users already have.
  • Fast setup: because the group exists in advance and the control frame comes round every multiframe, the standard's target is well under half a second.
  • Late entry: a radio switched on, or returning to coverage, during a group call joins it, because the group call is a state of the network rather than a connection between endpoints.
  • Priority and pre-emption: calls carry priorities, and an emergency call can take a channel from a lower-priority call in progress.

Direct mode

Direct mode operation lets two radios talk on a dedicated channel with no network at all: in a tunnel, a basement, a building whose coverage has failed, or an area with no infrastructure. The range is that of a handset, so hundreds of metres to a few kilometres, and one radio may act as a gateway, relaying direct-mode users into the network, or a repeater, extending direct-mode range.

Nothing in GSM or DECT corresponds to this, and it is the plainest statement of what TETRA is for: a system whose users cannot accept that the infrastructure is a single point of failure.

TETRA, computed

The program computes the frame hierarchy and the per-channel rate; prices trunking against separate channel groups; counts the channels a group call saves; lists direct mode's properties; and tabulates what a public network cannot offer.

# TETRA: the frame, what trunking buys, and why a group call is not many calls.
import math

# 1. The TDMA structure, from EN 300 392-2 clause 4.5.2.
FRAME_MS, SLOTS, MULTI, HYPER = 56.67, 4, 18, 60
print("TETRA's TDMA structure (EN 300 392-2, 4.5.2):")
print("  1 TDMA frame  = %d timeslots      = %.2f ms, so a slot is %.2f ms"
      % (SLOTS, FRAME_MS, FRAME_MS / SLOTS))
print("  1 multiframe  = %d TDMA frames   = %.2f s (frame 18 is the control frame)"
      % (MULTI, MULTI * FRAME_MS / 1000))
print("  1 hyperframe  = %d multiframes   = %.1f s" % (HYPER, HYPER * MULTI * FRAME_MS / 1000))
print("  modulation rate %d kbit/s in a %d kHz carrier with pi/4-DQPSK, 4 channels a carrier"
      % (36, 25))
print("  so one channel has about %.1f kbit/s gross, against GSM's %.1f in 200 kHz"
      % (36 / 4, 270.833 / 8))
print("  spectral efficiency: TETRA %.2f channels a kHz, GSM %.2f"
      % (4 / 25, 8 / 200))

# 2. Trunking: why a shared pool of channels beats fixed assignments, which is
#    the whole idea of a trunked radio system. Erlang B again.
def erlang_b(c, a):
    b = 1.0
    for k in range(1, c + 1):
        b = a * b / (k + a * b)
    return b

print("\nTrunking: 5 user groups, 4 channels each, against one pool of 20 (2 per cent blocking):")
def capacity(c, target=0.02):
    lo, hi = 0.0, 1000.0
    for _ in range(60):
        mid = (lo + hi) / 2
        lo, hi = (mid, hi) if erlang_b(c, mid) < target else (lo, mid)
    return lo
sep, pooled = 5 * capacity(4), capacity(20)
print("  five separate groups of 4 channels: %.2f erlangs in all" % sep)
print("  one trunked pool of 20 channels:    %.2f erlangs, %.0f%% more" % (pooled, 100 * (pooled / sep - 1)))
print("  that is trunking, and it is what the T in TETRA stands for.")

# 3. A group call against many individual calls: how many channels each needs.
print("\nOne fire crew of 12 talking to each other:")
print("  individual calls, everyone to everyone: %d pairs, so %d channels for all to hear all"
      % (12 * 11 // 2, 12 * 11 // 2))
print("  a conference bridged in the network: 12 channels, one per member")
print("  a TETRA group call: 1 downlink channel for all, 1 uplink for whoever is speaking: 2")
print("  and the call sets up in under half a second, because the group already exists")

# 4. Direct mode: two radios with no network at all.
print("\nDirect mode operation (DMO), and why it matters:")
for point in ("two radios talk directly, on their own channel, with no base station",
              "it works in a tunnel, a basement, or where the network has failed",
              "range is that of a handset, so a few hundred metres to a few kilometres",
              "a radio can act as a gateway, relaying direct-mode users into the network"):
    print("  - %s" % point)

# 5. What TETRA has that a public network does not.
print("\nWhat a public cellular network cannot offer, and TETRA does:")
for feature, why in (("group call and broadcast call", "one transmission heard by a whole talkgroup"),
                     ("call setup in under 0.5 s", "a group already exists; no per-call negotiation"),
                     ("push to talk, half duplex", "one speaks, many listen: the discipline of radio"),
                     ("pre-emptive priority", "an emergency call takes a channel from an ordinary one"),
                     ("direct mode", "works with no infrastructure at all"),
                     ("air interface encryption plus end to end", "the operator is not trusted with the content")):
    print("  %-38s %s" % (feature, why))
munotes.in742

TETRA

TETRA's TDMA structure (EN 300 392-2, 4.5.2):
  1 TDMA frame  = 4 timeslots      = 56.67 ms, so a slot is 14.17 ms
  1 multiframe  = 18 TDMA frames   = 1.02 s (frame 18 is the control frame)
  1 hyperframe  = 60 multiframes   = 61.2 s
  modulation rate 36 kbit/s in a 25 kHz carrier with pi/4-DQPSK, 4 channels a carrier
  so one channel has about 9.0 kbit/s gross, against GSM's 33.9 in 200 kHz
  spectral efficiency: TETRA 0.16 channels a kHz, GSM 0.04

Trunking: 5 user groups, 4 channels each, against one pool of 20 (2 per cent blocking):
  five separate groups of 4 channels: 5.46 erlangs in all
  one trunked pool of 20 channels:    13.18 erlangs, 141% more
  that is trunking, and it is what the T in TETRA stands for.

One fire crew of 12 talking to each other:
  individual calls, everyone to everyone: 66 pairs, so 66 channels for all to hear all
  a conference bridged in the network: 12 channels, one per member
  a TETRA group call: 1 downlink channel for all, 1 uplink for whoever is speaking: 2
  and the call sets up in under half a second, because the group already exists

Direct mode operation (DMO), and why it matters:
  - two radios talk directly, on their own channel, with no base station
  - it works in a tunnel, a basement, or where the network has failed
  - range is that of a handset, so a few hundred metres to a few kilometres
  - a radio can act as a gateway, relaying direct-mode users into the network

What a public cellular network cannot offer, and TETRA does:
  group call and broadcast call          one transmission heard by a whole talkgroup
  call setup in under 0.5 s              a group already exists; no per-call negotiation
  push to talk, half duplex              one speaks, many listen: the discipline of radio
  pre-emptive priority                   an emergency call takes a channel from an ordinary one
  direct mode                            works with no infrastructure at all
  air interface encryption plus end to end the operator is not trusted with the content
munotes.in743

TETRA

The frame. A slot is 14.17 ms, a frame 56.67 ms, a multiframe 1.02 s with a control frame at its end, a hyperframe 61.2 s. Four channels in 25 kHz gives 0.16 channels per kHz against GSM's 0.04: four times the spectral efficiency, because the channels are slower.

Trunking. 141 per cent more traffic from the same twenty channels, purely by pooling them. That single number is why trunked radio replaced allocated channels, and it is the same trunking effect that makes a large cell better than several small ones for a given load.

The group call. Twelve people, 66 pairs, against two channels. A conference bridge in a public network would consume twelve channels and take seconds to build; TETRA consumes one downlink and one uplink and is ready before the speaker has finished pressing the key.

munotes.in744

TETRA

What is missing elsewhere. Six features, and every one of them follows from the users: group working, speed, discipline, priority, independence from infrastructure, and distrust of the operator with the content.

Distinctions

TETRAGSM
UsersPolice, fire, ambulance, transport, utilitiesThe public
Carrier and channels25 kHz, 4 channels (TDMA)200 kHz, 8 channels
Channel rateAbout 9 kbit/s grossAbout 34 kbit/s gross
Natural callGroup, half duplex, push to talkOne to one, full duplex
SetupWell under half a secondSeconds
PriorityYes, with pre-emptionNo
Without a networkDirect modeNothing
EncryptionAir interface plus end to endAir interface only
Trunking modeThe channel is heldSuits
Message trunkedFor the whole conversationShort exchanges, simple systems
Transmission trunkedFor one transmission onlyBusy systems, where pauses are long
Quasi-transmission trunkedFor a short time after each transmissionA compromise, the usual choice
Individual callGroup callBroadcast call
Who hearsOne partyEveryone in the talkgroupEveryone addressed
Who may speakBothOne at a time, by push to talkThe caller only
ChannelsOne pairOne down, one up, per cellOne down
SetupFastFast: the group already existsFast

What it does not mean

Trunked does not mean high capacity per user. It means channels are pooled; each channel is slower than GSM's.

Push to talk is not a limitation of the radio. It is the working discipline of group communication, and the system is built around it.

Direct mode is not a fallback of last resort only. It is used routinely, for example inside a building where a crew is working.

A group call is not a conference call. It is one transmission to many radios, not many connections bridged together.

End-to-end encryption is not the same as air interface encryption. TETRA has both, and the second is what keeps the content from the operator as well as from an eavesdropper.

Quick revision

  • TETRA: professional mobile radio; TDMA, 4 channels per carrier, 25 kHz for phase modulation, 36 kbit/s with pi/4-DQPSK, so about 9 kbit/s a channel.
  • Frames: slot 14.17 ms, TDMA frame 4 slots = 56.67 ms, multiframe 18 frames = 1.02 s (frame 18 is the control frame), hyperframe 60 multiframes = 61.2 s.
  • Trunking: channels pooled and assigned per call. Modes: message, transmission, quasi-transmission trunked. Program: 20 channels pooled carry 141 per cent more than five groups of four.
  • Services: group call (one downlink for the whole talkgroup), broadcast call, push to talk, setup under half a second, late entry, priority and pre-emption, direct mode (DMO) with gateways and repeaters, voice plus data, end-to-end encryption.
  • Program: a crew of 12 needs 66 channels as individual calls, 12 as a bridged conference, 2 as a TETRA group call.
munotes.in745

TETRA

Test yourself

1. What is TETRA and who uses it? TETRA, Terrestrial Trunked Radio, is an ETSI standard for professional mobile radio, used by emergency services, transport operators, utilities and other organisations whose work depends on radio. It is a trunked system carrying voice plus data, designed around group working, immediate call setup, priority, and operation with or without infrastructure, none of which a public cellular network provides.

2. Give TETRA's TDMA structure. The access scheme is TDMA with four physical channels per carrier, and for phase modulation the carrier bandwidth is 25 kHz with a modulation rate of 36 kbit/s using pi/4-DQPSK. One TDMA frame is four timeslots and lasts about 56.67 ms, so a slot is about 14.17 ms; one multiframe is 18 TDMA frames, 1.02 seconds, of which the eighteenth is a control frame reserved for signalling; and one hyperframe is 60 multiframes, 61.2 seconds.

3. What is trunking, and what does it achieve? In a trunked system the radio channels are held in a common pool and assigned to a call when it is made, instead of each user group having permanently allocated channels. Because peaks in different groups' demands do not coincide, a pool carries far more traffic than the same number of channels divided up: in the chapter's calculation, five groups of four channels carry 5.46 erlangs at 2 per cent blocking while the same twenty channels pooled carry 13.18, 141 per cent more.

4. Distinguish message, transmission and quasi-transmission trunking. In a message trunked system a channel is assigned for the whole conversation and held through the pauses between transmissions, which is simple but wastes the pauses. In a transmission trunked system the channel is released at the end of each transmission and requested again for the next, which uses the pauses for other users but adds a small delay and a risk that no channel is free for the reply. Quasi-transmission trunking holds the channel for a short time after each transmission in case the conversation continues, which is a compromise between the two.

5. Why is a group call fundamentally different from a conference call? A group call is one transmission addressed to a talkgroup: the base station sends it once on a single downlink channel and every radio of that group in the cell receives it, whatever their number, while one radio at a time transmits on a single uplink channel. A conference call is many separate connections bridged together in the network, costing one channel per participant and taking as long to build as the slowest setup. For a crew of twelve, the chapter counts 66 channels for individual calls, 12 for a bridged conference and 2 for a TETRA group call, and the group call is ready in under half a second because the group already exists.

munotes.in746

TETRA

6. What is direct mode operation, and why does it matter? In direct mode two or more radios communicate with each other on a dedicated channel without any base station or network. It matters because the users cannot accept infrastructure as a single point of failure: direct mode works in tunnels, basements and buildings where coverage fails, in areas with no network at all, and when the network itself has been destroyed or overloaded. Its range is that of a handset, so a few hundred metres to a few kilometres, and a radio may act as a gateway into the network or as a repeater to extend the direct-mode range.

Contents This chapter on its own page

munotes.in747

Chapter Ninety-Eight

UMTS and IMT-2000

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: UMTS and IMT-2000"

In one line

IMT-2000 is not a system but a list of demands the ITU made of any third generation system, chief among them data rates that fall as the user moves faster, and UMTS is the European answer: a new radio interface, W-CDMA, and a new access network bolted onto an evolved GSM core, which then evolved through a decade of releases into HSPA, LTE and eventually 5G.

In the wording a student can write in an examination: IMT-2000 (International Mobile Telecommunications 2000) is the ITU's framework for third generation mobile systems, a family of radio interfaces meeting a common set of requirements rather than a single standard. Its headline requirements were data rates of at least 144 kbit/s for a user in a fast-moving vehicle, 384 kbit/s for a pedestrian and 2 Mbit/s for a nearly stationary indoor user, together with global roaming, support for packet and circuit services, and quality of service differentiation. Members of the family include UMTS with its W-CDMA radio interface (the European and Japanese choice), cdma2000 (the American evolution of IS-95) and TD-SCDMA (China's).

UMTS (Universal Mobile Telecommunications System) is specified by 3GPP. Its first release, Release 99, kept the GSM core network, extended for packets by GPRS, and replaced the radio with the UTRAN (UMTS terrestrial radio access network) based on W-CDMA in 5 MHz carriers. Its own architecture specification divides a network into domains: the user equipment domain (the UE, itself a mobile equipment domain and a USIM domain) and the infrastructure domain, split into the access network domain and the core network domain, with the Uu reference point between user equipment and infrastructure. The design principle stated in that specification is modularity, because "it is not possible to optimise UMTS to only one set of applications", so "UMTS must be built in such a way that it is flexible and possible to evolve so it will have a long technical lifetime."

Later releases added HSDPA and HSUPA (together HSPA), HSPA+, then LTE with the all-IP evolved packet core, LTE-Advanced, which met the ITU's next family IMT-Advanced (100 Mbit/s at high mobility, 1 Gbit/s at low), and 5G New Radio, which addresses IMT-2020 (20 Gbit/s peak, 1 ms latency, a million devices per square kilometre). IMT-2030 is the ITU's framework for the sixth generation.

IMT-2000: requirements, not a system

The ITU's role is to say what a generation must do, not how. IMT-2000 set the requirements, allocated spectrum around 2 GHz at the 1992 and 2000 world radiocommunication conferences, and then accepted several radio interfaces as members of the family, so that a "3G" phone in Europe and one in America could be built to different standards and still be third generation.

munotes.in748

UMTS and IMT-2000

Its rate requirements are worth reading as an engineering statement rather than a marketing one: 144 kbit/s moving fast, 384 kbit/s walking, 2 Mbit/s indoors and nearly still. The pattern is deliberate. A fast-moving user must be served by a large cell, because handovers would otherwise be constant; a large cell shares its capacity among many users and its edge is far away; so the rate per user must be lower. A stationary indoor user can be served by a small cell with few users and a short path. Rate and mobility trade against each other, and every generation since has set its targets in the same shape.

What UMTS kept and what it replaced

UMTS is frequently described as a new system; it is better understood as a new radio access network on an evolved core.

  • Kept: the core network, in Release 99, is the GSM core with its MSC, VLR, HLR and, for packets, the SGSN and GGSN of GPRS ([New Data Services: HSCSD, GPRS and EDGE]). The service model, the subscriber identities, the registers and the roaming machinery are GSM's.
  • Replaced: the radio. W-CDMA in 5 MHz carriers replaces GSM's 200 kHz carriers and time slots; every cell uses every carrier, so the cluster size is 1; soft handover becomes possible; and the SIM becomes the USIM, with mutual authentication answering [GSM Security]'s largest weakness.

That division is why an operator could deploy UMTS on its existing core, and why the two systems interwork: a call can be handed from UMTS to GSM as coverage runs out.

Domains

TS 23.101 defines the separation "common to all UMTS networks independent of their origin":

  • The user equipment domain, containing the mobile equipment domain and the USIM domain, which is the subscription.
  • The infrastructure domain, containing the access network domain, which is everything specific to the radio access technology, and the core network domain, which is not.

The point of the split is stated in the introduction: applications cannot be predicted, so the network must be modular and long-lived, and a module is "one or several physical network nodes that together implements some functionality". The proof of that principle is what happened next: the access network was replaced twice, by LTE's E-UTRAN and by 5G's NG-RAN, while the core evolved separately.

The releases

3GPP does not version its system by generation but by release, and the releases are how the subject is actually organised.

  • Release 99: UMTS as first deployed. W-CDMA, the UTRAN, the GSM and GPRS core.
  • Release 4: the MSC splits into a server (control) and a media gateway (traffic), the first step toward a packet core.
  • Release 5: HSDPA, which adds a fast shared downlink channel with adaptive modulation and fast scheduling, and the IP multimedia subsystem.
  • Release 6: HSUPA, the same for the uplink. Together, HSPA.
  • Release 7: HSPA+, higher order modulation and MIMO.
  • Release 8: LTE and the evolved packet core: OFDMA downlink, all-IP, no circuit domain at all.
  • Release 10: LTE-Advanced, carrier aggregation and more MIMO, which is what actually met IMT-Advanced.
  • Release 15: 5G New Radio.
munotes.in749

UMTS and IMT-2000

After IMT-2000

  • IMT-Advanced (2008): 100 Mbit/s in high mobility, 1 Gbit/s in low. LTE-Advanced and WirelessMAN-Advanced met it.
  • IMT-2020 (2015): 20 Gbit/s peak downlink, 1 ms user-plane latency, a million connected devices per square kilometre, and three usage scenarios, enhanced mobile broadband, ultra-reliable low-latency communication and massive machine-type communication. The last of those is where the sensor networks of Module 1 meet the cellular world: NB-IoT and LTE-M are its answers, beside the 6LoWPAN and TSCH of [Built on 802.15.4: Zigbee Routing, Security and the Later Amendments].
  • IMT-2030 (2023): the framework for the sixth generation.

The generations, computed

The program sets each generation's rate against a 5 MB photograph, restates IMT-2000's requirements by mobility, lists the releases, and lists the ITU's families.

# IMT-2000 and what came after it: the rate targets of each generation, and
# what those rates actually allow.
print("The rate a generation promised, and how long a 5 MB photograph takes at it:")
gens = [("GSM circuit data", 9.6e3), ("GPRS, in practice", 40e3), ("EDGE, in practice", 150e3),
        ("IMT-2000: moving vehicle", 144e3), ("IMT-2000: pedestrian", 384e3),
        ("IMT-2000: indoor", 2e6), ("IMT-Advanced: high mobility", 100e6),
        ("IMT-Advanced: low mobility", 1e9), ("IMT-2020: peak downlink", 20e9)]
for name, rate in gens:
    seconds = 5 * 1024 * 1024 * 8 / rate
    when = "%.0f s" % seconds if seconds >= 1 else "%.0f ms" % (seconds * 1000)
    if seconds > 120:
        when = "%.1f minutes" % (seconds / 60)
    print("  %-30s %10s %14s" % (name, "%g kb/s" % (rate / 1e3) if rate < 1e9 else "%g Gb/s" % (rate / 1e9), when))

print("\nWhat IMT-2000 asked for, by how fast the user is moving:")
for setting, rate in (("a vehicle at speed", 144e3), ("a pedestrian", 384e3), ("indoors, nearly still", 2e6)):
    print("  %-22s at least %6.0f kbit/s" % (setting, rate / 1e3))
print("  the rule behind it: the faster the user, the larger the cell must be, and the larger the")
print("  cell, the more users share the same spectrum, so the rate each can have falls.")

print("\nThe 3GPP releases, and what each added:")
for rel, year, what in (("Release 99", 1999, "UMTS: the W-CDMA radio interface and the UTRAN, on the GSM core"),
                        ("Release 4", 2001, "the split MSC: server and media gateway"),
                        ("Release 5", 2002, "HSDPA, and the IP multimedia subsystem"),
                        ("Release 6", 2004, "HSUPA, completing high speed packet access"),
                        ("Release 7", 2007, "HSPA+ : higher order modulation and MIMO"),
                        ("Release 8", 2008, "LTE and the evolved packet core: all IP, no circuits"),
                        ("Release 10", 2011, "LTE-Advanced, which met IMT-Advanced"),
                        ("Release 15", 2018, "5G New Radio, the first 5G release")):
    print("  %-12s %4d  %s" % (rel, year, what))

print("\nThe ITU's families, and what each demanded:")
for name, year, demand in (("IMT-2000", 1999, "144 kbit/s moving, 384 pedestrian, 2 Mbit/s indoors"),
                           ("IMT-Advanced", 2008, "100 Mbit/s high mobility, 1 Gbit/s low mobility"),
                           ("IMT-2020", 2015, "20 Gbit/s peak downlink, 1 ms latency, a million devices a square km"),
                           ("IMT-2030", 2023, "the framework for the sixth generation")):
    print("  %-14s %4d  %s" % (name, year, demand))
print("  each is a set of requirements from the ITU; the systems that meet them are built by others,")
print("  which is why UMTS is a member of IMT-2000 rather than the same thing as it.")
munotes.in750

UMTS and IMT-2000

The rate a generation promised, and how long a 5 MB photograph takes at it:
  GSM circuit data                 9.6 kb/s   72.8 minutes
  GPRS, in practice                 40 kb/s   17.5 minutes
  EDGE, in practice                150 kb/s    4.7 minutes
  IMT-2000: moving vehicle         144 kb/s    4.9 minutes
  IMT-2000: pedestrian             384 kb/s          109 s
  IMT-2000: indoor                2000 kb/s           21 s
  IMT-Advanced: high mobility    100000 kb/s         419 ms
  IMT-Advanced: low mobility         1 Gb/s          42 ms
  IMT-2020: peak downlink           20 Gb/s           2 ms

What IMT-2000 asked for, by how fast the user is moving:
  a vehicle at speed     at least    144 kbit/s
  a pedestrian           at least    384 kbit/s
  indoors, nearly still  at least   2000 kbit/s
  the rule behind it: the faster the user, the larger the cell must be, and the larger the
  cell, the more users share the same spectrum, so the rate each can have falls.

The 3GPP releases, and what each added:
  Release 99   1999  UMTS: the W-CDMA radio interface and the UTRAN, on the GSM core
  Release 4    2001  the split MSC: server and media gateway
  Release 5    2002  HSDPA, and the IP multimedia subsystem
  Release 6    2004  HSUPA, completing high speed packet access
  Release 7    2007  HSPA+ : higher order modulation and MIMO
  Release 8    2008  LTE and the evolved packet core: all IP, no circuits
  Release 10   2011  LTE-Advanced, which met IMT-Advanced
  Release 15   2018  5G New Radio, the first 5G release

The ITU's families, and what each demanded:
  IMT-2000       1999  144 kbit/s moving, 384 pedestrian, 2 Mbit/s indoors
  IMT-Advanced   2008  100 Mbit/s high mobility, 1 Gbit/s low mobility
  IMT-2020       2015  20 Gbit/s peak downlink, 1 ms latency, a million devices a square km
  IMT-2030       2023  the framework for the sixth generation
  each is a set of requirements from the ITU; the systems that meet them are built by others,
  which is why UMTS is a member of IMT-2000 rather than the same thing as it.
munotes.in751

UMTS and IMT-2000

What the rates mean. A 5 MB photograph takes 72.8 minutes over GSM circuit data, 4.7 minutes over EDGE in practice, 21 seconds at IMT-2000's indoor rate, 419 ms at IMT-Advanced's high-mobility rate and 2 ms at IMT-2020's peak. Reading down that column is the history of what people expect a phone to do: the jump from minutes to seconds is what made the mobile web possible, and the jump from seconds to milliseconds is what made video normal.

The mobility trade. 144 kbit/s in a vehicle, 384 walking, 2 Mbit/s indoors: a factor of fourteen between the fastest and the slowest user, decided not by the radio but by the geometry of cells.

The releases. Nine of them across nineteen years, and the pattern is worth noticing: the radio is replaced twice, the core once, and the service model, identities and roaming survive from GSM throughout. That is the modularity TS 23.101 asked for, working.

Distinctions

IMT-2000UMTS
IsA family of requirements from the ITUA system specified by 3GPP
DefinesRates, roaming, services a 3G system must offerA radio interface, an access network, a core
MembersUMTS (W-CDMA), cdma2000, TD-SCDMA and othersNot applicable
RelationshipThe standard to be metOne way of meeting it
GSMUMTS (Release 99)
Carrier200 kHz5 MHz
AccessTDMA and FDMAW-CDMA
Cluster size3, 4 or 71: every cell uses every carrier
HandoverHardSoft and softer, plus hard
CoreMSC, VLR, HLR, SGSN, GGSNThe same, evolved
Subscriber moduleSIMUSIM, with mutual authentication
FamilyYearRequires
IMT-20001999144 kbit/s vehicular, 384 pedestrian, 2 Mbit/s indoor
IMT-Advanced2008100 Mbit/s high mobility, 1 Gbit/s low
IMT-2020201520 Gbit/s peak, 1 ms latency, a million devices a square km
IMT-20302023The framework for the sixth generation

What it does not mean

IMT-2000 is not a synonym for UMTS. It is the family; UMTS is one member.

UMTS is not a replacement for the GSM core. Release 99 kept it; the core was replaced only with LTE's evolved packet core.

2 Mbit/s was not a promise to every user. It was the requirement for a nearly stationary indoor user, and the cell edge got a fraction of it.

A release is not a generation. 3GPP's releases run continuously; the generations are the ITU's labels on the results.

Meeting IMT-Advanced is not what LTE did. LTE Release 8 did not meet it; LTE-Advanced, Release 10, did, although the name "4G" was used for both.

munotes.in752

UMTS and IMT-2000

Quick revision

  • IMT-2000: the ITU's family of requirements for 3G. Rates: 144 kbit/s vehicular, 384 pedestrian, 2 Mbit/s indoor; plus roaming, packet and circuit services, quality of service. Members: UMTS (W-CDMA), cdma2000, TD-SCDMA.
  • UMTS: 3GPP's system. Release 99 keeps the GSM and GPRS core and replaces the radio with the UTRAN and W-CDMA in 5 MHz carriers; SIM becomes USIM.
  • Domains (TS 23.101): user equipment (mobile equipment + USIM) and infrastructure (access network + core network), joined at Uu; the design is modular because applications cannot be predicted.
  • Releases: 99 UMTS; 4 split MSC; 5 HSDPA and IMS; 6 HSUPA; 7 HSPA+; 8 LTE and the evolved packet core; 10 LTE-Advanced; 15 5G New Radio.
  • After: IMT-Advanced (100 Mbit/s and 1 Gbit/s), IMT-2020 (20 Gbit/s, 1 ms, a million devices a square km), IMT-2030.
  • Program: a 5 MB photograph takes 72.8 minutes on GSM data, 21 s at IMT-2000's indoor rate, 2 ms at IMT-2020's peak.

Test yourself

1. What is IMT-2000, and how does it relate to UMTS? IMT-2000 is the International Telecommunication Union's framework for third generation mobile systems: a family of radio interfaces that meet a common set of requirements, together with spectrum allocated for them, rather than a single standard. UMTS is the system specified by 3GPP that became the European and Japanese member of that family, using W-CDMA as its radio interface; cdma2000 and TD-SCDMA are other members. A system is therefore said to be an IMT-2000 system if it meets the requirements, and UMTS is one way of meeting them.

2. What data rates did IMT-2000 require, and why do they differ? At least 144 kbit/s for a user in a fast-moving vehicle, 384 kbit/s for a pedestrian and 2 Mbit/s for a nearly stationary user indoors. They differ because mobility and rate trade against each other: a fast user must be served by a large cell to avoid constant handovers, a large cell has a distant edge and shares its capacity among many users, so each can be given less; a stationary indoor user can be served by a small cell with a short path and few users, and can be given much more.

3. What did UMTS keep from GSM and what did it replace? It kept the core network: in Release 99 the MSC, VLR, HLR and the GPRS nodes SGSN and GGSN, together with the service model, the subscriber identities and the roaming machinery. It replaced the radio access network: W-CDMA in 5 MHz carriers instead of 200 kHz carriers with time slots, the UTRAN with its Node Bs and radio network controllers instead of the base station subsystem, a cluster size of one instead of three to seven, soft handover, and the USIM with mutual authentication in place of the SIM.

munotes.in753

UMTS and IMT-2000

4. Describe the domains of a UMTS network. The user equipment domain contains the mobile equipment domain, the terminal itself, and the USIM domain, which holds the subscription. The infrastructure domain contains the access network domain, which holds everything specific to the radio access technology, and the core network domain, which does not depend on it. The reference point between the user equipment and the infrastructure is Uu, the radio interface. The separation is deliberate and modular, because the standard's own introduction notes that the applications cannot be predicted and the network must be flexible enough to evolve and have a long technical lifetime.

5. Trace the 3GPP releases from UMTS to 5G. Release 99 introduced UMTS with W-CDMA and the UTRAN on the GSM and GPRS core. Release 4 split the MSC into a server and a media gateway. Release 5 added HSDPA and the IP multimedia subsystem, Release 6 added HSUPA, completing high speed packet access, and Release 7 added HSPA+ with higher order modulation and MIMO. Release 8 introduced LTE and the all-IP evolved packet core, Release 10 introduced LTE-Advanced, which met the ITU's IMT-Advanced requirements, and Release 15 introduced 5G New Radio, addressing IMT-2020.

6. What has the ITU defined since IMT-2000? IMT-Advanced in 2008, requiring 100 Mbit/s in high mobility and 1 Gbit/s in low mobility, met by LTE-Advanced; IMT-2020 in 2015, requiring a 20 Gbit/s peak downlink rate, one millisecond user-plane latency and support for a million devices per square kilometre, and describing three usage scenarios, enhanced mobile broadband, ultra-reliable low-latency communication and massive machine-type communication; and IMT-2030 in 2023, the framework for the sixth generation.

Contents This chapter on its own page

munotes.in754

Chapter Ninety-Nine

The UMTS System Architecture: UTRAN and the Core Network

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: UMTS system architecture"

In one line

UMTS keeps GSM's core and replaces its radio side: base stations become Node Bs, base station controllers become radio network controllers with far more to do, and one new interface between controllers, Iur, lets two subsystems serve one mobile at once, which is what makes soft handover possible without telling the core anything.

In the wording a student can write in an examination: a UMTS network is the UE, the UTRAN and the core network. The UE is the mobile equipment plus the USIM. The UTRAN consists of radio network subsystems (RNS); "A Radio Network Subsystem contains one RNC and is responsible for the resources" of its cells. A Node B is a "logical node in the RNS responsible for radio transmission / reception in one or more cells to/from the UE", and "The logical node terminates the Iub interface towards the RNC". A Radio Network Controller (RNC) is a "logical node in the RNS in charge of controlling the use and the integrity of the radio resources".

An RNC takes roles. The Controlling RNC has "the overall control of the logical resources of its node B's", and "There is only one Controlling RNC for any Node B". For a particular UE, one RNS is the Serving RNS, which "is in charge of the radio connection between a UE and the UTRAN" and "terminates the Iu for this UE", while another that lends cells is the Drift RNS, which "supports the Serving RNS with radio resources when the connection between the UTRAN and the UE need to use cell(s) controlled by this RNS".

The interfaces are Uu (UE to Node B, the W-CDMA radio interface), Iub (Node B to RNC), Iur (RNC to RNC) and Iu (RNC to core), the last split into Iu-CS toward the MSC and Iu-PS toward the SGSN. The core network keeps GSM's two domains: circuit switched (MSC, VLR, GMSC) and packet switched (SGSN, GGSN), over a common HLR, later the HSS.

The picture

A diagram in two dashed panels. On the left a UE, mobile equipment plus USIM, connects over Uu to two Node Bs, each in its own radio network subsystem with its own RNC, joined to it over Iub. The two RNCs are joined to each other over Iur. The upper RNC connects over Iu-CS to an MSC and VLR, which connects to a GMSC toward the PSTN; the lower RNC connects over Iu-PS to an SGSN, which connects to a GGSN toward the Internet. Notes say that one RNS is Serving for a given UE and terminates its Iu while another may be Drift, lending cells over Iur, and that the core keeps both domains

Figure 99.1 The UMTS architecture, with the interfaces of TS 25.401

Node B and RNC

The Node B is the base station, and the name is deliberately bland: during standardisation it was a placeholder that stuck. It serves "one or more cells", performs the W-CDMA transmission and reception, measures what it receives, and runs the fast inner loop of power control, which must react in under a millisecond and therefore cannot live further away.

The RNC is the BSC's successor with a much larger job, because W-CDMA's resources are not channels but interference budget and codes ([The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover]):

  • It controls the use and the integrity of the radio resources, in the standard's words.
  • It performs admission control: whether one more call can be accepted depends on how much interference it will add for everyone else, which is a judgement no BSC ever had to make.
  • It allocates codes, which are the scarce commodity.
  • It combines and splits the streams of a soft handover, because the same data arrives from and goes to several Node Bs at once and somebody must join them up.
  • It performs ciphering, which in GSM lived in the BTS: moving it to the RNC means the air interface and the Iub are both protected.
munotes.in755

The UMTS System Architecture: UTRAN and the Core Network

A radio network subsystem is one RNC and the Node Bs it controls, the analogue of a BSS.

Iur: the interface GSM never had

In GSM, two BSCs never talk. If a mobile moves between them, the MSC switches the call, and the two BSCs know nothing of each other.

W-CDMA cannot work that way, because of soft handover: a UE may be served by cells of two different RNCs at the same time, its transmissions being combined and its data duplicated. Somebody must hold the connection together, and that somebody is the Serving RNS. The other RNS, the Drift RNS, lends its cells and passes the data through over Iur.

Two consequences follow.

  • The core network sees one RNC, whatever is happening in the radio network. It talks to the Serving RNS over Iu and is told nothing about drift.
  • A change of serving RNS, SRNS relocation, is a rarer and heavier procedure, performed when the drift arrangement has become inefficient, and only then does the core participate.

The program sizes the difference: with Iur, a fraction of the boundary crossings become core procedures; without it, every one would.

The core network's two domains

The core is GSM's, and the split matters.

  • The circuit switched domain: MSC and VLR, with a GMSC toward the PSTN, exactly as in [The GSM System Architecture]. From Release 4 the MSC splits into a server for control and a media gateway for traffic.
  • The packet switched domain: SGSN and GGSN, exactly as in [New Data Services: HSCSD, GPRS and EDGE].
  • Shared: the HLR, becoming the HSS, with the AuC and the EIR.

A single UE can have a voice call in the circuit domain and a web session in the packet domain at the same time, over one Uu interface and one RNC, and the same USIM authenticates for both. That dual life is why the two Iu interfaces exist.

The architecture, computed

The program sets each UMTS entity beside its GSM counterpart, lists the interfaces, quantifies what Iur saves, describes a user in both domains at once, and lists what the RNC does that a BSC did not.

munotes.in756

The UMTS System Architecture: UTRAN and the Core Network

# The UMTS architecture at work: who does what, and what the Iur interface,
# which GSM had no equivalent of, actually buys.
print("The entities, and the GSM entity each replaces or resembles:")
for umts, gsm, does in (("UE (ME + USIM)", "MS (ME + SIM)", "the terminal and its subscription"),
                        ("Node B", "BTS", "radio transmission and reception in one or more cells"),
                        ("RNC", "BSC", "controls the use and the integrity of the radio resources"),
                        ("RNS", "BSS", "one RNC and the Node Bs it controls"),
                        ("MSC / VLR", "MSC / VLR", "the circuit switched domain"),
                        ("SGSN / GGSN", "SGSN / GGSN", "the packet switched domain"),
                        ("HLR (HSS)", "HLR", "the subscriber record")):
    print("  %-16s %-14s %s" % (umts, gsm, does))

print("\nThe interfaces:")
for name, between, what in (("Uu", "UE and Node B", "the W-CDMA radio interface"),
                            ("Iub", "Node B and RNC", "inside one radio network subsystem"),
                            ("Iur", "RNC and RNC", "between subsystems: no GSM equivalent"),
                            ("Iu-CS", "RNC and MSC", "to the circuit domain"),
                            ("Iu-PS", "RNC and SGSN", "to the packet domain")):
    print("  %-6s %-16s %s" % (name, between, what))

# 2. The serving and drift roles, and what Iur saves. A UE in soft handover
#    uses cells of two RNSs; without Iur, every such move would have to go
#    through the core.
print("\nA UE whose cells belong to two subsystems:")
print("  the Serving RNS 'is in charge of the radio connection' and 'terminates the Iu for this UE'")
print("  the Drift RNS 'supports the Serving RNS with radio resources'")
print("  so the core network still talks to one RNC only, however many cells serve the UE.")

moves, boundary = 40, 0.25
print("\n  a user making %d cell changes in a call, %.0f%% of them across an RNS boundary:" % (moves, 100 * boundary))
print("    with Iur:    %2d changes reach the core (only when the Serving RNS is relocated)"
      % int(moves * boundary * 0.2))
print("    without Iur: %2d changes reach the core (every boundary crossing becomes a core procedure)"
      % int(moves * boundary))

# 3. What the split of the core into two domains means for one user.
print("\nOne user, two domains at once:")
for domain, carries, nodes in (("circuit switched", "a voice call", "MSC server, media gateway, GMSC"),
                               ("packet switched", "a web session", "SGSN, GGSN")):
    print("  %-18s %-16s through %s" % (domain, carries, nodes))
print("  both reach the same RNC over Iu-CS and Iu-PS, and the same USIM authenticates for both.")

# 4. Why the RNC is more than a BSC: the functions that moved into it.
print("\nWhat the RNC does that a BSC did not:")
for f in ("combines and splits the streams of soft handover, which needs all the cells' data in one place",
          "runs fast outer-loop power control, hundreds of times a second",
          "admits or refuses calls by their effect on the interference floor (admission control)",
          "allocates codes, which are the scarce resource in W-CDMA",
          "encrypts: ciphering moved from the base station to the controller"):
    print("  - %s" % f)
munotes.in757

The UMTS System Architecture: UTRAN and the Core Network

The entities, and the GSM entity each replaces or resembles:
  UE (ME + USIM)   MS (ME + SIM)  the terminal and its subscription
  Node B           BTS            radio transmission and reception in one or more cells
  RNC              BSC            controls the use and the integrity of the radio resources
  RNS              BSS            one RNC and the Node Bs it controls
  MSC / VLR        MSC / VLR      the circuit switched domain
  SGSN / GGSN      SGSN / GGSN    the packet switched domain
  HLR (HSS)        HLR            the subscriber record

The interfaces:
  Uu     UE and Node B    the W-CDMA radio interface
  Iub    Node B and RNC   inside one radio network subsystem
  Iur    RNC and RNC      between subsystems: no GSM equivalent
  Iu-CS  RNC and MSC      to the circuit domain
  Iu-PS  RNC and SGSN     to the packet domain

A UE whose cells belong to two subsystems:
  the Serving RNS 'is in charge of the radio connection' and 'terminates the Iu for this UE'
  the Drift RNS 'supports the Serving RNS with radio resources'
  so the core network still talks to one RNC only, however many cells serve the UE.

  a user making 40 cell changes in a call, 25% of them across an RNS boundary:
    with Iur:     2 changes reach the core (only when the Serving RNS is relocated)
    without Iur: 10 changes reach the core (every boundary crossing becomes a core procedure)

One user, two domains at once:
  circuit switched   a voice call     through MSC server, media gateway, GMSC
  packet switched    a web session    through SGSN, GGSN
  both reach the same RNC over Iu-CS and Iu-PS, and the same USIM authenticates for both.

What the RNC does that a BSC did not:
  - combines and splits the streams of soft handover, which needs all the cells' data in one place
  - runs fast outer-loop power control, hundreds of times a second
  - admits or refuses calls by their effect on the interference floor (admission control)
  - allocates codes, which are the scarce resource in W-CDMA
  - encrypts: ciphering moved from the base station to the controller

The correspondence. UE for MS, Node B for BTS, RNC for BSC, RNS for BSS, and then the core unchanged: MSC and VLR, SGSN and GGSN, HLR. Anyone who has learned GSM's architecture has learned three quarters of UMTS's, which was the intention.

munotes.in758

The UMTS System Architecture: UTRAN and the Core Network

The interfaces. Uu, Iub, Iu-CS and Iu-PS all have GSM equivalents (Um, Abis, A, Gb). Iur has none, and that single addition is the architectural consequence of soft handover.

What Iur saves. In the model, a user making 40 cell changes during a call, a quarter of them across an RNS boundary, causes about 2 core procedures with Iur and 10 without: the drift arrangement absorbs most boundary crossings inside the UTRAN. The real saving is not the count but the kind: a core procedure involves the MSC or SGSN, signalling, and a risk to the call; a drift arrangement is local.

Both domains at once. A voice call through the MSC and a web session through the SGSN, over one radio connection, is ordinary in UMTS and impossible in GSM without two radio resources.

The RNC's new work. Five functions, and all five exist because W-CDMA shares one wideband carrier among everyone: combining soft handover streams, fast outer-loop power control, admission control against the interference floor, code allocation, and ciphering.

Distinctions

UMTSGSM counterpartDifference
UE (ME + USIM)MS (ME + SIM)Mutual authentication, larger applications
Node BBTSServes one or more cells; runs fast inner-loop power control
RNCBSCAlso: admission control, code allocation, soft handover combining, ciphering
RNSBSSOne RNC and its Node Bs
Iur(none)Lets two RNCs serve one UE
Iu-CS, Iu-PSA, GbThe same split of domains
Controlling RNCServing RNSDrift RNS
Defined forA set of Node BsOne UE's connectionOne UE's connection
HoldsOverall control of its Node Bs' logical resourcesThe radio connection; terminates Iu for that UECells lent to the Serving RNS
NumberOne per Node BOne per connected UEZero or more
Circuit domainPacket domain
NodesMSC, VLR, GMSCSGSN, GGSN
Interface from the RNCIu-CSIu-PS
CarriesCallsPacket sessions
Shared with itHLR/HSS, AuC, EIRThe same

What it does not mean

A Node B is not just a renamed BTS. It runs the fast power control loop, which GSM had no equivalent of.

The RNC is not a BSC with a new name either. Admission control, code allocation and soft handover combining are new kinds of work.

Iur is not a backup path. It is the normal way two subsystems cooperate while a UE is in soft handover.

The Serving RNS is not always the nearest. It is the one that terminates Iu for that UE, which may be drifting behind the UE's movement until a relocation.

UMTS did not abolish the circuit domain. Release 99 kept it; that happened only with LTE.

Quick revision

  • UE (ME + USIM), UTRAN, core network.
  • Node B: "logical node in the RNS responsible for radio transmission / reception in one or more cells to/from the UE", terminating Iub. RNC: "in charge of controlling the use and the integrity of the radio resources". RNS = one RNC + its Node Bs.
  • Roles: Controlling RNC, one per Node B, with "overall control of the logical resources"; Serving RNS, one per connected UE, which "terminates the Iu for this UE"; Drift RNS, which lends cells.
  • Interfaces: Uu, Iub, Iur (new: RNC to RNC), Iu-CS, Iu-PS.
  • Core: circuit (MSC, VLR, GMSC) and packet (SGSN, GGSN) domains over a common HLR/HSS, AuC, EIR.
  • The RNC's new work: soft handover combining, outer-loop power control, admission control, code allocation, ciphering.
munotes.in759

The UMTS System Architecture: UTRAN and the Core Network

Test yourself

1. Draw the UMTS architecture and name the interfaces. The user equipment, mobile equipment plus USIM, connects over the Uu radio interface to Node Bs. Each Node B connects over Iub to a radio network controller; an RNC together with its Node Bs forms a radio network subsystem, and two RNCs are connected to each other over Iur. Each RNC connects to the core network over Iu, which is Iu-CS toward the MSC and VLR of the circuit switched domain and Iu-PS toward the SGSN of the packet switched domain; the MSC reaches other networks through a GMSC and the SGSN reaches packet networks through a GGSN, with an HLR, an authentication centre and an equipment identity register shared.

2. Define the Node B, the RNC and the RNS as the standard does. A Node B is a logical node in the radio network subsystem responsible for radio transmission and reception in one or more cells to and from the user equipment, and it terminates the Iub interface towards the RNC. A radio network controller is a logical node in the radio network subsystem in charge of controlling the use and the integrity of the radio resources. A radio network subsystem contains one RNC and is responsible for the resources of its cells, and may be a whole UTRAN or part of one.

3. What are the serving and drift roles, and why do they exist? For each UE connected to the UTRAN, one radio network subsystem is the Serving RNS: it is in charge of the radio connection between the UE and the UTRAN and terminates the Iu interface for that UE. If the connection needs cells controlled by another subsystem, that subsystem acts as a Drift RNS, supporting the Serving RNS with radio resources. The roles exist because in W-CDMA a UE may be served by cells of two different RNCs at once during soft handover, and somebody must hold the connection together and present a single point of contact to the core network.

munotes.in760

The UMTS System Architecture: UTRAN and the Core Network

4. Why does UMTS need the Iur interface when GSM has nothing like it? Because soft handover lets a UE be connected to cells belonging to different RNCs at the same time, with its uplink transmissions combined and its downlink data duplicated. The Serving RNS must therefore exchange user data and control information directly with the drift RNS, which is what Iur carries. In GSM two BSCs never need to cooperate, because a mobile is in exactly one cell and the MSC switches the connection when it moves, so no equivalent interface exists.

5. What does the RNC do that a BSC did not? It combines and splits the data streams of soft handover, since the same connection passes through several Node Bs; it runs the outer loop of fast power control; it performs admission control, deciding whether an additional connection can be accepted given the interference it will add for everyone else; it allocates the spreading and scrambling codes, which are W-CDMA's scarce resource; and it performs ciphering, which in GSM was done in the base transceiver station.

6. How does the UMTS core network differ from GSM's? In Release 99 it barely differs: it keeps both domains, the circuit switched domain with its MSC, VLR and gateway MSC and the packet switched domain with its SGSN and GGSN, over a shared HLR, authentication centre and equipment identity register. The differences are that the RNC connects to the two domains over Iu-CS and Iu-PS rather than A and Gb, that a UE routinely uses both domains at once over a single radio connection, and that from Release 4 the MSC is split into a server for control and a media gateway for traffic.

Contents This chapter on its own page

munotes.in761

Chapter One Hundred

The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover

Syllabus topic Module 2, "Medium Access Control and Telecommunication Systems: UMTS radio interface"

In one line

W-CDMA gives every cell the whole 5 MHz carrier and separates everything by codes: users inside a cell by orthogonal codes from a tree, cells from each other by scrambling codes, and since every user is interference to every other, the network must command each mobile's power fifteen hundred times a second, while a mobile near a boundary simply talks to both cells at once.

In the wording a student can write in an examination: W-CDMA (wideband code division multiple access) uses a carrier of 5 MHz at a chip rate of 3.84 Mcps. Data is spread by channelisation codes, which are orthogonal variable spreading factor (OVSF) codes: "The channelisation codes ... are Orthogonal Variable Spreading Factor (OVSF) codes that preserve the orthogonality between a user's different physical channels", defined by a code tree. A spreading factor of SF means SF chips per symbol, so the symbol rate is 3.84 Mcps divided by SF: SF 4 gives 960 ksymbol/s, SF 128 gives 30, and the higher the rate a user needs the shorter the code it takes. Because the tree's codes are orthogonal only among those with no ancestor-descendant relation, allocating a short code blocks the whole subtree beneath it, which makes code allocation a real resource problem for the RNC.

Scrambling codes are applied on top: they do not spread further but separate cells in the downlink and mobiles in the uplink, because uplink transmissions cannot be kept orthogonal when they arrive at different times. Since all users share the band, each is interference to the rest, so fast closed-loop power control runs at 1500 Hz in steps of about 1 dB, answering the near-far problem: without it a mobile close to the base station drowns a distant one. The cell cluster size is 1, every cell using every carrier, so a mobile at a boundary hears two cells well and can be served by both: soft handover (cells of different Node Bs) and softer handover (sectors of one Node B), giving macro-diversity, combined by the RNC in the uplink and by the rake receiver in the mobile. The consequence is that capacity, coverage and quality are one resource: a loaded cell breathes inward, as [Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing] computed.

The carrier and the chip rate

A UMTS carrier is 5 MHz wide and carries 3.84 million chips a second. Everything else follows from those two numbers.

The bandwidth is the reason the system is called wideband CDMA: 5 MHz is wide enough that multipath echoes separated by more than one chip, 260 nanoseconds, or 78 metres of extra path, can be resolved and used ([Multipath, Fading and the Doppler Effect]).

munotes.in762

The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover

The chip rate is the budget. A user's symbol rate is the chip rate divided by its spreading factor, so a whole cell's users must share 3.84 Mcps between them, and the more a user takes, the fewer chips are left. That is a very different arithmetic from GSM's fixed slots.

The OVSF code tree

Channelisation codes come from a tree. The root is the single chip 1; each code of length n splits into two of length 2n, one being the parent repeated and the other the parent followed by its inverse. All the codes at one level are mutually orthogonal: correlate any two and the result is zero, which the program checks for every pair at four spreading factors.

The tree has one rule that matters more than any other: a code is orthogonal to another only if neither is an ancestor of the other. Using a short code therefore blocks every longer code descended from it, and the program shows the non-zero correlation that forces the rule. Allocating codes is consequently a packing problem: a user asking for a high rate takes a short code and removes a whole subtree from the pool, which is why a cell can run out of codes before it runs out of power, and why the RNC's allocator matters.

Scrambling codes

Orthogonality survives only if the codes arrive aligned. In the downlink all transmissions leave one Node B together, so they stay orthogonal, and a scrambling code per cell is added on top to tell cells apart. In the uplink transmissions arrive from different distances at different times, so orthogonality is lost; each mobile is given its own scrambling code instead, and the receiver separates them by correlation, accepting the residual interference.

This is the practical meaning of a cluster size of one: neighbouring cells use the same frequency and are distinguished by code, so there is no frequency plan, and a mobile can hear several cells at once, which is exactly what soft handover needs.

The near-far problem and power control

If every user is interference to every other, the loudest user dominates. The program computes the scale: a mobile at 100 m arrives 35 dB stronger than one at 1 km with a path-loss exponent of 3.5, and 56 dB if the distances are 50 m and 2 km. Spreading gain at SF 128 is only 21 dB, so the far mobile is simply not heard.

W-CDMA's answer is fast closed-loop power control: the Node B measures the received quality and sends a power control command 1500 times a second, and the mobile adjusts in steps of about 1 dB. An outer loop in the RNC adjusts the target that the inner loop aims at, according to the error rate actually achieved.

munotes.in763

The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover

Power control is therefore not an optimisation but a precondition: without it the uplink does not work at all. It is also what makes cell breathing inevitable, since the total received power at the Node B is the sum of everyone's, and the more users there are the higher every user must shout.

Soft and softer handover

Because every cell uses the same carrier, a mobile near a boundary can be received by two Node Bs at once, and can receive from both. That is soft handover, and it is make-before-break:

  • Uplink: both Node Bs receive the transmission and send their versions to the RNC, which selects or combines the better frames. That is why the Iur interface of [The UMTS System Architecture: UTRAN and the Core Network] exists.
  • Downlink: both transmit the same data with different codes, and the mobile's rake receiver combines them.
  • Softer handover is the same between two sectors of one Node B, where the combining happens in the Node B itself.

The gain is macro-diversity: the two paths fade independently, so both failing at once is much less likely than one failing, which the program measures. The cost is that a mobile in soft handover consumes resources in two cells, and typically 20 to 40 per cent of mobiles are in that state, which is a real capacity charge.

The rake receiver

A wideband signal resolves multipath: echoes arriving more than a chip apart look like separate, decodable copies. A rake receiver has several "fingers", each correlating with the code at a different delay, and combines their outputs. Multipath, the enemy of every other system in this book, becomes diversity here, which is the deepest reason for choosing a wide band.

The radio interface, computed

The program lists the spreading factors and the rates they leave; generates the OVSF codes and checks every pair for orthogonality; shows the ancestor rule breaking it; computes the near-far ratio against the spreading gain; and measures soft handover's diversity.

# W-CDMA: the OVSF code tree, proved orthogonal; what a spreading factor buys;
# the near-far problem power control exists for; and soft handover's gain.
import math
import random

CHIP_RATE = 3.84e6                                  # TS 25.213: "The modulating chip rate is 3.84 Mcps."

# 1. The OVSF code tree. Each code splits into two children, one the parent
#    repeated and one the parent followed by its inverse.
def ovsf(sf):
    codes = [[1]]
    while len(codes[0]) < sf:
        codes = [c + c for c in codes] + [c + [-x for x in c] for c in codes]
        codes = [c for pair in zip(codes[:len(codes) // 2], codes[len(codes) // 2:]) for c in pair]
    return codes

print("OVSF codes (TS 25.213, 4.3.1.1) and the rate each spreading factor leaves at %.2f Mcps:" % (CHIP_RATE / 1e6))
print("  SF   codes   rate a code carries   what it suits")
for sf, what in ((4, "high rate packet data"), (8, "384 kbit/s data"), (16, "a fast data channel"),
                 (32, "a slower data channel"), (128, "a speech call"), (256, "signalling")):
    print("  %3d %7d %18.1f kb/s   %s" % (sf, sf, CHIP_RATE / sf / 1000, what))

# Orthogonality: any two codes of the same length correlate to zero.
for sf in (4, 8, 16, 32):
    codes = ovsf(sf)
    worst = max(abs(sum(a * b for a, b in zip(x, y)))
                for i, x in enumerate(codes) for y in codes[i + 1:])
    print("  SF %3d: the largest correlation between any two of the %d codes is %d" % (sf, sf, worst))

# The tree's rule: using a code blocks its ancestors and its descendants.
codes4, codes8 = ovsf(4), ovsf(8)
parent = codes4[1]
children = [c for c in codes8 if c[:4] == parent or c[:4] == [-x for x in parent]]
print("\n  a code of SF 4 and one of its SF 8 descendants correlate to %d, not 0:"
      % abs(sum(a * b for a, b in zip(parent + parent, children[0]))))
print("  so taking a short code blocks every longer code below it: that is why code allocation is")
print("  the RNC's problem, and why a high-rate user costs more than its share of the tree.")

# 2. The near-far problem. Two mobiles, one at 100 m and one at 1 km, both
#    transmitting at the same power, with a path-loss exponent of 3.5.
print("\nThe near-far problem, path-loss exponent 3.5:")
for near, far in ((100, 1000), (50, 2000)):
    ratio = 10 * 3.5 * math.log10(far / near)
    print("  a mobile at %4d m arrives %5.1f dB stronger than one at %4d m" % (near, ratio, far))
print("  the far mobile's signal is below the near one's by far more than the spreading gain of")
print("  %4.1f dB at SF 128, so without power control the far mobile is simply not heard."
      % (10 * math.log10(128)))
print("  W-CDMA therefore commands power 1500 times a second, in steps of 1 dB.")

# 3. Soft handover: two cells receive the same transmission, and the RNC keeps
#    the better frame. What that is worth against fading.
rnd = random.Random(100)
print("\nSoft handover: two cells listening to one mobile, 20000 frames:")
print("   fading depth   one cell fails   both fail (soft handover fails)")
for sigma in (4.0, 6.0, 8.0):
    one = both = 0
    for _ in range(20000):
        a, b = rnd.gauss(0, sigma), rnd.gauss(0, sigma)    # independent shadowing, dB
        one += a < -6
        both += a < -6 and b < -6
    print("  %10.1f dB %14.2f%% %26.2f%%" % (sigma, 100 * one / 20000, 100 * both / 20000))
print("  the gain is that two independent paths rarely fade together, which is diversity;")
print("  the cost is that both cells spend resources on one call.")
munotes.in764

The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover

OVSF codes (TS 25.213, 4.3.1.1) and the rate each spreading factor leaves at 3.84 Mcps:
  SF   codes   rate a code carries   what it suits
    4       4              960.0 kb/s   high rate packet data
    8       8              480.0 kb/s   384 kbit/s data
   16      16              240.0 kb/s   a fast data channel
   32      32              120.0 kb/s   a slower data channel
  128     128               30.0 kb/s   a speech call
  256     256               15.0 kb/s   signalling
  SF   4: the largest correlation between any two of the 4 codes is 0
  SF   8: the largest correlation between any two of the 8 codes is 0
  SF  16: the largest correlation between any two of the 16 codes is 0
  SF  32: the largest correlation between any two of the 32 codes is 0

  a code of SF 4 and one of its SF 8 descendants correlate to 8, not 0:
  so taking a short code blocks every longer code below it: that is why code allocation is
  the RNC's problem, and why a high-rate user costs more than its share of the tree.

The near-far problem, path-loss exponent 3.5:
  a mobile at  100 m arrives  35.0 dB stronger than one at 1000 m
  a mobile at   50 m arrives  56.1 dB stronger than one at 2000 m
  the far mobile's signal is below the near one's by far more than the spreading gain of
  21.1 dB at SF 128, so without power control the far mobile is simply not heard.
  W-CDMA therefore commands power 1500 times a second, in steps of 1 dB.

Soft handover: two cells listening to one mobile, 20000 frames:
   fading depth   one cell fails   both fail (soft handover fails)
         4.0 dB           6.47%                       0.43%
         6.0 dB          16.25%                       2.62%
         8.0 dB          22.63%                       4.91%
  the gain is that two independent paths rarely fade together, which is diversity;
  the cost is that both cells spend resources on one call.
munotes.in765

The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover

Codes and rates. SF 4 leaves 960 ksymbol/s, SF 128 leaves 30 ksymbol/s, enough for a speech call after coding, and SF 256 leaves 15, enough for signalling. The tree's arithmetic is exact and unforgiving: the chip rate is fixed, so rate and the number of users trade directly.

Orthogonality, proved. At spreading factors 4, 8, 16 and 32, the largest correlation between any two codes of the same length is 0. That is the property the whole downlink rests on, and it is worth seeing checked rather than asserted.

munotes.in766

The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover

And its limit. A code of SF 4 correlates with its own SF 8 descendant to 8, not 0. So the allocator may not give out a code and its descendants at the same time, and a single SF 4 code costs a quarter of the tree.

Near and far. 35 dB of imbalance against 21 dB of spreading gain at SF 128: the distant mobile loses by 14 dB and is inaudible. Fifteen hundred power control commands a second are what close that gap, and the reason the Node B, not the RNC, runs the inner loop is that a command must arrive within two thirds of a millisecond.

Soft handover. With 6 dB shadowing, one cell's link fails 16.25 per cent of the time, but both fail only 2.62 per cent of the time: the failure rate falls by a factor of six because the two paths fade independently. At 8 dB the figures are 22.63 and 4.91 per cent. That factor is what a mobile at a cell boundary gains by being served by both cells, and it is why the boundary, the worst place in a GSM network, is not the worst place in a UMTS one.

Distinctions

GSMW-CDMA
Carrier200 kHz5 MHz at 3.84 Mcps
Users separated byFrequency and time slotCode
Cluster size3, 4 or 71
Capacity limitChannels: hardInterference: soft
Power controlSlow, about twice a secondFast, 1500 Hz, about 1 dB steps
HandoverHardSoft, softer, and hard between carriers
MultipathAn enemy, answered by an equaliserResolved and combined by a rake receiver
Channelisation (OVSF) codesScrambling codes
PurposeSeparate channels of one sourceSeparate cells (downlink) or mobiles (uplink)
Effect on bandwidthSpreads: SF chips per symbolNone: same rate
OrthogonalYes, within a level and with no ancestor relationNo: low cross-correlation only
Allocated byThe RNC, from the treePlanned per cell, or per mobile
Soft handoverSofter handoverHard handover
BetweenCells of different Node BsSectors of one Node BCarriers, or to GSM
Combined atThe RNCThe Node BNot combined
BreakNoneNoneYes
CostsResources in two cells, and IurResources in two sectorsNothing extra

What it does not mean

A spreading factor is not a data rate by itself. The chip rate divided by SF is the symbol rate before coding and control overhead.

OVSF codes are not all orthogonal to one another. Only those with no ancestor-descendant relation are, which is why the tree must be managed.

munotes.in767

The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover

Power control is not about saving battery. It is about making the uplink work at all; saving battery is a by-product.

Soft handover is not free. The mobile occupies resources in two cells, and the network carries its data twice.

Cluster size 1 does not mean no interference planning. It means no frequency planning; interference is managed by power control and admission control instead.

Quick revision

  • W-CDMA: 5 MHz carrier, 3.84 Mcps. Symbol rate = chip rate / SF: SF 4 gives 960 k, 128 gives 30 k, 256 gives 15 k.
  • OVSF codes: a tree; all codes of one level are orthogonal (program: largest correlation 0), but a code and its descendant are not (program: 8), so a short code blocks its subtree.
  • Scrambling codes: separate cells downlink and mobiles uplink; no extra spreading.
  • Near-far: 100 m against 1 km is 35 dB at exponent 3.5, more than the 21 dB spreading gain at SF 128; hence fast power control at 1500 Hz, about 1 dB steps, with an outer loop in the RNC.
  • Soft handover (different Node Bs, combined in the RNC), softer (one Node B, combined there), macro-diversity: program, at 6 dB shadowing one link fails 16.25 per cent of the time and both 2.62 per cent.
  • Rake receiver: fingers at different delays turn multipath into diversity.
  • Cluster size 1; capacity, coverage and quality are one resource, so the cell breathes.

Test yourself

1. What are the basic parameters of the W-CDMA radio interface? A carrier bandwidth of 5 MHz with a modulating chip rate of 3.84 Mcps, using code division: every cell uses every carrier, so the cluster size is one. User data is spread by orthogonal variable spreading factor codes, whose spreading factor sets the symbol rate as the chip rate divided by the factor, and a scrambling code is applied on top to distinguish cells in the downlink and mobiles in the uplink.

2. What are OVSF codes, and what rule governs their allocation? They are channelisation codes arranged in a tree, each code of length n producing two children of length 2n, the parent repeated and the parent followed by its inverse; they preserve orthogonality between a user's different physical channels. All codes of the same length are mutually orthogonal, so any two users at the same spreading factor do not interfere. However, a code is not orthogonal to its own ancestors or descendants, so allocating a code blocks the entire subtree beneath it and all the codes above it on its branch. A high-rate user therefore takes a short code and removes a large part of the tree, which makes code allocation a real resource management problem for the RNC.

munotes.in768

The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover

3. Why are scrambling codes needed as well as channelisation codes? Channelisation codes spread the signal and keep the channels of one transmitter orthogonal, which works in the downlink where everything leaves the Node B together. It does not work in the uplink, because transmissions from different mobiles arrive at different times and orthogonality is lost. Scrambling codes, which do not spread further, are therefore applied: one per cell in the downlink so that neighbouring cells using the same channelisation codes can be told apart, and one per mobile in the uplink so that the Node B can separate them by correlation.

4. What is the near-far problem, and how does W-CDMA solve it? Since all users share the same band and are separated only by codes, every user is interference to the others, and a mobile close to the base station can arrive so much stronger than a distant one that the distant one is buried. With a path-loss exponent of 3.5, a mobile at 100 m arrives about 35 dB stronger than one at a kilometre, while the spreading gain at spreading factor 128 is only about 21 dB, so the far mobile is lost. W-CDMA answers with fast closed-loop power control: the Node B measures quality and sends power commands 1500 times a second, in steps of about a decibel, with an outer loop in the RNC setting the quality target.

5. Explain soft and softer handover and the gain they give. In soft handover a mobile near a cell boundary communicates with two or more cells belonging to different Node Bs at the same time: in the uplink each Node B receives the transmission and sends it to the RNC, which selects or combines the frames, and in the downlink both transmit and the mobile's rake receiver combines them. Softer handover is the same between sectors of one Node B, with the combining done in the Node B. The gain is macro-diversity: the paths fade independently, so both failing together is far less likely than one failing. In the chapter's simulation with 6 dB shadowing, one link failed 16.25 per cent of the time but both failed only 2.62 per cent. The cost is that the mobile consumes resources in two cells.

6. Why is multipath an advantage in W-CDMA when it is a problem elsewhere? Because the signal is wide: at 3.84 Mcps a chip lasts about 260 nanoseconds, so echoes whose paths differ by more than about 78 metres arrive more than one chip apart and are not correlated with the code at the same delay. A rake receiver correlates at several delays at once, one finger per resolvable path, and combines the outputs, so each echo contributes energy instead of interference. In a narrowband system the same echoes would smear symbols into one another and would have to be removed by an equaliser.

Contents This chapter on its own page

munotes.in769

Chapter One Hundred One

A Web Request Over a Cellular Network

Syllabus topic Module 2, the paired practical, "Cellular Network Simulation with Client-Server Communication: Simulate a mobile network consisting of a cell tower, central office server, web server, and web browser, and analyze packet flow during a data request-response cycle"

In one line

A web page fetched on a phone crosses two different worlds: from the browser to the gateway it is a mobile network's problem, with bearers, contexts, ciphering and tunnels that keep the user's address fixed while the user moves, and from the gateway onward it is an ordinary TCP connection to a server that knows nothing about any of it.

In the wording a student can write in an examination: the practical's web browser is an application on the handset over TCP/IP; the cell tower is the Node B (or BTS) with its RNC (or BSC) behind it; the central office server is the operator's core network, whose packet side is the SGSN and the GGSN; and the web server is an ordinary host on the Internet.

The request travels: browser to the handset's IP stack (an HTTP GET inside TCP inside IP); over the radio bearer to the cell tower; over Iub or Abis to the controller; over Iu-PS or Gb to the SGSN, ciphered all the way from the handset; through a GTP tunnel to the GGSN; and out to the Internet with the address the GGSN gave the handset. The response returns to the GGSN, is tunnelled to whichever SGSN now serves the handset, and is delivered through the controller and the cell the handset is in.

Two ideas carry the whole chapter. First, the PDP context and its tunnel mean the handset keeps one IP address wherever it goes, so mobility is invisible above the GGSN. Second, the time a page takes is bytes divided by rate plus round trips times latency, and as bearers got faster the second term came to dominate, which is why later generations were designed for latency as much as for rate.

The four parties, and what they really are

The practical's names are the textbook ones; each stands for something specific.

  • The web browser is a TCP/IP application. Nothing about it is mobile: it opens connections, sends requests and renders replies.
  • The cell tower is the Node B or BTS, and behind it the RNC or BSC that controls it. Everything that makes the radio work, the power control, the scheduling, the handovers, lives here.
  • The central office server is the core network. For packets that is the SGSN, which serves the mobile, and the GGSN, which is the door to the outside; for calls it would be the MSC ([The GSM System Architecture]).
  • The web server is a host on the Internet, reached by ordinary routing.

The packet flow

The program lists the twelve steps. Three of them deserve attention.

Ciphering runs from the handset to the core. In GPRS and UMTS the packet is ciphered between the handset and the SGSN (in UMTS, the RNC), not merely over the air, which is one of the improvements over [GSM Security].

munotes.in770

A Web Request Over a Cellular Network

The tunnel between SGSN and GGSN is the mechanism that makes mobility invisible. The handset's IP address belongs to the GGSN; packets for it always arrive there, and the GGSN forwards them through a tunnel to whichever SGSN currently serves the handset. When the handset moves to another SGSN, the tunnel is re-pointed and nothing outside notices. This is the same idea as Mobile IP's home agent, built into the operator's network.

The web server sees an ordinary client. Every mobile-specific thing, the context, the bearer, the ciphering, the tunnel, stops at the GGSN. That is what makes the mobile Internet the Internet and not a separate network.

Where the time goes

A page is not one object. It is an HTML file and then dozens of images, scripts and stylesheets, fetched over a few parallel connections. The time is therefore:

total = (bytes / rate) + (round trips x latency) + setup

and which term dominates changes completely as bearers improve. The program computes it for four generations, and the answer is the most useful thing in this chapter: on GPRS the bytes take 200 seconds and the round trips 10; on HSPA the bytes take 2.2 seconds and the round trips 1.2. As the rate rises by ninety times, the balance shifts from almost all bytes to nearly half latency.

That is why the later generations pursued latency: LTE's 1 ms target and 5G's, and before them HSPA's shorter transmission time interval, are worth more to a page load than another doubling of rate. And it is why TCP's behaviour matters: a slow start ([Traditional Transport Control Protocols: TCP and UDP]) costs round trips at exactly the moment they are most expensive.

What each party remembers

A running session is a chain of state, and the practical's instruction to analyse the packet flow is best answered by naming it:

  • The handset holds a PDP context, an IP address and a radio bearer.
  • The RNC or BSC holds the radio bearer and knows which cells serve it.
  • The SGSN holds the mobile's location, its context and its ciphering state.
  • The GGSN holds the tunnel and the address it gave out.
  • The web server holds an ordinary TCP connection and knows none of the rest.

If the user moves

The practical's scenario is static; a real one is not, and the chapter's last table is the interesting case.

  • A cell change within one RNC moves the radio bearer; nothing above the RNC notices.
  • A change of RNC, with Iur available, is absorbed by the drift arrangement of [The UMTS System Architecture: UTRAN and the Core Network]; the core sees nothing.
  • A change of SGSN makes the GGSN re-point its tunnel; the IP address does not change, so TCP survives.
  • A change of GGSN would change the address and break every connection, which is why it does not happen during a session.
munotes.in771

A Web Request Over a Cellular Network

The design puts the anchor at the GGSN precisely so that everything below it can move.

The request, computed

The program names the four parties, lists the twelve hops of a request and response, computes where the time goes for four bearers, lists the state each party holds, and says what happens when the user moves.

# The practical: a browser on a handset fetches a page through a cell tower,
# the operator's core and a web server. Every packet, and where the time goes.
print("The four parties the practical names, and what each is in a real network:")
for name, real in (("web browser", "an application on the handset, over TCP/IP"),
                   ("cell tower", "the Node B or BTS, and behind it the RNC or BSC"),
                   ("central office server", "the operator's core: SGSN and GGSN (or MSC for calls)"),
                   ("web server", "a host on the Internet, reached through the GGSN")):
    print("  %-22s %s" % (name, real))

# 1. The packet flow, request and response.
print("\nOne request and one response, hop by hop:")
flow = [("browser", "handset IP stack", "an HTTP GET inside TCP inside IP"),
        ("handset IP stack", "radio link", "the packet is queued for the radio bearer"),
        ("radio link", "cell tower", "sent on the uplink, spread or in a slot, power controlled"),
        ("cell tower", "controller", "over Iub or Abis to the RNC or BSC"),
        ("controller", "SGSN", "over Iu-PS or Gb, ciphered from the handset to here"),
        ("SGSN", "GGSN", "through the operator's tunnel (GTP), which hides the mobile's movement"),
        ("GGSN", "the Internet", "the packet leaves with the address the GGSN gave the handset"),
        ("the Internet", "web server", "ordinary routing, ordinary TCP"),
        ("web server", "GGSN", "the response, addressed to the handset's address"),
        ("GGSN", "SGSN", "tunnelled to whichever SGSN now serves the handset"),
        ("SGSN", "controller", "and on to the cell the handset is in"),
        ("cell tower", "browser", "the downlink, then up the handset's stack to the page")]
for i, (a, b, what) in enumerate(flow, 1):
    print("  %2d. %-18s -> %-16s %s" % (i, a, b, what))

# 2. Where the time goes. A page of 40 objects over three generations.
print("\nFetching a page of 40 objects, 25 kB each, and where the time goes:")
print("  bearer            rate     round trip   transfer   setup+RTT cost   total")
for name, kbps, rtt_ms, setup_ms in (("GPRS", 40, 600, 2500), ("EDGE", 150, 400, 2000),
                                     ("UMTS R99", 384, 150, 1000), ("HSPA", 3600, 70, 300)):
    transfer = 40 * 25 * 8 / kbps                       # seconds
    # six objects at a time, so the round trips are paid in batches
    latency = setup_ms / 1000 + (40 / 6) * 2 * rtt_ms / 1000
    print("  %-14s %7.0f kb/s %8.0f ms %9.1f s %14.1f s %8.1f s"
          % (name, kbps, rtt_ms, transfer, latency, transfer + latency))
print("  on the slow bearers the bytes dominate; on the fast ones the round trips do, which is why")
print("  latency, not rate, is what later generations chased.")

# 3. What the network must remember for the session to work.
print("\nWhat the network is holding while the page loads:")
for who, state in (("the handset", "a PDP context, an IP address, a radio bearer"),
                   ("the RNC or BSC", "the radio bearer and which cells serve it"),
                   ("the SGSN", "the mobile's location, its context, its ciphering state"),
                   ("the GGSN", "the tunnel to the SGSN and the address given to the handset"),
                   ("the web server", "an ordinary TCP connection: it knows none of the above")):
    print("  %-16s %s" % (who, state))
print("  the web server sees a normal client. Every mobile-specific thing stops at the GGSN.")

# 4. What happens if the user moves mid-page.
print("\nIf the user moves while the page loads:")
for event, effect in (("cell change inside one RNC", "the radio bearer moves; nothing above notices"),
                      ("change of RNC, Iur available", "the drift RNS serves the cells; the core sees nothing"),
                      ("change of SGSN", "the GGSN retargets its tunnel; the IP address does not change"),
                      ("change of GGSN", "the address would change, so this does not happen mid-session")):
    print("  %-30s %s" % (event, effect))
print("  the IP address belongs to the GGSN, which is why mobility is invisible to the web server.")
munotes.in772

A Web Request Over a Cellular Network

The four parties the practical names, and what each is in a real network:
  web browser            an application on the handset, over TCP/IP
  cell tower             the Node B or BTS, and behind it the RNC or BSC
  central office server  the operator's core: SGSN and GGSN (or MSC for calls)
  web server             a host on the Internet, reached through the GGSN

One request and one response, hop by hop:
   1. browser            -> handset IP stack an HTTP GET inside TCP inside IP
   2. handset IP stack   -> radio link       the packet is queued for the radio bearer
   3. radio link         -> cell tower       sent on the uplink, spread or in a slot, power controlled
   4. cell tower         -> controller       over Iub or Abis to the RNC or BSC
   5. controller         -> SGSN             over Iu-PS or Gb, ciphered from the handset to here
   6. SGSN               -> GGSN             through the operator's tunnel (GTP), which hides the mobile's movement
   7. GGSN               -> the Internet     the packet leaves with the address the GGSN gave the handset
   8. the Internet       -> web server       ordinary routing, ordinary TCP
   9. web server         -> GGSN             the response, addressed to the handset's address
  10. GGSN               -> SGSN             tunnelled to whichever SGSN now serves the handset
  11. SGSN               -> controller       and on to the cell the handset is in
  12. cell tower         -> browser          the downlink, then up the handset's stack to the page

Fetching a page of 40 objects, 25 kB each, and where the time goes:
  bearer            rate     round trip   transfer   setup+RTT cost   total
  GPRS                40 kb/s      600 ms     200.0 s           10.5 s    210.5 s
  EDGE               150 kb/s      400 ms      53.3 s            7.3 s     60.7 s
  UMTS R99           384 kb/s      150 ms      20.8 s            3.0 s     23.8 s
  HSPA              3600 kb/s       70 ms       2.2 s            1.2 s      3.5 s
  on the slow bearers the bytes dominate; on the fast ones the round trips do, which is why
  latency, not rate, is what later generations chased.

What the network is holding while the page loads:
  the handset      a PDP context, an IP address, a radio bearer
  the RNC or BSC   the radio bearer and which cells serve it
  the SGSN         the mobile's location, its context, its ciphering state
  the GGSN         the tunnel to the SGSN and the address given to the handset
  the web server   an ordinary TCP connection: it knows none of the above
  the web server sees a normal client. Every mobile-specific thing stops at the GGSN.

If the user moves while the page loads:
  cell change inside one RNC     the radio bearer moves; nothing above notices
  change of RNC, Iur available   the drift RNS serves the cells; the core sees nothing
  change of SGSN                 the GGSN retargets its tunnel; the IP address does not change
  change of GGSN                 the address would change, so this does not happen mid-session
  the IP address belongs to the GGSN, which is why mobility is invisible to the web server.
munotes.in773

A Web Request Over a Cellular Network

The hops. Twelve, of which one crosses the air, three cross the operator's own network and the rest are ordinary Internet routing. The asymmetry is the point: the difficult, expensive, carefully engineered part is the first three hops.

The time. GPRS: 200.0 s of bytes and 10.5 s of latency, total 210.5 s, and nobody browsed the web on GPRS for pleasure. EDGE: 60.7 s. UMTS Release 99: 23.8 s. HSPA: 2.2 s of bytes and 1.2 s of latency, total 3.5 s. The ratio of latency to total rises from 5 per cent to 34 per cent, and on a modern bearer it is higher still. Rate stopped being the bottleneck; round trips became it.

munotes.in774

A Web Request Over a Cellular Network

The state. Five parties, four of which hold something about this particular mobile, and one, the web server, which holds nothing. That one asymmetry is the whole architecture of the mobile Internet.

Movement. Three of the four kinds of move are invisible above the layer that handles them, and the fourth is avoided by design.

Distinctions

The practical's nameThe real entityHolds
Web browserAn application on the handsetA TCP connection and a page
Cell towerNode B or BTS, with its RNC or BSCThe radio bearer, the cells serving it
Central office serverSGSN and GGSN (packets); MSC (calls)The context, location, ciphering, tunnel, address
Web serverAn Internet hostAn ordinary TCP connection
The mobile sideThe Internet side
BoundaryThe GGSNThe GGSN
AddressingContexts, tunnels, a fixed addressOrdinary IP routing
MobilityHandled by re-pointing tunnelsNot visible
SecurityCiphered to the SGSN or RNCWhatever the application provides
BearerBytes (40 objects of 25 kB)Round trips and setupTotal
GPRS, 40 kb/s200.0 s10.5 s210.5 s
EDGE, 150 kb/s53.3 s7.3 s60.7 s
UMTS R99, 384 kb/s20.8 s3.0 s23.8 s
HSPA, 3.6 Mb/s2.2 s1.2 s3.5 s

What it does not mean

The cell tower is not the network. It is the radio end; the decisions are in the controller and the core.

The IP address does not belong to the handset. It belongs to the GGSN, which is why it survives the handset's movement.

Faster is not the same as quicker. A page is round trips as well as bytes, and past a certain rate the round trips dominate.

The web server is not aware of the mobile. It sees an ordinary TCP client behind an ordinary address.

Ciphering is not end to end. It protects the handset to the core; beyond that it is the application's business, which is why HTTPS matters.

Quick revision

  • Browser (an app over TCP/IP), cell tower (Node B or BTS + RNC or BSC), central office (SGSN and GGSN), web server (an Internet host).
  • Request: browser, IP stack, radio bearer, cell tower, Iub/Abis, controller, Iu-PS/Gb, SGSN, GTP tunnel, GGSN, Internet, server. Response returns the same way, tunnelled to whichever SGSN now serves the handset.
  • The GGSN owns the address, so mobility is invisible above it; ciphering runs handset to SGSN or RNC.
  • Time = bytes / rate + round trips x latency + setup. Program: 210.5 s (GPRS), 60.7 (EDGE), 23.8 (UMTS R99), 3.5 (HSPA), with latency's share rising from 5 to 34 per cent.
  • State: handset (context, address, bearer), controller (bearer, cells), SGSN (location, context, ciphering), GGSN (tunnel, address), server (nothing).
  • Moving: within an RNC, invisible; between RNCs, absorbed by Iur; between SGSNs, the tunnel is re-pointed; between GGSNs, avoided.
munotes.in775

A Web Request Over a Cellular Network

Test yourself

1. Identify the four parties of the practical with their real counterparts. The web browser is an ordinary TCP/IP application on the handset. The cell tower is the Node B in UMTS or the BTS in GSM, together with the radio network controller or base station controller behind it, which own the radio bearer and take the radio decisions. The central office server is the operator's core network, which for packet data means the serving GPRS support node and the gateway GPRS support node, and for calls the mobile switching centre. The web server is a host on the Internet reached by ordinary routing.

2. Trace one request and response through the network. The browser issues an HTTP GET inside TCP inside IP; the handset's stack queues it on the radio bearer; it is transmitted on the uplink to the cell tower; the cell tower passes it over Iub or Abis to the controller; the controller passes it over Iu-PS or Gb to the SGSN, ciphered all the way from the handset; the SGSN tunnels it to the GGSN; the GGSN sends it onto the Internet with the address it gave the handset; ordinary routing carries it to the web server. The response is addressed to that same address, so it arrives at the GGSN, which tunnels it to whichever SGSN now serves the handset, and it is delivered through the controller and the serving cell to the browser.

3. Why does the handset's IP address not change as it moves? Because the address belongs to the GGSN, not to the handset or to any cell. Packets for the handset are always routed to the GGSN, which holds a tunnel to the SGSN currently serving the handset and forwards them along it. When the handset moves to a cell of another controller or to another SGSN, only the lower layers or the tunnel endpoint change; the address is untouched, so TCP connections survive and the web server never learns that anything happened.

4. Where does the time go in fetching a page, and how has that changed? The total is the bytes divided by the rate, plus the number of round trips multiplied by the latency, plus setup. On slow bearers the bytes dominate: in the chapter's model a page of forty 25 kB objects takes 200 seconds of transfer and 10.5 seconds of latency over GPRS. On fast bearers the balance reverses: over HSPA the same page takes 2.2 seconds of transfer and 1.2 seconds of latency, so a third of the time is round trips. That is why later generations set targets for latency, and why TCP's slow start matters most on the fastest bearers.

munotes.in776

A Web Request Over a Cellular Network

5. What state does each party hold during the session? The handset holds a PDP context, an IP address and a radio bearer. The controller holds the radio bearer and knows which cells serve it. The SGSN holds the mobile's location, its context and its ciphering state. The GGSN holds the tunnel to the SGSN and the address it allocated. The web server holds an ordinary TCP connection and knows nothing about any of the rest, which is exactly the property that lets the mobile Internet be the Internet.

6. What happens if the user moves while the page is loading? A change of cell within one controller moves the radio bearer and nothing above notices. A change of controller, where Iur exists, is handled inside the radio access network by the serving and drift arrangement, so the core network sees nothing. A change of SGSN causes the GGSN to re-point its tunnel, and the IP address is unchanged, so the TCP connections survive. A change of GGSN would change the address and break the connections, which is why the GGSN is kept as the anchor for the life of the session.

Contents This chapter on its own page

munotes.in777

Chapter One Hundred Two

The History of Satellite Systems

Syllabus topic Module 2, "Satellite Systems: History"

In one line

Every satellite communication system descends from a four-page article written in 1945 by a radar officer who worked out that a body 42,000 km from the earth's centre would take exactly a day to go round, and would therefore hang motionless above one spot, and the history since is the story of building that, discovering what half a second of delay costs, and going back to low orbits to avoid it.

In the wording a student can write in an examination: in October 1945 Arthur C. Clarke published "Extra-Terrestrial Relays" in Wireless World, proposing that satellites in a particular orbit could give worldwide radio coverage. He observed that "one orbit, with a radius of 42,000 km, has a period of exactly 24 hours. A body in such an orbit, if its plane coincided with that of the earth's equator, would revolve with the earth and would thus be stationary above the same spot on the planet." Three such stations, he argued, could cover the whole inhabited world.

The hardware followed: Sputnik 1 (1957), the first artificial satellite; Telstar 1 (1962), which relayed live television across the Atlantic from a low, moving orbit; Early Bird, or Intelsat I (1965), the first commercial geostationary satellite; India's INSAT series from 1983, a multipurpose system carrying telecommunications, television and meteorology together; Iridium (1998), a low earth orbit constellation with inter-satellite links designed for handheld telephony; and, from about 2019, constellations of hundreds and then thousands of small satellites for broadband.

The pattern is a swing: low orbits first because nothing else could be reached, then geostationary because one satellite covers a third of the earth and needs no tracking, then back to low orbits because a geostationary hop costs about 240 ms each way and broadband users notice.

Clarke's article, and what it got right

Clarke's problem was the one [Signal Propagation: Ranges, Path Loss and How a Signal Travels] describes: "long-distance communication is greatly hampered by the peculiarities of the ionosphere, and there are even occasions when it may be impossible." Short waves bounce off the ionosphere unreliably; VHF and above go straight through and are limited by the horizon. A relay above the horizon would fix both.

His reasoning is worth following because it is elementary and exact. As the orbit's radius grows, the period grows: "the velocity decreases, since gravity is diminishing and less centrifugal force is needed to balance it." Somewhere there is a radius at which the period equals a day, and a satellite there, over the equator, "would remain fixed in the sky of a whole hemisphere and unlike all other heavenly bodies would neither rise nor set."

The program checks his three figures against the physics, and the result is the best argument for reading the original: 42,000 km gives 23.795 hours, within 0.4 per cent of the true geostationary radius of 42,164 km; his 8 km/s for the lowest orbit is 7.78 km/s at 200 km; his "about 90 minutes" is 88.5. All three were computed with a slide rule, from published constants, by an officer in the Royal Air Force.

munotes.in778

The History of Satellite Systems

He also saw what it was for: not point-to-point links but broadcasting, "A true broadcast service, giving constant field strength at all times over the whole globe", and he described the stations as crewed, which is the one part that did not happen.

What the hardware proved, one thing at a time

  • Sputnik 1, 1957: that anything could be put in orbit at all, and that its signal could be heard on the ground. It carried no relay; it beeped.
  • Telstar 1, 1962: that television could be relayed by satellite. Its orbit was low and elliptical, so it was visible from both sides of the Atlantic for only about twenty minutes in each orbit, and the ground stations had to track it. That inconvenience is exactly the argument for Clarke's orbit.
  • Early Bird (Intelsat I), 1965: the first commercial geostationary satellite. One satellite, always in the same place, no tracking, continuous service: the business of satellite communication begins here.
  • INSAT, from 1983: India's own multipurpose geostationary system, combining telecommunications, television broadcasting and meteorology on one platform, and later joined by the GSAT series. For a country of India's size and terrain, a satellite reaches villages that no cable would, and carries the weather pictures that watch the monsoon.
  • Iridium, 1998: a constellation of low orbit satellites with inter-satellite links, so that a call could be routed through space rather than dropped to a ground station at every hop, aimed at handheld telephones anywhere on earth. It was a commercial failure at first, for reasons that had nothing to do with the engineering, and survives.
  • The new constellations, from about 2019: hundreds and then thousands of small satellites in low orbits, for broadband, made possible by cheap launch and mass-produced spacecraft.

Why the orbit keeps changing

Two quantities decide it, and the program computes both.

Delay. A geostationary satellite is 35,786 km up, so a signal takes about 119 ms to reach it and the same to come down: 239 ms for one hop, and nearly half a second before an answer can start coming back. For telephony that is at the edge of tolerable, and for anything interactive it is worse than any terrestrial path. A satellite at 780 km costs 2.6 ms each way.

munotes.in779

The History of Satellite Systems

Coverage and number. One geostationary satellite sees about a third of the earth, so three cover the inhabited world, exactly as Clarke said. A satellite at 780 km sees a small patch and moves, so dozens are needed for continuous coverage, and they must hand over to one another constantly ([Localization and Handover in Satellite Systems]).

The whole history is the trade between those two: few satellites and much delay, or many satellites and little delay. Every era picks whichever it can afford.

The history, computed

The program checks Clarke's numbers against the physics, tabulates four orbits with their periods and speeds, computes the delay each costs, and lists the milestones with what each first showed.

# Clarke's 1945 proposal, checked against the physics, and the systems that
# followed it. The orbital arithmetic is Kepler's third law with the earth's
# gravitational parameter.
import math

GM = 3.986004418e14                    # m^3/s^2, the earth's gravitational parameter
R_EARTH = 6378137.0                    # m, equatorial radius
SIDEREAL_DAY = 86164.0905              # s, one rotation of the earth against the stars

period = lambda r: 2 * math.pi * math.sqrt(r ** 3 / GM)
radius = lambda t: (GM * (t / (2 * math.pi)) ** 2) ** (1 / 3)
speed = lambda r: math.sqrt(GM / r)

print("Clarke, 1945: 'one orbit, with a radius of 42,000 km, has a period of exactly 24 hours'")
print("  checked: a radius of 42,000 km gives a period of %.3f hours" % (period(42e6) / 3600))
print("  the exact 24-hour orbit has a radius of %.0f km, an altitude of %.0f km"
      % (radius(86400) / 1000, (radius(86400) - R_EARTH) / 1000))
print("  and against the sidereal day, which is what actually matters, radius %.0f km, altitude %.0f km"
      % (radius(SIDEREAL_DAY) / 1000, (radius(SIDEREAL_DAY) - R_EARTH) / 1000))
print("  so Clarke's 42,000 km is within %.1f per cent of the right answer, in 1945, with no computer."
      % (100 * abs(42e6 - radius(SIDEREAL_DAY)) / radius(SIDEREAL_DAY)))

print("\nHis other numbers, checked:")
print("  'a velocity of 8 km/sec. applies only to the closest possible orbit': at %.0f km altitude, %.2f km/s"
      % (200, speed(R_EARTH + 200e3) / 1000))
print("  'the period of revolution would be about 90 minutes': %.1f minutes"
      % (period(R_EARTH + 200e3) / 60))

print("\nOrbits and their periods, computed:")
for name, alt_km in (("the space station's height", 400), ("a LEO constellation", 780),
                     ("a MEO navigation satellite", 20200), ("geostationary", 35786)):
    r = R_EARTH + alt_km * 1000
    p = period(r)
    unit = "%.1f minutes" % (p / 60) if p < 7200 else "%.2f hours" % (p / 3600)
    print("  %-28s %6d km  period %-14s speed %.2f km/s" % (name, alt_km, unit, speed(r) / 1000))

print("\nThe delay a signal pays, one way, at the speed of light:")
C = 299792458.0
for name, alt_km in (("LEO, overhead", 780), ("MEO, overhead", 20200), ("GEO, overhead", 35786),
                     ("GEO, near the horizon", 41679)):
    print("  %-26s %8.1f ms one way, %6.1f ms there and back" % (name, alt_km * 1000 / C * 1000,
                                                                 2 * alt_km * 1000 / C * 1000))
print("  a telephone call through a GEO satellite pays about half a second before anyone answers,")
print("  which is why voice moved to fibre and LEO constellations came back.")

print("\nThe milestones, and what each first showed:")
for year, what, first in ((1945, "Clarke's article in Wireless World", "the geostationary orbit as a relay"),
                          (1957, "Sputnik 1", "that an artificial satellite could be put up at all"),
                          (1962, "Telstar 1", "live television across the Atlantic, in a low orbit"),
                          (1965, "Early Bird (Intelsat I)", "the first commercial geostationary satellite"),
                          (1983, "INSAT-1B", "India's own multipurpose system: telecom, television, weather"),
                          (1998, "Iridium", "a LEO constellation with inter-satellite links"),
                          (2019, "the new constellations", "hundreds then thousands of small LEO satellites")):
    print("  %4d  %-34s %s" % (year, what, first))
munotes.in780

The History of Satellite Systems

Clarke, 1945: 'one orbit, with a radius of 42,000 km, has a period of exactly 24 hours'
  checked: a radius of 42,000 km gives a period of 23.795 hours
  the exact 24-hour orbit has a radius of 42241 km, an altitude of 35863 km
  and against the sidereal day, which is what actually matters, radius 42164 km, altitude 35786 km
  so Clarke's 42,000 km is within 0.4 per cent of the right answer, in 1945, with no computer.

His other numbers, checked:
  'a velocity of 8 km/sec. applies only to the closest possible orbit': at 200 km altitude, 7.78 km/s
  'the period of revolution would be about 90 minutes': 88.5 minutes

Orbits and their periods, computed:
  the space station's height      400 km  period 92.6 minutes   speed 7.67 km/s
  a LEO constellation             780 km  period 100.5 minutes  speed 7.46 km/s
  a MEO navigation satellite    20200 km  period 11.98 hours    speed 3.87 km/s
  geostationary                 35786 km  period 23.93 hours    speed 3.07 km/s

The delay a signal pays, one way, at the speed of light:
  LEO, overhead                   2.6 ms one way,    5.2 ms there and back
  MEO, overhead                  67.4 ms one way,  134.8 ms there and back
  GEO, overhead                 119.4 ms one way,  238.7 ms there and back
  GEO, near the horizon         139.0 ms one way,  278.1 ms there and back
  a telephone call through a GEO satellite pays about half a second before anyone answers,
  which is why voice moved to fibre and LEO constellations came back.

The milestones, and what each first showed:
  1945  Clarke's article in Wireless World the geostationary orbit as a relay
  1957  Sputnik 1                          that an artificial satellite could be put up at all
  1962  Telstar 1                          live television across the Atlantic, in a low orbit
  1965  Early Bird (Intelsat I)            the first commercial geostationary satellite
  1983  INSAT-1B                           India's own multipurpose system: telecom, television, weather
  1998  Iridium                            a LEO constellation with inter-satellite links
  2019  the new constellations             hundreds then thousands of small LEO satellites
munotes.in781

The History of Satellite Systems

Clarke checked. 42,000 km gives 23.795 hours; the true figure, against the sidereal day of 86,164 s rather than the solar day, is a radius of 42,164 km and an altitude of 35,786 km. He was out by 0.4 per cent, and he used the solar day, which is the small part of the error. The lowest orbit's speed and period come out at 7.78 km/s and 88.5 minutes against his 8 km/s and "about 90 minutes".

The orbits. 400 km, 92.6 minutes; 780 km, 100.5 minutes; 20,200 km, 11.98 hours, which is why navigation satellites repeat their ground track twice a day; 35,786 km, 23.93 hours, which is the sidereal day and the reason geostationary satellites stay put.

The delay. 2.6 ms from a LEO satellite overhead, 67.4 ms from MEO, 119.4 ms from GEO overhead and 139 ms from a GEO satellite near the horizon. Doubling those for a round trip gives the half second that shaped fifty years of satellite telephony, and explains both why long-distance voice moved to fibre and why the new constellations are low.

Distinctions

Telstar (1962)Early Bird (1965)
OrbitLow, elliptical, movingGeostationary
VisibilityAbout 20 minutes an orbitAlways, from a third of the earth
Ground stationsMust track itFixed dishes
ShowedThat relaying television was possibleThat a commercial service was possible
GeostationaryLow earth orbit
Altitude35,786 kmHundreds of km
One-way delayAbout 119 msA few ms
Satellites for global coverage3Dozens to thousands
Ground antennaFixedTracking, or electronically steered
HandoverNoneConstant
SuitsBroadcasting, wide coverageInteractive traffic, handheld terminals
MilestoneYearFirst showed
Clarke's article1945The geostationary orbit as a relay
Sputnik 11957That an artificial satellite was possible
Telstar 11962Live television across an ocean
Early Bird (Intelsat I)1965Commercial geostationary service
INSAT-1B1983India's multipurpose system: telecom, TV, weather
Iridium1998A LEO constellation with inter-satellite links
New constellationsfrom 2019Thousands of small satellites for broadband

What it does not mean

Clarke did not invent the orbit. The mathematics was known; he saw what it was for, and published it.

Geostationary is not stationary. The satellite moves at 3.07 km/s; it is stationary relative to the ground because the earth turns underneath at the same angular rate.

A 24-hour orbit is not quite right. The period must match the sidereal day, 86,164 s, not the solar day, which is why the altitude is 35,786 km and not 35,863.

munotes.in782

The History of Satellite Systems

Low orbits are not new. They were first because nothing else could be reached; they returned because delay matters more than it used to.

Iridium's failure was not technical. The constellation worked; the business did not, at first.

Quick revision

  • Clarke, Wireless World, October 1945: "one orbit, with a radius of 42,000 km, has a period of exactly 24 hours"; such a body "would revolve with the earth and would thus be stationary above the same spot"; three stations cover the world. Checked: 42,000 km gives 23.795 hours; the true radius is 42,164 km, altitude 35,786 km, against the sidereal day.
  • Milestones: Sputnik 1 (1957), Telstar 1 (1962, low and tracked), Early Bird / Intelsat I (1965, first commercial GEO), INSAT (from 1983, India's multipurpose GEO), Iridium (1998, LEO with inter-satellite links), new constellations (from 2019).
  • Orbits: 400 km, 92.6 min; 780 km, 100.5 min; 20,200 km, 11.98 h; 35,786 km, 23.93 h.
  • Delay one way: LEO 2.6 ms, MEO 67.4 ms, GEO 119.4 ms overhead, 139 ms near the horizon.
  • The trade: 3 satellites and half a second, or dozens and milliseconds.

Test yourself

1. What did Clarke propose in 1945, and why? He proposed that satellites, which he called rocket stations, be placed in an orbit whose period equalled the earth's rotation so that they would appear stationary above a fixed point, and used as radio relays giving worldwide coverage. His reason was that long-distance communication then depended on the ionosphere, which is unreliable and sometimes fails altogether, while frequencies that pass through it are limited by the horizon; a relay high above the horizon would remove both problems and give, in his words, a true broadcast service with constant field strength over the whole globe.

2. What orbit did he identify, and how close was he? He stated that an orbit with a radius of 42,000 km has a period of exactly 24 hours, and that a body in such an orbit over the equator would revolve with the earth and remain above the same spot. The true geostationary radius is 42,164 km, giving an altitude of 35,786 km, because the period must equal the sidereal day of 86,164 seconds rather than the solar day of 86,400. His figure is therefore within about 0.4 per cent, and his other estimates, 8 km/s and about 90 minutes for the lowest orbit, compare with computed values of 7.78 km/s and 88.5 minutes.

3. What did Telstar and Early Bird each demonstrate? Telstar 1, in 1962, showed that live television could be relayed across the Atlantic by satellite, but it was in a low elliptical orbit, visible from both sides for only about twenty minutes in each orbit, and the ground stations had to track it. Early Bird, or Intelsat I, in 1965, was the first commercial geostationary satellite: it stayed in one place in the sky, so ground stations could use fixed dishes and service was continuous, which is what made satellite communication a business rather than a demonstration.

munotes.in783

The History of Satellite Systems

4. Why did satellite systems move from low orbits to geostationary ones and then back? Early satellites were low because nothing else could be reached. Geostationary orbit was then preferred because one satellite covers about a third of the earth, three cover the inhabited world, and the ground antennas can be fixed. Its drawback is distance: 35,786 km costs about 119 ms of propagation each way, nearly half a second for a question and its answer, which is intolerable for interactive traffic. As launch became cheap enough to orbit hundreds of satellites, low orbits returned, giving a few milliseconds of delay at the cost of many satellites, constant handovers and tracking or steerable antennas.

5. What was distinctive about Iridium? It was a constellation in low earth orbit intended to serve handheld telephones anywhere on the planet, and, unlike most systems, its satellites had inter-satellite links, so a call could be routed from satellite to satellite through space rather than being dropped to a ground station at every hop. That made it independent of the availability of ground stations over oceans and remote regions. Its engineering worked; its original business did not, and the system was later bought and continued.

6. Why is INSAT significant for India? From 1983 INSAT gave India its own multipurpose geostationary system, carrying telecommunications, television broadcasting and meteorological instruments on the same platform, later extended by the GSAT series. For a country of that size, with mountains, deserts, islands and a very large rural population, a satellite reaches places no cable economically can, delivers education and television broadcasting nationally, and supplies the weather imagery on which monsoon and cyclone warnings depend.

Contents This chapter on its own page

munotes.in784

Chapter One Hundred Three

Applications of Satellite Systems

Syllabus topic Module 2, "Satellite Systems: Applications"

In one line

A satellite is worth its cost wherever geography defeats cable: it sees a third of the earth at once, so it broadcasts to a continent from one transmitter and watches a monsoon continuously; it is at a known place at a known time, so it tells a receiver where and when it is; and it needs nothing on the ground, so it reaches a ship, a desert, a mountain or a disaster where every other network has stopped.

In the wording a student can write in an examination: the main applications are:

  1. Weather forecasting and earth observation. Geostationary satellites watch a whole hemisphere continuously and image it every few minutes; polar orbiters pass twice a day from a much lower altitude with finer resolution. India's INSAT and dedicated meteorological satellites carry the imagery behind monsoon and cyclone warnings.
  2. Broadcasting. One satellite covers a third of the earth, so a single uplink reaches a continent: direct-to-home television, radio, and educational broadcasting. India's INSAT and GSAT series carry it.
  3. Military applications. Communications that do not depend on any infrastructure in the theatre, reconnaissance, early warning, and navigation.
  4. Navigation. A receiver hearing four satellites solves for three coordinates and its own clock error: GPS, GLONASS, Galileo, BeiDou, and India's regional NavIC.
  5. Global telephone backbones. The original business, carrying intercontinental calls, now largely taken by submarine fibre, which is cheaper and has a tenth of the delay; satellites remain for routes and places fibre does not serve.
  6. Remote or distributed connections. Islands, ships, aircraft, oil platforms, pipelines, remote villages, VSAT networks for banks and businesses, and disaster relief when the terrestrial network is destroyed.
  7. Mobile communication by satellite. Handheld and vehicle terminals through constellations such as Iridium and Inmarsat, and the broadband constellations in low orbit.

Coverage is the property everything else rests on

A geostationary satellite, with users required to see it at 10 degrees above the horizon or better, covers about 174 million square kilometres, which is 34 per cent of the earth. Three such satellites therefore cover the inhabited world, as Clarke said. A satellite at 780 km covers 13.4 million square kilometres, 2.6 per cent, so dozens are needed, and they move.

That single table decides which orbit each application uses. Broadcasting and weather want one satellite to see everything at once, so they are geostationary. Navigation wants many satellites visible from anywhere with good geometry, so it uses medium orbits. Handheld communication wants low delay and a signal a small antenna can reach, so it uses low orbits and many satellites.

Broadcasting

The comparison the program makes is the whole argument. India is 3.29 million square kilometres, 1.9 per cent of one geostationary footprint. Covering it with terrestrial television transmitters of 80 km radius would take about 163 of them, each with a mast, power, maintenance and a feed; covering it with mobile cells of 3 km radius would take about 116,000. One satellite, one uplink station, and every dish in the country receives the same signal.

munotes.in785

Applications of Satellite Systems

Broadcasting is where the satellite's economics are unanswerable, because the cost does not grow with the audience. It is also what Clarke foresaw: not point-to-point links but "A true broadcast service, giving constant field strength at all times over the whole globe."

For India the argument has an extra edge: the same satellite that reaches Mumbai reaches a village in Arunachal Pradesh, and the educational broadcasting that began with SITE in the 1970s and continues through GSAT-based services exists because of it.

Weather

A geostationary satellite sees the same third of the earth all the time, so it can take an image every few minutes and make a film of a developing system: this is how a cyclone is tracked hour by hour across the Bay of Bengal. A polar orbiter at 830 km sees a narrow swath and passes over a given place about twice a day, but from forty times closer, so its instruments resolve much finer detail.

The two are complementary, and every serious meteorological service uses both: the geostationary satellite watches, the polar orbiter inspects.

Navigation

A navigation receiver measures the time a signal took to arrive from each satellite. Each measurement gives a sphere of possible positions, so three would fix a point, but the receiver's own clock is not accurate enough to know when the signal left. Its clock error is therefore a fourth unknown, and a fourth satellite is needed to solve for it.

The program shows why the clocks matter: one nanosecond of timing error is 30 cm of position error, 10 ns is 3 m, and a microsecond is 300 m. That is why the satellites carry atomic clocks, why the system continuously corrects them from the ground, and why the corrections include relativity: at 20,200 km the satellite's clocks run fast by tens of microseconds a day compared with the ground, which would be kilometres of error within a day if it were ignored.

India's NavIC is a regional system, using satellites in geostationary and inclined geosynchronous orbits to serve India and the region rather than the whole globe, which is a cheaper way to get the same service where it is wanted.

The uses that changed

Telephone backbones were the original commercial use, and they have largely gone to submarine fibre, which carries far more capacity at a fraction of the delay: [The History of Satellite Systems] computed 239 ms for a geostationary hop, against about 70 ms across an ocean by cable. Satellites keep the routes fibre does not serve and act as backup when a cable is cut.

munotes.in786

Applications of Satellite Systems

Remote and distributed connections remain: VSAT networks connect thousands of small terminals, typically for banks, fuel stations and government offices, through a hub; ships and aircraft are connected only by satellite; and disaster relief depends on terminals that need nothing on the ground, which is exactly what a cyclone or an earthquake leaves.

Mobile communication by satellite is a different trade: a handheld terminal has a poor antenna and little power, so it needs a satellite that is close, which means a low orbit and a constellation.

Why a dish

The program's last table is the physical reason satellite terminals look as they do. At 12 GHz over 36,000 km, free space loss is 205.2 dB. With a strong satellite transmitter, a half-wave dipole would deliver about -123 dBm, far below any receiver's sensitivity; a 60 cm dish delivers -90.2 dBm and a 2.4 m dish -78.2 dBm.

The dish is not a convenience: it is 33 dB, a factor of two thousand, and without it the link does not exist. It is also why a handheld satellite phone cannot work through a geostationary satellite at those frequencies, and why handheld systems use lower frequencies and much lower orbits.

The applications, computed

The program computes the footprint of four orbits, compares one satellite with terrestrial transmitters for a country the size of India, shows what a timing error costs a navigation receiver, contrasts geostationary and polar weather satellites, and computes the link budget that forces a dish.

# What satellites are used for, and the numbers behind each use.
import math

GM, R_EARTH = 3.986004418e14, 6378137.0
C = 299792458.0

# 1. Coverage: what one satellite sees, by altitude. The half-angle at the
#    earth's centre for a minimum elevation angle e.
def footprint(alt_km, elev_deg=10.0):
    r = R_EARTH + alt_km * 1000
    e = math.radians(elev_deg)
    lam = math.acos(R_EARTH / r * math.cos(e)) - e            # earth central half-angle
    area = 2 * math.pi * R_EARTH ** 2 * (1 - math.cos(lam))   # spherical cap
    return math.degrees(lam), area, 100 * area / (4 * math.pi * R_EARTH ** 2)

print("What one satellite can see, with users at 10 degrees elevation or better:")
print("  altitude      half-angle      area seen        share of the earth   satellites for global cover")
for alt in (780, 1200, 20200, 35786):
    lam, area, share = footprint(alt)
    print("  %6d km %10.1f deg %13.2f million km2 %14.1f%% %20.0f"
          % (alt, lam, area / 1e12, share, math.ceil(100 / share)))
print("  (the last column ignores geometry, so a real constellation needs more; it is a floor.)")

# 2. Broadcasting: one satellite against terrestrial transmitters.
print("\nBroadcasting to a country the size of India (3.29 million km2):")
INDIA = 3.287e12
lam, area, share = footprint(35786)
print("  one geostationary satellite covers %.1f million km2, so India is %.1f%% of its footprint"
      % (area / 1e12, 100 * INDIA / area))
for name, radius_km in (("high-power TV transmitters of 80 km radius", 80),
                        ("mobile cells of 3 km radius", 3)):
    n = INDIA / (math.pi * (radius_km * 1000) ** 2)
    print("  the same area needs about %8.0f %s" % (n, name))

# 3. Navigation: why four satellites, and what a timing error costs.
print("\nNavigation: why a receiver needs four satellites")
print("  three unknowns of position, plus one unknown clock error, so four equations, so four satellites")
for err_ns in (1, 10, 100, 1000):
    print("  a clock error of %5d ns puts the position out by %7.2f m" % (err_ns, err_ns * 1e-9 * C))
print("  which is why navigation satellites carry atomic clocks and why the system corrects for")
print("  relativity: at 20,200 km the clocks run fast by tens of microseconds a day.")

# 4. Weather: what a geostationary satellite can watch, and how often.
print("\nWeather from geostationary orbit:")
print("  the satellite sees the same %.0f%% of the earth continuously, so it can image every few minutes"
      % share)
print("  a polar orbiter at 830 km passes over a given place about twice a day, but from much closer")
print("  so the two are complementary: one watches, the other inspects.")

# 5. The link budget for a small terminal: why a dish and not a whip antenna.
print("\nWhy a geostationary terminal needs a dish (downlink, 12 GHz, 36,000 km):")
lam_m = C / 12e9
fspl = 20 * math.log10(4 * math.pi * 36e6 / lam_m)
print("  free space loss over 36,000 km at 12 GHz: %.1f dB" % fspl)
for name, gain in (("a half-wave dipole", 2.15), ("a 60 cm dish", 35.0), ("a 2.4 m dish", 47.0)):
    print("  %-20s gain %5.1f dBi, so a 50 dBW satellite gives %7.1f dBm at the receiver"
          % (name, gain, 50 + 30 - fspl + gain))
print("  a dipole leaves the signal far below any receiver's sensitivity; a dish is not a luxury.")
munotes.in787

Applications of Satellite Systems

What one satellite can see, with users at 10 degrees elevation or better:
  altitude      half-angle      area seen        share of the earth   satellites for global cover
     780 km       18.7 deg         13.43 million km2            2.6%                   39
    1200 km       24.0 deg         22.13 million km2            4.3%                   24
   20200 km       66.3 deg        152.99 million km2           29.9%                    4
   35786 km       71.4 deg        174.21 million km2           34.1%                    3
  (the last column ignores geometry, so a real constellation needs more; it is a floor.)

Broadcasting to a country the size of India (3.29 million km2):
  one geostationary satellite covers 174.2 million km2, so India is 1.9% of its footprint
  the same area needs about      163 high-power TV transmitters of 80 km radius
  the same area needs about   116254 mobile cells of 3 km radius

Navigation: why a receiver needs four satellites
  three unknowns of position, plus one unknown clock error, so four equations, so four satellites
  a clock error of     1 ns puts the position out by    0.30 m
  a clock error of    10 ns puts the position out by    3.00 m
  a clock error of   100 ns puts the position out by   29.98 m
  a clock error of  1000 ns puts the position out by  299.79 m
  which is why navigation satellites carry atomic clocks and why the system corrects for
  relativity: at 20,200 km the clocks run fast by tens of microseconds a day.

Weather from geostationary orbit:
  the satellite sees the same 34% of the earth continuously, so it can image every few minutes
  a polar orbiter at 830 km passes over a given place about twice a day, but from much closer
  so the two are complementary: one watches, the other inspects.

Why a geostationary terminal needs a dish (downlink, 12 GHz, 36,000 km):
  free space loss over 36,000 km at 12 GHz: 205.2 dB
  a half-wave dipole   gain   2.1 dBi, so a 50 dBW satellite gives  -123.0 dBm at the receiver
  a 60 cm dish         gain  35.0 dBi, so a 50 dBW satellite gives   -90.2 dBm at the receiver
  a 2.4 m dish         gain  47.0 dBi, so a 50 dBW satellite gives   -78.2 dBm at the receiver
  a dipole leaves the signal far below any receiver's sensitivity; a dish is not a luxury.
munotes.in788

Applications of Satellite Systems

Footprints. 2.6 per cent of the earth from 780 km, 4.3 from 1200 km, 29.9 from 20,200 km and 34.1 from geostationary orbit, with a floor of 39, 24, 4 and 3 satellites respectively for global coverage. The floor is optimistic, since a real constellation needs overlap and good geometry, but the shape of the answer is right: coverage is bought with altitude, and delay is the price.

Broadcasting. One satellite against 163 television transmitters or 116,254 mobile cells for the same area. That ratio has not changed since Early Bird, and it is why satellite television still exists in an age of fibre.

Navigation. Four satellites because the clock is a fourth unknown; 30 cm per nanosecond; atomic clocks and relativistic corrections as a consequence.

The dish. 205.2 dB of free space loss, and 33 dB of antenna gain between a dipole and a 60 cm dish. Every satellite dish is that number made of metal.

munotes.in789

Applications of Satellite Systems

Distinctions

ApplicationOrbit usedWhy that orbit
BroadcastingGeostationaryOne satellite covers a third of the earth; fixed dishes
Weather: watchingGeostationaryContinuous view of a hemisphere
Weather: inspectingPolar, lowFine resolution, global coverage over a day
NavigationMedium (about 20,200 km)Several satellites visible everywhere with good geometry
Handheld mobileLowShort delay and a link a small antenna can close
Backbone telephonyGeostationary (historically)Wide coverage; now largely replaced by fibre
SatelliteTerrestrial
Cost with areaNearly independentGrows with the area
Cost with usersNearly independent (broadcast)Grows with the users
Delay239 ms (GEO, round trip)Milliseconds
ReachesAnywhere in the footprintWhere it is built
Best atBroadcasting, remote places, mobility over oceansDense, interactive, high-capacity traffic
Geostationary weather satellitePolar orbiter
Altitude35,786 kmAbout 830 km
SeesA third of the earth, alwaysA swath, twice a day
ResolutionCoarserFiner
Used forWatching systems developDetailed measurement

What it does not mean

A satellite is not a high tower. Its advantage is coverage and independence from ground infrastructure, not height as such, and it pays for them in delay and in link budget.

Three satellites do not cover the whole earth. They cover the inhabited world; the poles are not visible from geostationary orbit.

Navigation satellites do not know where the receiver is. They broadcast their own position and time; the receiver computes.

Satellite broadcasting is not obsolete. Its cost does not grow with the audience or the area, which no terrestrial system can match.

A VSAT is not a small version of a broadcast dish. It transmits as well as receives, which is why it is licensed and carefully pointed.

Quick revision

  • Applications: weather (GEO watches, polar inspects), broadcasting (one satellite, a continent), military, navigation (four satellites: three coordinates plus the clock), telephone backbones (largely lost to fibre), remote connections (VSAT, ships, aircraft, disaster relief), mobile by satellite (Iridium, Inmarsat, LEO broadband).
  • Footprint at 10 degrees elevation: 780 km, 2.6 per cent; 1200 km, 4.3; 20,200 km, 29.9; GEO, 34.1, so at least 39 / 24 / 4 / 3 satellites for global coverage.
  • India: 1.9 per cent of one GEO footprint; the same area needs about 163 TV transmitters or 116,254 mobile cells. INSAT, GSAT, NavIC.
  • Navigation: 1 ns = 30 cm; atomic clocks; relativistic correction of tens of microseconds a day at 20,200 km.
  • Link: 205.2 dB free space loss at 12 GHz over 36,000 km; a dipole gives -123 dBm, a 60 cm dish -90.2 dBm: the dish is 33 dB.
munotes.in790

Applications of Satellite Systems

Test yourself

1. List the main applications of satellite systems. Weather forecasting and earth observation; broadcasting of television and radio; military communications, reconnaissance and early warning; navigation; global telephone backbones; remote and distributed connections such as VSAT networks, ships, aircraft and disaster relief; and mobile communication by satellite for handheld and vehicle terminals.

2. Why is broadcasting the application satellites suit best? Because a satellite's cost does not grow with the area covered or with the number of receivers. One geostationary satellite covers about 34 per cent of the earth, so a single uplink reaches a continent; the same area covered terrestrially would need many transmitters, and in the chapter's calculation India alone, which is under 2 per cent of one footprint, would need about 163 high-power television transmitters or 116,254 mobile cells. Every additional viewer costs the satellite operator nothing.

3. Why does a navigation receiver need four satellites? Each satellite's signal gives the distance to that satellite, which would place the receiver on a sphere, so three spheres would fix a point in three dimensions. But the distance is computed from the time the signal took, and the receiver's own clock is not accurate enough to know when the signal was sent, so the clock offset is a fourth unknown. A fourth satellite provides the fourth equation, and the receiver solves for its position and its clock error together, which also gives it a very accurate time.

4. Why must navigation satellites carry atomic clocks and correct for relativity? Because position is computed from timing: one nanosecond of error is about 30 centimetres of position error, ten nanoseconds is three metres, and a microsecond is 300 metres. Ordinary clocks drift far more than that in a day. At 20,200 km the satellites' clocks also run measurably faster than clocks on the ground, by tens of microseconds a day when the effects of gravitation and of velocity are combined, which would accumulate into kilometres of error if it were not corrected.

5. How do geostationary and polar weather satellites complement each other? A geostationary satellite sees the same third of the earth continuously, so it can image every few minutes and follow a storm as it develops, but from 35,786 km its resolution is coarse and it cannot see the poles. A polar orbiter at about 830 km sees only a narrow swath at a time and passes over a given place roughly twice a day, but it is forty times closer, so its instruments resolve much finer detail and it covers the whole earth including the poles. Meteorological services use both: the geostationary satellite watches, the polar orbiter inspects.

munotes.in791

Applications of Satellite Systems

6. Why does a satellite terminal need a dish? Because of the path loss. At 12 GHz over 36,000 km the free space loss is about 205 dB, so even a powerful satellite transmitter leaves a very weak signal at the ground. In the chapter's link budget a half-wave dipole would receive about -123 dBm, far below any receiver's sensitivity, while a 60 cm dish, with about 35 dBi of gain, receives about -90 dBm and a 2.4 m dish about -78 dBm. The dish supplies more than 30 dB, a factor of over a thousand, without which the link simply does not close.

Contents This chapter on its own page

munotes.in792

Chapter One Hundred Four

Satellite Basics: Orbits, Periods, Elevation and Footprints

Syllabus topic Module 2, "Satellite Systems: Basics"

In one line

Choose the altitude and everything else follows: the period and the speed come from Kepler's third law, the footprint and the elevation angle come from a triangle, and the delay comes from the distance, so the whole design of a satellite system is the consequence of one decision.

Why it stays up

A satellite is not held up by anything. It is falling, and missing. At the right speed for its radius, the earth's gravity supplies exactly the centripetal acceleration that a circular path of that radius needs, no more and no less, so the satellite keeps turning without getting closer or further away.

Write that as an equation and the speed drops out of it. Gravity pulls with acceleration GM over r squared, where GM is the earth's gravitational parameter and r the distance from the earth's centre, not from its surface. A circle of radius r at speed v needs v squared over r. Setting them equal gives v equal to the square root of GM over r, and going round once at that speed takes

T equal to two pi times the square root of r cubed over GM,

which is Kepler's third law: the square of the period is proportional to the cube of the semi-major axis. The program prints both sides of the balance at three altitudes and they agree to four decimal places, which is the whole of orbital mechanics in one line of arithmetic.

Two consequences matter for the rest of the chapter. Higher means slower: at 780 km a satellite travels at about 7,462 metres a second and goes round in about 100 minutes; at geostationary height it travels at about 3,075 metres a second and takes about 1,436 minutes, which is 23 hours and 56 minutes. And the period depends only on the radius, not on the satellite's mass, its shape or what it carries.

The vocabulary a question expects

TermWhat it means
ApogeeThe point of the orbit furthest from the earth
PerigeeThe point closest to the earth
Semi-major axisHalf the long axis of the ellipse; for a circle, the radius
EccentricityHow far from circular the orbit is: 0 is a circle, near 1 is a long thin ellipse
InclinationThe angle between the orbital plane and the equator: 0 is equatorial, 90 is polar
Sub-satellite pointThe point on the ground directly below the satellite
Elevation angleHow high above the user's own horizon the satellite appears
FootprintThe area of the earth that can reach the satellite at or above a chosen minimum elevation
Line of sightAn unobstructed straight path between the antenna and the satellite
UplinkGround to satellite
DownlinkSatellite to ground
TransponderThe unit on board that receives an uplink channel, shifts and amplifies it, and retransmits it on the downlink
Inter-satellite linkA link from one satellite directly to another, without touching the ground
Gateway linkThe link between a satellite and a fixed earth station that joins it to the terrestrial network
Mobile user linkThe link between a satellite and a mobile or handheld terminal
munotes.in793

Satellite Basics: Orbits, Periods, Elevation and Footprints

Elevation, and why the minimum matters

A user standing under the satellite sees it at 90 degrees, straight overhead. Move away and the angle falls, at first slowly and then fast, until the satellite is on the horizon and then below it.

A system does not use the whole geometric horizon. Close to the horizon the signal travels a long way through the atmosphere, it is blocked by buildings, hills and trees, and it arrives with multipath from the ground: [Multipath, Fading and the Doppler Effect] is at its worst there. So a system specifies a minimum elevation angle, commonly 10 degrees, and the footprint is the area from which the satellite is at least that high.

That minimum is expensive. The program computes what it costs: from geostationary orbit, insisting on 10 degrees rather than 0 shrinks the footprint radius from 9,050 km to 7,952 km, and insisting on 30 degrees shrinks it to 5,841 km. A handheld terminal, which cannot tolerate obstruction at all, may want 30 degrees or more, and that is one of the reasons handheld systems need many satellites.

A diagram of the elevation geometry. A gently curved ground line runs across the picture. A satellite sits above it, with a dashed vertical line down to a marked sub-satellite point on the ground. A user stands to the right of the sub-satellite point, with a solid line drawn from the user up to the satellite labelled slant range, and a faint straight line through the user labelled the user's horizon. A small arc between the horizon and the slant line is labelled elevation. Two dashed lines run from the satellite down to two open circles far out on either side, marked footprint edge. A note says the satellite is drawn close so the angles can be seen, and another says that past the dashed edges the satellite is under ten degrees up, so it is out of the footprint

Figure 104.1 The elevation angle is measured from the user's own horizon, not from the vertical

The four orbits

A quarter of the earth drawn at the bottom left, with three arcs above it drawn to the same scale. The first arc hugs the surface and is labelled LEO, 780 km, imaging. The second is much further out and is labelled MEO, 20,200 km, navigation. The third is furthest and is labelled GEO, 35,786 km, broadcasting and weather. Two shaded bands lie between the arcs. A note says the drawing is to scale, that LEO almost grazes the surface and that GEO is 5.6 earth radii above it, and a second note says the shaded bands are the radiation belts, roughly 2,000 to 6,000 km and 15,000 to 30,000 km

Figure 104.2 The orbit bands to scale, with the radiation belts that shape the choice

GEO, the geostationary orbit, at 35,786 km above the equator, with a period of one day so that the satellite appears fixed in the sky. MEO, medium earth orbit, a few thousand to about 20,000 km, used by navigation systems. LEO, low earth orbit, a few hundred to about 1,500 km, used by imaging satellites and by communication constellations. And HEO, a highly elliptical orbit, which is not a height at all but a shape: a low perigee and a very high apogee.

The elliptical orbit is worth the program's last section. With a perigee of 1,000 km and an apogee of 39,400 km the period is 12 hours, and because a body moves slowly when it is far away, the satellite spends 8.8 hours of every 12 above 20,000 km. It therefore hangs over one hemisphere for most of its orbit and rushes through the other half in a couple of hours. Three such satellites, spaced in time, keep one of them always high over a high latitude, which is what a geostationary satellite cannot do: from a high latitude, a satellite over the equator is always low in the sky or below the horizon.

munotes.in794

Satellite Basics: Orbits, Periods, Elevation and Footprints

Why not every altitude

Two things rule out most of the space between the bands.

The first is the radiation belts, regions where the earth's magnetic field traps charged particles. They are usually given as roughly 2,000 to 6,000 km and roughly 15,000 to 30,000 km, though they have no sharp edges and different books draw them differently. A satellite that spends its life inside one needs heavier shielding and radiation-hardened electronics, so a designer either keeps below the first, aims for the gap, or builds for the radiation.

The second is delay, and it is the price of altitude. The program's last table gives the round trip: 5.2 ms through a satellite at 780 km, 134.8 ms at 20,200 km, 238.7 ms at geostationary height. A conversation over a geostationary link has a noticeable pause; a protocol that waits for an acknowledgement before sending more, as [Traditional Transport Control Protocols: TCP and UDP] does, runs at a fraction of its rate over one.

How long a satellite stays up in your sky

A geostationary satellite never sets: it turns with the earth, so a dish is aimed once and bolted down. Every other orbit moves through the sky, and the program computes how long it is usable. At 780 km a satellite passing straight overhead is above 10 degrees for only 10.4 minutes out of its 100 minute orbit; at 1,200 km, for 14.6 minutes; at 20,200 km, for 264.8 minutes.

Ten minutes is the whole reason a low constellation is hard. A call must be handed from one satellite to the next every few minutes, whether or not the user has moved at all: this is the inter-satellite handover of [Localization and Handover in Satellite Systems], and it is a handover caused by the network moving rather than the user. Many satellites must be in orbit for one to be there at all times, and the ground station must track.

The bands satellites work in

Satellite bands are named by letters that came from wartime radar and stuck. In common use: L around 1 to 2 GHz and S around 2 to 4 GHz for mobile and handheld terminals, where a small antenna can work; C with an uplink near 6 GHz and a downlink near 4 GHz, the oldest fixed-service band, robust in rain; Ku with an uplink near 14 GHz and a downlink near 11 to 12 GHz, used for direct-to-home television, where the higher frequency lets a small dish have the gain that [Antennas: Radiators, Dipoles and Radiation Patterns] computes; and Ka near 30 GHz up and 20 GHz down, which gives the most bandwidth and suffers the most from rain.

munotes.in795

Satellite Basics: Orbits, Periods, Elevation and Footprints

The pattern is the one [Frequencies for Radio Transmission] sets out: higher frequency means more bandwidth and smaller antennas, and worse weather. The uplink is always the higher of a pair, because the earth station can afford the power and the satellite cannot.

Orbits, computed

# Orbits: the period, the speed, the elevation angle and how long a satellite stays up.
import math

GM = 3.986004418e14        # earth's gravitational parameter, m3/s2
R = 6378137.0              # equatorial radius, m
C = 299792458.0

def period(alt_km):
    """Kepler's third law for a circular orbit of the given altitude."""
    r = R + alt_km * 1000
    return 2 * math.pi * math.sqrt(r ** 3 / GM)

def speed(alt_km):
    return math.sqrt(GM / (R + alt_km * 1000))

# 1. Why it stays up: at the right speed, gravity supplies exactly the
#    centripetal acceleration the circle needs, and nothing is left over.
print("Why a satellite stays up: gravity is the centripetal force, no more and no less.")
for alt in (780, 20200, 35786):
    r = R + alt * 1000
    v = speed(alt)
    print("  at %6d km: gravity pulls at %.4f m/s2, a circle at %.0f m/s needs %.4f m/s2"
          % (alt, GM / r ** 2, v, v ** 2 / r))
print("  too slow and the orbit falls towards the earth; too fast and it climbs away.")

print()
print("The period and the speed of a circular orbit:")
print("  altitude        period            speed     one-way delay to the sub-satellite point")
for alt in (200, 780, 1200, 10000, 20200, 35786):
    T, v = period(alt), speed(alt)
    print("  %6d km   %8.1f min   %8.0f m/s   %6.1f ms"
          % (alt, T / 60, v, alt * 1000 / C * 1000))

# 2. The elevation angle. A user at earth-central angle gamma from the
#    sub-satellite point sees the satellite at this angle above the horizon.
def elevation(alt_km, gamma_deg):
    r = R + alt_km * 1000
    g = math.radians(gamma_deg)
    # from the plane triangle earth-centre, user, satellite
    return math.degrees(math.atan2(math.cos(g) - R / r, math.sin(g)))

print()
print("How the elevation angle falls as the user moves away from the sub-satellite point:")
print("  ground distance       elevation seen from the ground")
print("  from the sub-point      780 km    20200 km    35786 km")
for gamma in (0, 5, 10, 20, 40, 60, 71):
    arc_km = math.radians(gamma) * R / 1000
    cells = []
    for alt in (780, 20200, 35786):
        e = elevation(alt, gamma)
        cells.append("%8s" % ("%.1f deg" % e if e >= 0 else "below"))
    print("  %5.1f deg %7.0f km %s" % (gamma, arc_km, " ".join(cells)))

# 3. Footprint: the earth-central half-angle for a minimum elevation, and the
#    ground radius that gives.
def footprint(alt_km, elev_deg):
    r = R + alt_km * 1000
    e = math.radians(elev_deg)
    gamma = math.acos(R / r * math.cos(e)) - e
    return math.degrees(gamma), gamma * R / 1000

print()
print("What a minimum elevation costs in footprint (the user must see this high):")
print("  altitude    0 deg minimum       10 deg minimum      30 deg minimum")
for alt in (780, 1200, 20200, 35786):
    out = []
    for e in (0, 10, 30):
        g, radius = footprint(alt, e)
        out.append("%5.1f deg, %5.0f km" % (g, radius))
    print("  %6d km   %s" % (alt, "   ".join(out)))

# 4. How long one satellite stays visible. It sweeps its own footprint at its
#    orbital rate; the earth's rotation is a small correction, ignored here.
print()
print("How long one satellite stays above 10 degrees, passing straight overhead:")
for alt in (780, 1200, 20200):
    T = period(alt)
    g, _ = footprint(alt, 10)
    visible = T * (2 * g / 360.0)
    print("  %6d km   %5.1f min of a %6.1f min orbit" % (alt, visible / 60, T / 60))
print("   35786 km   the satellite keeps pace with the earth, so it never sets")

# 5. An elliptical orbit, the kind used to serve high latitudes. Kepler's
#    equation gives the time spent above any chosen radius.
def time_from_perigee(a, e, r0):
    """Seconds from perigee to the point where the radius reaches r0."""
    cosE = (1 - r0 / a) / e
    cosE = max(-1.0, min(1.0, cosE))
    E = math.acos(cosE)
    return math.sqrt(a ** 3 / GM) * (E - e * math.sin(E))

perigee_km, apogee_km = 1000, 39400
a = R + (perigee_km + apogee_km) / 2.0 * 1000
e = ((apogee_km - perigee_km) * 1000 / 2.0) / a
T = 2 * math.pi * math.sqrt(a ** 3 / GM)
print()
print("An elliptical orbit for high latitudes, perigee %d km and apogee %d km:"
      % (perigee_km, apogee_km))
print("  semi-major axis %.0f km, eccentricity %.3f, period %.1f h"
      % (a / 1000, e, T / 3600))
for above_km in (20000, 30000, 35786):
    t = time_from_perigee(a, e, R + above_km * 1000)
    frac = 1 - 2 * t / T
    print("  it is above %5d km for %.1f h of each orbit, which is %.0f%% of the period"
          % (above_km, frac * T / 3600, frac * 100))
print("  near apogee it moves slowly and hangs over one hemisphere, which is the point:")
print("  a few such satellites serve high latitudes that a geostationary one cannot reach.")

# 6. The delay, and the radiation belts that shape the choice of altitude.
print()
print("Round-trip delay through a satellite, and the altitude bands in use:")
rows = (("LEO", 780, "below the inner radiation belt"),
        ("MEO", 20200, "above the inner belt, in the lower outer belt region"),
        ("GEO", 35786, "above the strongest part of the outer belt, over the equator"))
for name, alt, note in rows:
    hop = 2 * alt * 1000 / C * 1000
    print("  %-4s %6d km   %6.1f ms up and down   %s" % (name, alt, hop, note))
print("  the inner belt is usually given as roughly 2,000 to 6,000 km and the outer as")
print("  roughly 15,000 to 30,000 km; the bands in use are chosen around them, and a")
print("  satellite that must sit inside one is built to withstand the radiation.")
munotes.in796

Satellite Basics: Orbits, Periods, Elevation and Footprints

Why a satellite stays up: gravity is the centripetal force, no more and no less.
  at    780 km: gravity pulls at 7.7793 m/s2, a circle at 7462 m/s needs 7.7793 m/s2
  at  20200 km: gravity pulls at 0.5643 m/s2, a circle at 3873 m/s needs 0.5643 m/s2
  at  35786 km: gravity pulls at 0.2242 m/s2, a circle at 3075 m/s needs 0.2242 m/s2
  too slow and the orbit falls towards the earth; too fast and it climbs away.

The period and the speed of a circular orbit:
  altitude        period            speed     one-way delay to the sub-satellite point
     200 km       88.5 min       7784 m/s      0.7 ms
     780 km      100.5 min       7462 m/s      2.6 ms
    1200 km      109.4 min       7252 m/s      4.0 ms
   10000 km      347.7 min       4933 m/s     33.4 ms
   20200 km      718.7 min       3873 m/s     67.4 ms
   35786 km     1436.1 min       3075 m/s    119.4 ms

How the elevation angle falls as the user moves away from the sub-satellite point:
  ground distance       elevation seen from the ground
  from the sub-point      780 km    20200 km    35786 km
    0.0 deg       0 km 90.0 deg 90.0 deg 90.0 deg
    5.0 deg     557 km 50.3 deg 83.4 deg 84.1 deg
   10.0 deg    1113 km 28.4 deg 76.9 deg 78.2 deg
   20.0 deg    2226 km  8.1 deg 64.0 deg 66.5 deg
   40.0 deg    4453 km    below 39.3 deg 43.7 deg
   60.0 deg    6679 km    below 16.7 deg 21.9 deg
   71.0 deg    7904 km    below  5.2 deg 10.4 deg

What a minimum elevation costs in footprint (the user must see this high):
  altitude    0 deg minimum       10 deg minimum      30 deg minimum
     780 km    27.0 deg,  3005 km    18.7 deg,  2077 km     9.5 deg,  1057 km
    1200 km    32.7 deg,  3639 km    24.0 deg,  2674 km    13.2 deg,  1470 km
   20200 km    76.1 deg,  8473 km    66.3 deg,  7384 km    48.0 deg,  5344 km
   35786 km    81.3 deg,  9050 km    71.4 deg,  7952 km    52.5 deg,  5841 km

How long one satellite stays above 10 degrees, passing straight overhead:
     780 km    10.4 min of a  100.5 min orbit
    1200 km    14.6 min of a  109.4 min orbit
   20200 km   264.8 min of a  718.7 min orbit
   35786 km   the satellite keeps pace with the earth, so it never sets

An elliptical orbit for high latitudes, perigee 1000 km and apogee 39400 km:
  semi-major axis 26578 km, eccentricity 0.722, period 12.0 h
  it is above 20000 km for 8.8 h of each orbit, which is 73% of the period
  it is above 30000 km for 6.3 h of each orbit, which is 53% of the period
  it is above 35786 km for 4.0 h of each orbit, which is 33% of the period
  near apogee it moves slowly and hangs over one hemisphere, which is the point:
  a few such satellites serve high latitudes that a geostationary one cannot reach.

Round-trip delay through a satellite, and the altitude bands in use:
  LEO     780 km      5.2 ms up and down   below the inner radiation belt
  MEO   20200 km    134.8 ms up and down   above the inner belt, in the lower outer belt region
  GEO   35786 km    238.7 ms up and down   above the strongest part of the outer belt, over the equator
  the inner belt is usually given as roughly 2,000 to 6,000 km and the outer as
  roughly 15,000 to 30,000 km; the bands in use are chosen around them, and a
  satellite that must sit inside one is built to withstand the radiation.
munotes.in797

Satellite Basics: Orbits, Periods, Elevation and Footprints

Read the tables against each other and the design of every satellite system is visible in them. Period and speed fall with altitude. Elevation falls away from the sub-satellite point, fast for a low satellite and slowly for a high one: from 780 km a user 2,226 km away sees the satellite at only 8.1 degrees, while from geostationary height the same user sees it at 66.5 degrees. Footprint grows with altitude and shrinks with the minimum elevation demanded. Visibility is minutes for a low satellite and forever for a geostationary one. And delay grows with altitude, which is the bill.

munotes.in798

Satellite Basics: Orbits, Periods, Elevation and Footprints

Distinctions

LEOMEOGEOHEO
AltitudeA few hundred to about 1,500 kmUp to about 20,000 km35,786 kmLow perigee, very high apogee
PeriodAbout 90 to 110 minAbout 6 to 12 h23 h 56 minChosen, often 12 h
Round trip delayAbout 5 msAbout 135 msAbout 239 msVaries through the orbit
Visible forAbout 10 to 15 minHoursAlwaysHours near apogee
Satellites neededMany tensAbout 10 to 303 for the inhabited world3 for one high latitude region
Used forImaging, handheld constellationsNavigationBroadcasting, weather, fixed linksHigh latitude coverage
munotes.in799

Satellite Basics: Orbits, Periods, Elevation and Footprints

Elevation angleEarth-central angle
Measured atThe user, from the local horizonThe earth's centre
Large whenThe satellite is nearly overheadThe user is far from the sub-satellite point
Used forDeciding whether the link is usableComputing the footprint's size
UplinkDownlink
DirectionGround to satelliteSatellite to ground
FrequencyThe higher of the pairThe lower
WhyThe earth station can afford the power and the bigger antennaThe satellite is limited in power and mass

What it does not mean

A satellite is not weightless because gravity is absent. Gravity at geostationary height is still about 0.22 metres per second squared, as the program prints; the satellite is in free fall, which is a different thing.

The period does not depend on the satellite's mass. Two satellites at the same altitude, one a cubesat and one a broadcasting platform, keep exactly the same period.

Elevation is not measured from the vertical. It is measured up from the user's own horizon, so straight overhead is 90 degrees, not 0.

The footprint is not a fixed property of the satellite. It depends on the minimum elevation the system insists on, and the same satellite has a larger footprint for a dish on a roof than for a handset in a street.

A geostationary satellite is not stationary. It travels at about 3,075 metres a second; it only appears fixed because the earth turns underneath it at the same angular rate.

HEO is not a fourth altitude. It is a shape: an ellipse whose apogee is used and whose perigee is passed through quickly.

Quick revision

  • Why it stays up: gravity supplies the centripetal acceleration. Speed is the square root of GM over r; period T equals two pi times the square root of r cubed over GM, Kepler's third law. r is measured from the earth's centre.
  • Numbers: 780 km gives 7,462 m/s and about 100 min; 20,200 km gives 3,873 m/s and about 719 min; 35,786 km gives 3,075 m/s and 1,436 min, which is 23 hours and 56 minutes.
  • Elevation is measured from the user's own horizon; systems set a minimum, often 10 degrees, because low angles mean atmosphere, obstruction and multipath.
  • Footprint radius at 10 degrees: 2,077 km from 780 km, 7,384 km from 20,200 km, 7,952 km from geostationary orbit. Demanding 30 degrees instead cuts the geostationary figure to 5,841 km.
  • Visibility: 10.4 min at 780 km, 14.6 min at 1,200 km, 264.8 min at 20,200 km, never-setting at geostationary height.
  • Delay, round trip: 5.2 ms, 134.8 ms, 238.7 ms for LEO, MEO and GEO.
  • Four orbits: GEO, MEO, LEO, HEO. Belts roughly 2,000 to 6,000 km and 15,000 to 30,000 km.
  • Bands: L and S for handhelds, C at 6 up and 4 down, Ku at 14 up and 11 down, Ka at 30 up and 20 down. Uplink is always the higher.
  • Vocabulary: apogee, perigee, semi-major axis, eccentricity, inclination, sub-satellite point, elevation, footprint, line of sight, uplink, downlink, transponder, inter-satellite link, gateway link, mobile user link.
munotes.in800

Satellite Basics: Orbits, Periods, Elevation and Footprints

Test yourself

1. Why does a satellite stay in orbit, and what decides its period? It is in free fall around the earth. At the right speed for its radius the earth's gravity supplies exactly the centripetal acceleration a circular path of that radius needs, so the satellite keeps turning at a constant distance. Equating the two gives a speed equal to the square root of GM divided by r, and therefore a period equal to two pi times the square root of r cubed divided by GM, which is Kepler's third law. The period depends only on the orbit's radius, measured from the earth's centre, and not at all on the satellite's mass.

2. What is the elevation angle, and why does a system set a minimum for it? It is the angle between the user's own horizon and the direction of the satellite, so a satellite straight overhead is at 90 degrees and one on the horizon is at 0. A system sets a minimum, commonly 10 degrees, because at low angles the signal travels a long way through the atmosphere, is easily blocked by buildings, hills and trees, and picks up multipath from the ground. Below that angle the link is not reliable, so it is treated as unavailable.

3. What is a footprint, and what makes it larger or smaller? It is the part of the earth from which the satellite can be seen at or above the chosen minimum elevation, and therefore the area the satellite can serve. It grows with altitude, because a higher satellite sees more of the earth, and it shrinks as the minimum elevation is raised. From geostationary orbit the footprint radius is about 9,050 km if any elevation above the horizon will do, about 7,952 km at a 10 degree minimum, and about 5,841 km at 30 degrees.

4. Compare LEO, MEO and GEO. LEO lies a few hundred to about 1,500 kilometres up, with a period of roughly 90 to 110 minutes, a round trip delay of about 5 milliseconds, and about 10 to 15 minutes of visibility per pass, so it needs many tens of satellites and constant handover; it suits handheld terminals and imaging. MEO reaches up to about 20,000 kilometres, with a period of several hours, about 135 milliseconds of delay and hours of visibility, and about ten to thirty satellites give global coverage; it suits navigation. GEO sits at 35,786 kilometres over the equator with a period of 23 hours 56 minutes, so the satellite appears fixed and a dish never moves, three satellites cover the inhabited world, but the round trip costs about 239 milliseconds and the poles are not covered.

munotes.in801

Satellite Basics: Orbits, Periods, Elevation and Footprints

5. What is a highly elliptical orbit for? For serving high latitudes, which a geostationary satellite cannot do because from a high latitude a satellite over the equator is low in the sky or below the horizon. An orbit with a low perigee and a very high apogee has a period that can be chosen, for example twelve hours, and a body moves slowly when it is far from the earth: the chapter's example, with a perigee of 1,000 km and an apogee of 39,400 km, is above 20,000 km for 8.8 hours of each 12 hour orbit. The satellite therefore hangs over one hemisphere for most of its orbit, and a few such satellites keep one always high over the region served.

6. Why are some altitudes avoided? Because of the radiation belts, regions where the earth's magnetic field traps charged particles, usually given as roughly 2,000 to 6,000 km and roughly 15,000 to 30,000 km. A satellite that spends its life inside a belt needs heavier shielding and radiation-hardened parts, which cost mass and money, so designers stay below the first belt, aim for the gap between them, or deliberately build the satellite to withstand the radiation where the orbit is worth it.

7. Why is the uplink frequency always higher than the downlink frequency in a band pair? Because the loss grows with frequency and the two ends are not equally able to pay for it. The earth station can use a large antenna and as much transmitter power as it needs, so it can afford the higher frequency and its greater path loss; the satellite is limited in mass, in antenna size and above all in the power its solar panels provide, so the easier, lower frequency is given to the direction it has to transmit in.

Contents This chapter on its own page

munotes.in802

Chapter One Hundred Five

GEO: The Geostationary Orbit

Syllabus topic Module 2, "Satellite Systems: GEO"

In one line

Put a satellite 35,786 km above the equator and it turns with the earth, so it hangs motionless in the sky: that one property buys the fixed dish, the continent-wide footprint and the three-satellite world, and it is paid for with a quarter-second delay and no coverage above about 71 degrees of latitude.

The altitude is not a choice

Everything about GEO follows from a single requirement: the satellite must go round in exactly the time the earth takes to turn once. Then it stays over the same spot, and an antenna on the ground can be aimed once and left.

The period wanted is the sidereal day, 23 hours 56 minutes 4 seconds, which is one turn of the earth against the stars, not the 24 hour solar day we live by. The four minute difference exists because the earth also moves along its orbit around the sun each day, so it must turn a little further than one full rotation to bring the sun back overhead.

Invert Kepler's third law for that period and the radius comes out at 42,164 km from the earth's centre, which is 35,786 km above the surface. That number is not a design decision and not a convention: it is arithmetic, and the program does it rather than quoting it.

The program also prices the mistake of using 24 hours. The radius would be 77 km too high, and the satellite would slip 0.986 degrees of longitude every day, which is a full circuit of the earth in a year. A satellite in it would be geosynchronous in no useful sense at all.

Three more conditions

A period of one sidereal day is necessary but not sufficient. For the satellite to appear fixed rather than merely to return to the same place each day, three more things must hold.

The orbit must be circular. In an ellipse the satellite moves fast at perigee and slowly at apogee, so even with the right period it would run ahead of the earth and fall behind it in turn, tracing an east-west oscillation across the sky.

The inclination must be zero, so the orbit lies in the equatorial plane. A geosynchronous orbit tilted by some angle carries the satellite that far north and then that far south each day, and it traces a figure of eight in the sky, which a fixed dish cannot follow.

The direction must be eastward, the same way the earth turns.

Meet all four and the satellite is geostationary. Meet only the period and it is geosynchronous, which is a weaker thing: it comes back to the same place once a day but does not stay there.

munotes.in803

GEO: The Geostationary Orbit

Station keeping, and why it ends

The orbit does not hold itself. The earth is not a perfect sphere, and its equatorial bulge is not even: the effect is that the ring has a few longitudes a satellite drifts towards and others it drifts away from, so east-west corrections are needed to hold a slot. The sun and the moon pull the orbital plane out of the equator, so north-south corrections are needed too, and they are the expensive ones, consuming the greater part of the fuel a satellite carries. Solar radiation pressure pushes on the panels.

So a geostationary satellite fires thrusters through its whole life to stay where it is said to be, and its life ends when the fuel does, not when the electronics fail. Before the last of it is gone, the operator raises the satellite into a disposal orbit above the ring, so that a dead satellite does not drift through the slots of living ones.

What it reaches, and what it never will

A latitude and longitude map. Three dots sit on the equator, evenly spaced, marked as three satellites 120 degrees apart. Around each dot is a dashed lens-shaped outline, widest at the equator and closing to a point at 71.4 degrees north and at 71.4 degrees south. The three outlines overlap along the equator. Bands across the top above 71.4 degrees north and across the bottom below 71.4 degrees south are shaded and labelled never covered. A note says each dashed outline is the true edge of one satellite's ten degree footprint, that they overlap along the equator and close over it, but that the shaded bands are beyond every one of them

Figure 105.1 The true ten degree footprints of three geostationary satellites, and the polar bands none of them reaches

The program computes how high the satellite sits in the sky from various places. On the equator, under the satellite, it is overhead. From Mumbai's latitude it is still 67.7 degrees up if the longitudes match. At 60 degrees of latitude it has fallen to 21.9 degrees, and at 71 degrees it is 10.4 degrees, scraping the minimum. A binary search in the program finds the limit exactly: 71.4 degrees north or south, on the satellite's own longitude, and less than that anywhere else.

That limit is absolute. It does not improve with a bigger dish or a stronger transmitter, because the satellite is geometrically too low in the sky, and beyond the limit it is under the horizon altogether. No geostationary satellite, and no number of them, covers the poles. That is the single fact that keeps other orbits in business, and the chapter [LEO and MEO] is largely about the systems that exist because of it.

Along the equator the picture is the opposite. One satellite reaches 71.4 degrees of longitude either side of its own, so neighbours parked 120 degrees apart still overlap by 22.9 degrees. Three satellites cover the inhabited world, exactly as [The History of Satellite Systems] records Clarke working out in 1945.

Pointing a dish

Because the satellite never moves, a receiving dish has two numbers and no motor. The program works them out for a real case: from Mumbai at 19.08 degrees north and 72.88 degrees east, to a satellite parked at 83.0 degrees east, the dish must be raised to an elevation of 64.8 degrees and turned to an azimuth of 151.4 degrees from true north, which is south by south-east. The satellite is 36,305 km away, a little further than the 35,786 km directly below it, because the path runs at a slant.

munotes.in804

GEO: The Geostationary Orbit

Set those two angles once, tighten the bolts, and the installation is finished for the life of the dish. Every other orbit needs a tracking mount or an electronically steered array.

What the delay costs

The distance has a price, and it is paid in time. At 36,305 km the one-way trip takes 121.1 ms, so up and down again is 242.2 ms, and a double hop through two satellites is 484.4 ms.

For speech that is noticeable. A speaker hears a reply about a quarter of a second late on one hop, so the natural rhythm of a conversation breaks down and the two ends begin to talk over each other; on two hops it is half a second and the call is hard work. This is why an international call is routed through at most one satellite hop where a choice exists, and why submarine fibre took the traffic, as [Applications of Satellite Systems] sets out.

For data the cost is subtler and often larger. A protocol that may have only one window of data unacknowledged in flight is limited to a window per round trip, no matter how fast the link is. The program computes it over a 272 ms round trip: the classic 64 kB window gives 1.93 Mbit/s, a 256 kB window gives 7.70 Mbit/s, and a megabyte window gives 30.82 Mbit/s. The link's own capacity never appears in that arithmetic. A 50 Mbit/s satellite link can deliver under 2 Mbit/s to a connection with a small window, which is the same reasoning as in [Traditional Transport Control Protocols: TCP and UDP] and the reason satellite links use large windows, window scaling, and sometimes a performance-enhancing proxy that acknowledges locally.

Slots, and why they are scarce

There is only one geostationary ring, and every satellite in it must be far enough from its neighbours that a receiving dish can tell them apart. The program counts what the spacings give: 2 degrees apart, 180 slots, with 1,472 km between neighbours; 1 degree apart, 360 slots; half a degree apart, 720 slots.

Fourteen hundred kilometres sounds like room to spare, and in space it is. The constraint is not space but angle: from the ground, two satellites 2 degrees apart are 2 degrees apart in the sky, and a small dish has a beam wide enough to hear both at once, as [Antennas: Radiators, Dipoles and Radiation Patterns] shows. The dish would receive one satellite's programme with the next one's interference laid over it. So the spacing is set by the smallest dish the service expects, the slots are allocated and coordinated internationally, and a slot over a populous longitude is a genuinely scarce asset.

munotes.in805

GEO: The Geostationary Orbit

The orbit, computed

# The geostationary orbit: where it has to be, what it reaches, and what it costs.
import math

GM, R = 3.986004418e14, 6378137.0
C = 299792458.0
SIDEREAL_DAY = 86164.0905      # one turn of the earth against the stars, seconds
SOLAR_DAY = 86400.0            # one turn against the sun

# 1. The altitude is not chosen. It is whatever makes the period one sidereal
#    day, so invert Kepler's third law rather than quoting the answer.
def radius_for_period(T):
    return (GM * (T / (2 * math.pi)) ** 2) ** (1.0 / 3.0)

r_geo = radius_for_period(SIDEREAL_DAY)
print("Where the geostationary orbit has to be:")
print("  the earth turns once against the stars in %.1f s, which is 23 h 56 min 4 s"
      % SIDEREAL_DAY)
print("  the orbit with that period has radius   %8.0f km" % (r_geo / 1000))
print("  so its altitude above the surface is    %8.0f km" % ((r_geo - R) / 1000))
print("  and the satellite travels at %.0f m/s, eastward, over the equator"
      % math.sqrt(GM / r_geo))

# 2. Why the sidereal day and not the 24 hour solar day: the mistake, in km.
r_wrong = radius_for_period(SOLAR_DAY)
print()
print("Why the sidereal day and not the 24 hour day we live by:")
print("  a 24 h period needs a radius of %.0f km, which is %.0f km too high"
      % (r_wrong / 1000, (r_wrong - r_geo) / 1000))
drift_deg = 360.0 * (SOLAR_DAY - SIDEREAL_DAY) / SIDEREAL_DAY
print("  a satellite with a 24 h period slips %.3f degrees of longitude a day," % drift_deg)
print("  which is about %.0f degrees a year: it would not be stationary at all."
      % (drift_deg * 365.25))

# 3. What a geostationary satellite reaches. A user at latitude lat and
#    longitude offset dlon sees it at this elevation.
def look(lat_deg, dlon_deg, r_sat=None):
    """Elevation, azimuth from true north, and slant range, by vectors.

    dlon_deg is the satellite's longitude minus the user's, so a positive
    value puts the satellite east of the user.
    """
    r_sat = r_sat or r_geo
    lat, dlon = math.radians(lat_deg), math.radians(-dlon_deg)
    # the user on a spherical earth, with the satellite's longitude as zero
    u = (R * math.cos(lat) * math.cos(dlon), R * math.cos(lat) * math.sin(dlon),
         R * math.sin(lat))
    s = (r_sat, 0.0, 0.0)
    d = tuple(s[i] - u[i] for i in range(3))
    # the local east, north and up directions at the user
    up = (math.cos(lat) * math.cos(dlon), math.cos(lat) * math.sin(dlon), math.sin(lat))
    east = (-math.sin(dlon), math.cos(dlon), 0.0)
    north = (-math.sin(lat) * math.cos(dlon), -math.sin(lat) * math.sin(dlon), math.cos(lat))
    dot = lambda a, b: sum(a[i] * b[i] for i in range(3))
    e, n, u2 = dot(d, east), dot(d, north), dot(d, up)
    rng = math.sqrt(e * e + n * n + u2 * u2)
    return math.degrees(math.asin(u2 / rng)), math.degrees(math.atan2(e, n)) % 360, rng

print()
print("How high a geostationary satellite sits in the sky, by latitude and longitude offset:")
print("  latitude   same longitude   30 deg away      60 deg away")
for lat in (0, 19, 40, 60, 71, 75):
    cells = []
    for dlon in (0, 30, 60):
        el, _, _ = look(lat, dlon)
        cells.append("%-16s" % ("%.1f deg up" % el if el > 0 else "below the horizon"))
    print("  %5d deg  %s" % (lat, "".join(cells).rstrip()))
print("  the highest latitude that sees it ten degrees up, on its own longitude, is")
lo, hi = 0.0, 89.0
for _ in range(60):
    mid = (lo + hi) / 2
    if look(mid, 0)[0] >= 10.0:
        lo = mid
    else:
        hi = mid
print("  %.1f degrees north or south, so the poles are never covered." % lo)

# 4. Pointing a dish. One real place, one real orbital slot.
print()
print("Pointing a dish from Mumbai, 19.08 N 72.88 E, at a satellite parked at 83.0 E:")
el, az, rng = look(19.076, 83.0 - 72.877)
print("  elevation %.1f degrees, azimuth %.1f degrees from true north, range %.0f km"
      % (el, az, rng / 1000))
print("  the dish is aimed once and bolted: the satellite does not move in that sky.")

# 5. What the distance costs in time.
print()
print("The delay a geostationary link pays:")
one_way = rng / C
print("  one way, Mumbai to the satellite:        %6.1f ms" % (one_way * 1000))
print("  up and down again, one hop:              %6.1f ms" % (2 * one_way * 1000))
print("  a double hop, through two satellites:    %6.1f ms" % (4 * one_way * 1000))
print("  a speaker hears a reply a quarter of a second late on one hop, and half a")
print("  second late on two, which is why a double hop is avoided in a phone call.")

# 6. What the delay costs a protocol that waits for acknowledgements.
print()
print("What that delay costs TCP, which may only have a window in flight:")
rtt = 2 * one_way + 0.030            # the satellite hop plus 30 ms of terrestrial tails
for window_kb, label in ((64, "the classic 64 kB window"),
                         (256, "a 256 kB window"),
                         (1024, "a 1 MB window")):
    rate = window_kb * 1024 * 8 / rtt
    print("  %-24s gives %6.2f Mbit/s over a %.0f ms round trip"
          % (label, rate / 1e6, rtt * 1000))
print("  the link's own speed never enters this: the window and the round trip decide it.")

# 7. How many satellites fit, and how far apart they really are.
print()
print("Orbital slots along the geostationary ring:")
circumference = 2 * math.pi * r_geo
for sep in (2.0, 1.0, 0.5):
    print("  %.1f degrees apart: %3d slots in the ring, %5.0f km between neighbours"
          % (sep, int(360 / sep), circumference * sep / 360 / 1000))
print("  they look close from the ground, so a small dish with a wide beam hears both:")
print("  that, not space, is what makes a slot scarce.")

# 8. Three satellites, and the gap they leave.
print()
print("Three satellites, 120 degrees apart, at a 10 degree minimum elevation:")
e = math.radians(10.0)
gamma = math.degrees(math.acos(R / r_geo * math.cos(e)) - e)
cap = 2 * math.pi * R ** 2 * (1 - math.cos(math.radians(gamma)))
print("  each covers %.1f degrees of earth-central angle, or %.1f million km2"
      % (gamma, cap / 1e12))
print("  so one reaches %.1f degrees of longitude either side of its own along the equator"
      % gamma)
print("  and neighbours parked 120 degrees apart still overlap by %.1f degrees there,"
      % (2 * gamma - 120))
print("  but nothing above %.1f degrees of latitude is covered, however many are launched."
      % lo)
munotes.in806

GEO: The Geostationary Orbit

Where the geostationary orbit has to be:
  the earth turns once against the stars in 86164.1 s, which is 23 h 56 min 4 s
  the orbit with that period has radius      42164 km
  so its altitude above the surface is       35786 km
  and the satellite travels at 3075 m/s, eastward, over the equator

Why the sidereal day and not the 24 hour day we live by:
  a 24 h period needs a radius of 42241 km, which is 77 km too high
  a satellite with a 24 h period slips 0.986 degrees of longitude a day,
  which is about 360 degrees a year: it would not be stationary at all.

How high a geostationary satellite sits in the sky, by latitude and longitude offset:
  latitude   same longitude   30 deg away      60 deg away
      0 deg  90.0 deg up     55.0 deg up     21.9 deg up
     19 deg  67.7 deg up     49.3 deg up     20.0 deg up
     40 deg  43.7 deg up     34.4 deg up     14.1 deg up
     60 deg  21.9 deg up     17.4 deg up     5.8 deg up
     71 deg  10.4 deg up     7.8 deg up      0.7 deg up
     75 deg  6.4 deg up      4.3 deg up      below the horizon
  the highest latitude that sees it ten degrees up, on its own longitude, is
  71.4 degrees north or south, so the poles are never covered.

Pointing a dish from Mumbai, 19.08 N 72.88 E, at a satellite parked at 83.0 E:
  elevation 64.8 degrees, azimuth 151.4 degrees from true north, range 36305 km
  the dish is aimed once and bolted: the satellite does not move in that sky.

The delay a geostationary link pays:
  one way, Mumbai to the satellite:         121.1 ms
  up and down again, one hop:               242.2 ms
  a double hop, through two satellites:     484.4 ms
  a speaker hears a reply a quarter of a second late on one hop, and half a
  second late on two, which is why a double hop is avoided in a phone call.

What that delay costs TCP, which may only have a window in flight:
  the classic 64 kB window gives   1.93 Mbit/s over a 272 ms round trip
  a 256 kB window          gives   7.70 Mbit/s over a 272 ms round trip
  a 1 MB window            gives  30.82 Mbit/s over a 272 ms round trip
  the link's own speed never enters this: the window and the round trip decide it.

Orbital slots along the geostationary ring:
  2.0 degrees apart: 180 slots in the ring,  1472 km between neighbours
  1.0 degrees apart: 360 slots in the ring,   736 km between neighbours
  0.5 degrees apart: 720 slots in the ring,   368 km between neighbours
  they look close from the ground, so a small dish with a wide beam hears both:
  that, not space, is what makes a slot scarce.

Three satellites, 120 degrees apart, at a 10 degree minimum elevation:
  each covers 71.4 degrees of earth-central angle, or 174.2 million km2
  so one reaches 71.4 degrees of longitude either side of its own along the equator
  and neighbours parked 120 degrees apart still overlap by 22.9 degrees there,
  but nothing above 71.4 degrees of latitude is covered, however many are launched.
munotes.in807

GEO: The Geostationary Orbit

Advantages and disadvantages

The examination asks for these by name, so here they are as a list, each one traceable to a number above.

munotes.in808

GEO: The Geostationary Orbit

Advantages. The satellite appears fixed, so ground antennas need no tracking and no motor, and installation is cheap. There is no handover caused by the satellite moving, so a link lasts as long as the equipment does. Three satellites cover the inhabited world. The footprint is huge, about a third of the earth, which is what makes broadcasting economic. There is almost no Doppler shift for a fixed user, because the satellite does not move relative to the dish, which removes the problem [Multipath, Fading and the Doppler Effect] describes. The satellite is always available, with no gaps to schedule around.

Disadvantages. The delay is about a quarter of a second round trip, which speech notices and acknowledged protocols pay for. No coverage above about 71 degrees of latitude, and none at all at the poles. The distance means high path loss, so both ends need power and gain, and a handheld terminal cannot close the link at broadcast frequencies. Launching to 35,786 km is expensive, and the satellite must be large and long-lived to be worth it. Slots and frequencies are scarce and must be coordinated. Station keeping consumes fuel, and the fuel decides the satellite's life. And a single satellite is a single point of failure for everything under it.

munotes.in809

GEO: The Geostationary Orbit

Distinctions

GeostationaryGeosynchronous
PeriodOne sidereal dayOne sidereal day
InclinationZero, in the equatorial planeAny
EccentricityZero, circularAny
Seen from the groundMotionlessReturns daily, tracing a figure of eight or an oscillation
AntennaFixedTracking
Sidereal daySolar day
Measured againstThe starsThe sun
Length23 h 56 min 4 s24 h
Radius that gives it42,164 km42,241 km
Used for GEOYesNo: the satellite would slip 0.986 degrees a day
GEOThe alternatives
AntennaFixed, aimed onceTracking or steered
DelayAbout 242 ms round tripAbout 5 ms in low orbit
PolesNeverCovered by polar and inclined orbits
Satellites for global service3, and not the polesTens in low orbit
HandoverNoneEvery few minutes in low orbit

What it does not mean

Geostationary does not mean motionless. The satellite travels at about 3,075 metres a second; it only appears fixed because the earth turns underneath it at the same rate.

It does not stay put by itself. Without station keeping it drifts in longitude and its plane tilts, so it fires thrusters throughout its life and its life ends with the fuel.

The period is not 24 hours. It is the sidereal day, and using 24 hours would put the satellite 77 km too high and let it slip a degree of longitude a day.

Three satellites do not cover the earth. They cover the inhabited world between about 71 degrees north and south. The poles are outside every geostationary footprint that exists or ever will.

A bigger dish does not extend the coverage. Beyond the latitude limit the satellite is too low in the sky or below the horizon, and gain cannot raise it.

The delay is not a fault of the equipment. It is the distance divided by the speed of light, and no improvement in electronics will reduce it.

Quick revision

  • Altitude 35,786 km, radius 42,164 km, over the equator, eastward, circular, zero inclination. The period is the sidereal day, 23 h 56 min 4 s.
  • A 24 h period would need 42,241 km and would slip 0.986 degrees of longitude a day.
  • Geostationary needs period, circularity, zero inclination and direction; geosynchronous needs only the period.
  • Latitude limit 71.4 degrees at a 10 degree minimum elevation, on the satellite's own longitude: the poles are never covered.
  • One satellite reaches 71.4 degrees of longitude either side, so three 120 degrees apart overlap by 22.9 degrees at the equator and cover the inhabited world.
  • From Mumbai to a satellite at 83 degrees east: elevation 64.8 degrees, azimuth 151.4 degrees, range 36,305 km.
  • Delay: 121.1 ms one way, 242.2 ms one hop, 484.4 ms double hop.
  • Throughput over a 272 ms round trip: 1.93 Mbit/s with a 64 kB window, 7.70 with 256 kB, 30.82 with 1 MB.
  • Slots: 2 degrees apart gives 180 slots and 1,472 km of separation; the limit is the angle a small dish can resolve, not the space.
  • Advantages: fixed antenna, no handover, huge footprint, three satellites, no Doppler, always available. Disadvantages: quarter-second delay, no polar coverage, high path loss, expensive launch, scarce slots, station keeping, single point of failure.
munotes.in810

GEO: The Geostationary Orbit

Test yourself

1. Why is the geostationary altitude 35,786 km? Because the satellite must go round in exactly the time the earth takes to turn once against the stars, which is the sidereal day of 23 hours 56 minutes 4 seconds. Kepler's third law gives the period from the orbit's radius; inverting it for that period gives a radius of 42,164 kilometres from the earth's centre, and subtracting the earth's radius leaves 35,786 kilometres above the surface. The altitude is therefore arithmetic, not a design choice.

2. What is the difference between a geostationary and a geosynchronous orbit? A geosynchronous orbit only has to have a period of one sidereal day, so the satellite returns to the same place in the sky once a day. A geostationary orbit must also be circular, lie in the equatorial plane and run eastward. Without circularity the satellite runs ahead and falls behind, oscillating east and west; with any inclination it swings north and south and traces a figure of eight. Only when all four conditions hold does the satellite appear motionless, which is what lets a dish be fixed.

3. Why is the sidereal day used rather than the 24 hour day? Because the sidereal day is one true rotation of the earth, while the solar day is about four minutes longer, since the earth must turn a little past one rotation to bring the sun overhead again as it moves along its own orbit. A satellite matched to 24 hours would be 77 kilometres too high, would slip 0.986 degrees of longitude each day, and would drift right round the earth in a year, so it would not be stationary over anything.

munotes.in811

GEO: The Geostationary Orbit

4. Why can a geostationary satellite never serve the poles? Because it sits over the equator, so the further a user is from the equator the lower it appears. The chapter's search finds the limit exactly: at 71.4 degrees of latitude, on the satellite's own longitude, the satellite is only ten degrees above the horizon, and beyond that it falls below the usable angle and then below the horizon. The limit is geometric, so no dish size, transmitter power or number of satellites can change it.

5. List the advantages and disadvantages of GEO. Advantages: the satellite appears fixed, so ground antennas need no tracking and installation is cheap; there is no handover caused by satellite movement; three satellites cover the inhabited world; the footprint of about a third of the earth makes broadcasting economic; there is almost no Doppler shift for a fixed user; and the satellite is always available. Disadvantages: a round trip delay of about 242 milliseconds, which speech notices and acknowledged protocols pay for; no coverage above about 71 degrees of latitude and none at the poles; very high path loss, so both ends need power and gain; an expensive launch and a large satellite; scarce, internationally coordinated slots and frequencies; station keeping that consumes fuel and so decides the satellite's life; and a single satellite as a single point of failure.

6. Why does a 50 Mbit/s geostationary link often deliver far less than 50 Mbit/s? Because a protocol that may only have one window of unacknowledged data in flight can send a window per round trip and no more, whatever the link's speed. Over the chapter's 272 millisecond round trip, a 64 kilobyte window yields 1.93 Mbit/s, a 256 kilobyte window 7.70 Mbit/s and a one megabyte window 30.82 Mbit/s. The remedy is a larger window, window scaling, or a proxy that acknowledges data locally, not a faster link.

7. Why are geostationary slots scarce when the satellites are more than a thousand kilometres apart? Because what matters on the ground is angle, not distance. Two satellites two degrees apart along the ring are 1,472 kilometres apart in space but only two degrees apart as seen from a dish, and a small dish has a beam wide enough to receive both at once, so one satellite's signal interferes with the other's. The usable spacing is therefore set by the smallest dish the service expects, the ring holds only as many slots as that allows, and a slot over a populous longitude is genuinely valuable.

Contents This chapter on its own page

munotes.in812

Chapter One Hundred Six

LEO and MEO

Syllabus topic Module 2, "Satellite Systems: LEO, MEO"

In one line

Come down from geostationary orbit and the link gets easier and the delay gets shorter, but the satellite stops standing still: low orbit buys a handheld telephone at the price of constant handover and a constellation of dozens, and medium orbit buys several satellites in view at once, which is exactly what a navigation fix requires.

Why anything flies below geostationary orbit

Two numbers from the program answer it.

The first is path loss. At 1,600 MHz, the frequency a handheld terminal can use, a satellite directly overhead at 780 km costs 154.4 dB and a geostationary satellite costs 187.6 dB. The difference is 33.2 dB, a factor of about 2,100. A telephone with a stub antenna and a battery cannot find 33 dB from anywhere, so a handheld satellite phone must talk to a low satellite. The reasoning is the one in [Signal Propagation: Ranges, Path Loss and How a Signal Travels]: loss grows with the square of the distance, and distance is the only term a low orbit changes.

The second is latitude. A geostationary satellite is below ten degrees of elevation beyond 71.4 degrees north or south, so the poles are unreachable. An inclined or polar low orbit passes over everywhere.

Against those, low orbit gives up the one property that made GEO valuable: the satellite moves.

The trade the altitude makes

A chart with altitude in kilometres on a logarithmic horizontal axis from about 500 to 35,786, and a logarithmic vertical axis from 1 to 300. A solid falling curve is labelled satellites needed for global cover, starting above 100 at low altitude and descending in steps to 3 at the right. A dashed rising curve is labelled round trip delay in milliseconds, starting near 2 and rising to about 240 at the right. The two curves cross at about 2,000 kilometres. A shaded vertical band from 2,000 to 6,000 kilometres covers the crossing. Dotted vertical lines mark LEO, MEO and GEO. A note says both axes are logarithmic, that the two costs cross at about 2,000 kilometres and that nothing flies there because the inner radiation belt begins, and that the orbits in use lie on either side

Figure 106.1 The two costs of altitude cross exactly where the inner radiation belt begins

Altitude is bought and paid for in two currencies at once, and they move in opposite directions. Go lower and the delay falls but the number of satellites rises; go higher and the reverse. The figure plots both from the same formulas the program uses, and they cross at about 2,000 km.

What makes the picture interesting is that nothing flies at the crossing. Around 2,000 km the inner radiation belt begins, so the altitude where the two costs balance is the one altitude a designer must avoid. The orbits in use are pushed to either side of it, which is why there are three named bands and not a smooth continuum.

Low earth orbit

LEO is roughly a few hundred kilometres to about 1,500 km. From the program, at 780 km the period is 100.5 minutes and the satellite is above ten degrees for at most 10.4 minutes of it; at 1,414 km the period is 114.1 minutes and the best pass is 16.7 minutes; at 550 km, where the broadband constellations fly, a pass is 7.9 minutes.

Coverage. One satellite at 780 km sees 2.6 per cent of the earth. Even if footprints tiled a sphere perfectly, which circles cannot do, 39 would be needed for global coverage, and the program shows Iridium flies 66, which is 1.7 times that floor. The excess pays for the fact that orbits are planes rather than a free scatter, and that coverage has to hold at every instant, not on average.

munotes.in813

LEO and MEO

Handover, constantly. Because a satellite is usable for at best ten minutes, a thirty minute call on a 780 km constellation is handed to the next satellite at least three times, and more when the pass is not overhead. This is a handover with no counterpart in a terrestrial network: the user has not moved at all, the network has. [Localization and Handover in Satellite Systems] gives it a name and a procedure.

Doppler, constantly. A satellite closing at 6,548 metres a second shifts a 1,620 MHz carrier by 35.4 kHz, and the shift reverses sign as the satellite passes over. A receiver must search a wide band to find the carrier at all and then track it as it slides, which is real work and real power: [Multipath, Fading and the Doppler Effect] sets out the mechanism, and a fixed dish on a geostationary satellite never meets it.

The systems. The classification in common use divides them by what they carry. Little LEO systems work below 1 GHz at low data rates, for messaging, tracking and telemetry. Big LEO systems carry voice: Iridium, 66 satellites in near-polar orbit at about 780 km, notable for carrying traffic between satellites rather than dropping it to the ground at every hop, which [Routing in Satellite Systems] takes up; and Globalstar, 48 satellites at about 1,414 km in inclined orbits, which does the opposite and relays each call straight down to a gateway, so a call only works where a gateway is also in view. Broadband LEO constellations fly lower still, around 550 km, with thousands of satellites and steered beams, trading an enormous constellation for low delay and high capacity.

Medium earth orbit

MEO is the band between the belts and below geostationary orbit, and in practice it means about 19,000 to 23,000 km, where the navigation systems are. From the program, GPS at 20,200 km has a period of 718.7 minutes, which is close to half a day, and a satellite is usable for 264.8 minutes, over four hours at a time. One satellite sees 29.9 per cent of the earth.

That last number is the point of MEO, and it is a different point from LEO's. A navigation receiver does not want one satellite in view; it wants four, for the reason [Applications of Satellite Systems] gives, three coordinates and the receiver's own clock error. Spread 24 satellites evenly and the expected number above ten degrees is, by the program, 7.2. That is a comfortable margin over four, which is why GPS is 24 satellites and not 12: the requirement is not coverage but redundancy of geometry, everywhere, all the time.

munotes.in814

LEO and MEO

The systems: GPS, nominally 24 satellites in six planes at about 20,200 km; GLONASS at about 19,100 km; Galileo at about 23,222 km, where the program gives 7.5 satellites in view on average; BeiDou, which mixes medium, geostationary and inclined geosynchronous orbits; and India's NavIC, which is regional and uses geostationary and inclined geosynchronous satellites to serve India and its neighbourhood rather than the world. A medium orbit is also used for broadband by a small constellation at about 8,000 km, which sits between the delay of geostationary orbit and the constellation size of a low one.

The orbits, computed

# Low and medium orbits: what they buy, what they cost, and how many it takes.
import math

GM, R = 3.986004418e14, 6378137.0
C = 299792458.0

def period(alt_km):
    return 2 * math.pi * math.sqrt((R + alt_km * 1000) ** 3 / GM)

def cap(alt_km, elev_deg=10.0):
    """Earth-central half-angle, and the fraction of the earth, in view."""
    r = R + alt_km * 1000
    e = math.radians(elev_deg)
    g = math.acos(R / r * math.cos(e)) - e
    return g, (1 - math.cos(g)) / 2.0

def slant(alt_km, gamma):
    r = R + alt_km * 1000
    return math.sqrt(R * R + r * r - 2 * R * r * math.cos(gamma))

# 1. The orbits themselves.
print("The orbits, with a ten degree minimum elevation:")
print("  system          altitude   period    visible per pass   share of the earth in view")
for name, alt in (("Iridium", 780), ("Globalstar", 1414), ("a broadband LEO", 550),
                  ("GPS", 20200), ("GLONASS", 19100), ("Galileo", 23222)):
    T = period(alt)
    g, frac = cap(alt)
    print("  %-15s %6d km %7.1f min %10.1f min %14.1f%%"
          % (name, alt, T / 60, T * (2 * math.degrees(g) / 360.0) / 60, 100 * frac))

# 2. How many are overhead. Spread N satellites evenly over the sphere and the
#    expected number in view is N times the share in view. Real constellations
#    are not evenly spread, so this is an average, not a guarantee.
print()
print("How many satellites a receiver can expect to see at once, above 10 degrees:")
for name, alt, n in (("Iridium, 66 satellites", 780, 66),
                     ("Globalstar, 48 satellites", 1414, 48),
                     ("GPS, 24 satellites", 20200, 24),
                     ("Galileo, 24 satellites", 23222, 24)):
    frac = cap(alt)[1]
    print("  %-26s %5.1f on average, and a navigation fix needs four"
          % (name, n * frac))
print("  a low constellation gives about one satellite at a time, so a call is handed on;")
print("  a medium one gives several at once, which is what a position fix requires.")

# 3. The floor on constellation size, and why the real number is larger.
print()
print("The fewest satellites that could cover the earth, if footprints tiled perfectly:")
for name, alt, real in (("Iridium", 780, 66), ("Globalstar", 1414, 48), ("GPS", 20200, 24)):
    frac = cap(alt)[1]
    print("  %-11s floor %3d, actually flies %2d, which is %.1f times the floor"
          % (name, math.ceil(1 / frac), real, real / math.ceil(1 / frac)))
print("  circles do not tile a sphere, orbits are planes rather than a free scatter, and")
print("  coverage must hold at every moment, so the real number is always the larger one.")

# 4. Handover: how often, over a call.
print()
print("How often a call must be handed to the next satellite:")
for name, alt in (("Iridium", 780), ("Globalstar", 1414)):
    T = period(alt)
    g, _ = cap(alt)
    best = T * (2 * math.degrees(g) / 360.0) / 60
    print("  %-11s at best %4.1f min on one satellite, so a 30 min call is handed on at"
          % (name, best))
    print("              least %d times, and more when the pass is not overhead"
          % math.ceil(30.0 / best))
print("  and that is for a user standing still: the network moves, not the caller.")

# 5. What the shorter distance is worth. Free space loss, Recommendation
#    ITU-R P.525: 32.44 dB plus 20 log10 of MHz plus 20 log10 of km.
def fsl(f_mhz, d_km):
    return 32.44 + 20 * math.log10(f_mhz) + 20 * math.log10(d_km)

print()
print("Why a handheld terminal needs a low orbit, at 1600 MHz:")
for name, alt in (("LEO at 780 km", 780), ("LEO at 1414 km", 1414),
                  ("MEO at 20200 km", 20200), ("GEO at 35786 km", 35786)):
    g, _ = cap(alt)
    d_over = alt
    d_edge = slant(alt, g) / 1000
    print("  %-16s overhead %8.1f dB, at the footprint edge %8.1f dB"
          % (name, fsl(1600, d_over), fsl(1600, d_edge)))
gain = fsl(1600, 35786) - fsl(1600, 780)
print("  the low orbit is %.1f dB easier than geostationary overhead, a factor of %.0f,"
      % (gain, 10 ** (gain / 10)))
print("  which is the whole of the difference between a dish and a telephone in a hand.")

# 6. Doppler, the price of a moving satellite. The satellite sweeps the central
#    angle at its orbital rate; the range changes fastest at the horizon.
print()
print("Doppler shift, for a satellite passing overhead:")
for name, alt, f_mhz in (("Iridium, 1620 MHz", 780, 1620.0),
                         ("Globalstar, 1610 MHz", 1414, 1610.0),
                         ("GPS L1, 1575 MHz", 20200, 1575.42)):
    r = R + alt * 1000
    g, _ = cap(alt)
    rate = 2 * math.pi / period(alt)                  # radians per second of central angle
    d = slant(alt, g)
    closing = R * r * math.sin(g) / d * rate          # metres per second, at the edge
    print("  %-22s closes at %5.0f m/s, so the carrier shifts by %6.1f kHz"
          % (name, closing, closing / C * f_mhz * 1000))
print("  the receiver must search for the carrier and track it as it slides, and the shift")
print("  reverses sign as the satellite passes over: that is work a fixed dish never does.")

# 7. The three orbits, side by side.
print()
print("The three orbits, side by side:")
print("  %-22s %10s %10s %10s" % ("", "LEO", "MEO", "GEO"))
ALTS = (780, 20200, 35786)
def row(label, fn, fmt=" %10.1f"):
    print(("  %-22s" + fmt * 3) % ((label,) + tuple(fn(a) for a in ALTS)))
row("altitude, km", lambda a: float(a), " %10.0f")
row("period, minutes", lambda a: period(a) / 60)
row("earth in view, per cent", lambda a: 100 * cap(a)[1])
row("round trip delay, ms", lambda a: 2 * a * 1000 / C * 1000)
row("loss at 1600 MHz, dB", lambda a: fsl(1600, a))
print("  %-22s %10.1f %10.1f %10s"
      % ("visible for, minutes", period(780) * (2 * math.degrees(cap(780)[0]) / 360.0) / 60,
         period(20200) * (2 * math.degrees(cap(20200)[0]) / 360.0) / 60, "always"))
munotes.in815

LEO and MEO

The orbits, with a ten degree minimum elevation:
  system          altitude   period    visible per pass   share of the earth in view
  Iridium            780 km   100.5 min       10.4 min            2.6%
  Globalstar        1414 km   114.1 min       16.7 min            5.2%
  a broadband LEO    550 km    95.6 min        7.9 min            1.7%
  GPS              20200 km   718.7 min      264.8 min           29.9%
  GLONASS          19100 km   674.5 min      246.3 min           29.4%
  Galileo          23222 km   844.7 min      317.9 min           31.1%

How many satellites a receiver can expect to see at once, above 10 degrees:
  Iridium, 66 satellites       1.7 on average, and a navigation fix needs four
  Globalstar, 48 satellites    2.5 on average, and a navigation fix needs four
  GPS, 24 satellites           7.2 on average, and a navigation fix needs four
  Galileo, 24 satellites       7.5 on average, and a navigation fix needs four
  a low constellation gives about one satellite at a time, so a call is handed on;
  a medium one gives several at once, which is what a position fix requires.

The fewest satellites that could cover the earth, if footprints tiled perfectly:
  Iridium     floor  39, actually flies 66, which is 1.7 times the floor
  Globalstar  floor  20, actually flies 48, which is 2.4 times the floor
  GPS         floor   4, actually flies 24, which is 6.0 times the floor
  circles do not tile a sphere, orbits are planes rather than a free scatter, and
  coverage must hold at every moment, so the real number is always the larger one.

How often a call must be handed to the next satellite:
  Iridium     at best 10.4 min on one satellite, so a 30 min call is handed on at
              least 3 times, and more when the pass is not overhead
  Globalstar  at best 16.7 min on one satellite, so a 30 min call is handed on at
              least 2 times, and more when the pass is not overhead
  and that is for a user standing still: the network moves, not the caller.

Why a handheld terminal needs a low orbit, at 1600 MHz:
  LEO at 780 km    overhead    154.4 dB, at the footprint edge    163.9 dB
  LEO at 1414 km   overhead    159.5 dB, at the footprint edge    167.4 dB
  MEO at 20200 km  overhead    182.6 dB, at the footprint edge    184.4 dB
  GEO at 35786 km  overhead    187.6 dB, at the footprint edge    188.7 dB
  the low orbit is 33.2 dB easier than geostationary overhead, a factor of 2105,
  which is the whole of the difference between a dish and a telephone in a hand.

Doppler shift, for a satellite passing overhead:
  Iridium, 1620 MHz      closes at  6548 m/s, so the carrier shifts by   35.4 kHz
  Globalstar, 1610 MHz   closes at  5765 m/s, so the carrier shifts by   31.0 kHz
  GPS L1, 1575 MHz       closes at   915 m/s, so the carrier shifts by    4.8 kHz
  the receiver must search for the carrier and track it as it slides, and the shift
  reverses sign as the satellite passes over: that is work a fixed dish never does.

The three orbits, side by side:
                                LEO        MEO        GEO
  altitude, km                  780      20200      35786
  period, minutes             100.5      718.7     1436.1
  earth in view, per cent        2.6       29.9       34.1
  round trip delay, ms          5.2      134.8      238.7
  loss at 1600 MHz, dB        154.4      182.6      187.6
  visible for, minutes         10.4      264.8     always
munotes.in816

LEO and MEO

Distinctions

LEOMEOGEO
AltitudeAbout 500 to 1,500 kmAbout 8,000 to 23,000 km35,786 km
Period95 to 115 min6 to 14 h23 h 56 min
In view of one satellite1.7 to 5.2 per cent of the earthAbout 30 per cent34.1 per cent
Usable for8 to 17 min a passOver 4 h a passAlways
Round trip delayAbout 5 msAbout 135 msAbout 239 ms
Loss at 1600 MHz, overhead154.4 dB182.6 dB187.6 dB
DopplerTens of kilohertzA few kilohertzEffectively none
Satellites for global serviceDozens to thousands24 to 303, and no poles
HandoverEvery few minutesEvery few hoursNone
Ground antennaTracking or steeredTrackingFixed
Used forHandheld voice, broadband, imagingNavigation, some broadbandBroadcasting, weather, fixed links
munotes.in817

LEO and MEO

IridiumGlobalstar
AltitudeAbout 780 kmAbout 1,414 km
Satellites6648
InclinationNear-polarInclined, not polar
Between satellitesLinks satellite to satelliteNone: every call goes straight down
ConsequenceWorks far from any gateway, including the polesNeeds a gateway in the same footprint
CostComplex satellites and routingSimpler satellites, incomplete coverage
munotes.in818

LEO and MEO

Why LEO needs many satellitesWhy MEO needs many satellites
The requirementAt least one in view, everywhere, alwaysAt least four in view with good geometry, everywhere, always
One satellite sees2.6 per cent of the earth29.9 per cent
Driven bySmall footprintsThe fix, not the coverage

What it does not mean

LEO is not simply a cheaper GEO. It is a different design with different problems: the satellites are cheaper to launch but there must be dozens, they must be replaced far more often, and the ground segment and the handover machinery are far more complex.

A big constellation is not needed because space is crowded. It is needed because a low satellite sees very little of the earth and does not stay put.

MEO is not chosen for its delay. At about 135 ms round trip it is much closer to GEO than to LEO. It is chosen because several satellites are visible at once from everywhere.

Doppler is not a detail. Tens of kilohertz at 1.6 GHz forces the receiver to search and track, and a system designed as if the satellite were fixed would not acquire the carrier at all.

Globalstar's lack of links between satellites is not an oversight. It is a deliberate trade: simpler, cheaper satellites, in exchange for needing a gateway within the same footprint as the user.

NavIC is not a smaller GPS. It is regional by design, and it uses geostationary and inclined geosynchronous orbits rather than the medium orbits GPS uses.

Quick revision

  • LEO: about 500 to 1,500 km. At 780 km, period 100.5 min, best pass 10.4 min, 2.6 per cent of the earth in view, round trip 5.2 ms, loss at 1600 MHz overhead 154.4 dB.
  • MEO: about 19,000 to 23,000 km. At 20,200 km, period 718.7 min, pass 264.8 min, 29.9 per cent in view, round trip 134.8 ms, loss 182.6 dB.
  • Why low: 33.2 dB easier than geostationary at 1600 MHz, a factor of about 2,100, which is what a handheld needs; and the poles are reachable.
  • Why many: one low satellite sees 2.6 per cent, so the floor for global coverage is 39 and Iridium flies 66.
  • Why 24 for GPS: the need is four in view with good geometry, everywhere; 24 gives 7.2 on average.
  • Handover: at best every 10.4 min in LEO, so a 30 min call is handed on at least three times, caused by the network moving.
  • Doppler: 35.4 kHz at 1620 MHz in LEO, 4.8 kHz at GPS L1, effectively none at GEO.
  • Systems: Iridium 66 at 780 km with inter-satellite links; Globalstar 48 at 1,414 km without them; broadband constellations near 550 km; GPS, GLONASS, Galileo, BeiDou; NavIC regional.
munotes.in819

LEO and MEO

Test yourself

1. Why does a handheld satellite telephone need a low orbit? Because of the link budget. At 1,600 MHz a satellite directly overhead at 780 kilometres costs 154.4 dB of free space loss, while a geostationary satellite costs 187.6 dB. The difference of 33.2 dB is a factor of about 2,100, and a handset with a stub antenna and a small battery has no way to find it. The low orbit supplies that margin by being close, which no amount of design in the handset could do.

2. Why does a low constellation need so many satellites? Because one satellite at 780 kilometres sees only 2.6 per cent of the earth and does not stay over any part of it. Even if footprints tiled a sphere perfectly, 39 would be needed for continuous global coverage, and since circles do not tile a sphere, orbits are confined to planes, and the coverage must hold at every instant rather than on average, the real number is larger: Iridium flies 66.

3. What kind of handover does a low constellation force, and why is it unusual? Inter-satellite handover, caused by the satellite setting rather than by the user moving. A satellite at 780 kilometres is above ten degrees for at most 10.4 minutes, so a thirty minute call is passed to the next satellite at least three times even if the caller stands perfectly still. In a terrestrial network handover means the user has moved; here the network moves and the user need not.

4. Why is MEO the orbit for navigation? Because a navigation fix needs at least four satellites in view at once, for three position coordinates and the receiver's clock error, and it needs them with good geometry everywhere on earth at all times. One satellite at 20,200 kilometres sees 29.9 per cent of the earth, so a constellation of 24 puts 7.2 in view on average, a comfortable margin over four. A low orbit would need hundreds of satellites to give the same simultaneous visibility, and a geostationary ring could not serve high latitudes at all.

5. Why is Doppler shift a problem in LEO and not in GEO? Because the shift is proportional to how fast the distance is changing. A satellite at 780 kilometres closes on a user at up to about 6,548 metres a second, which shifts a 1,620 MHz carrier by 35.4 kHz and reverses sign as the satellite passes overhead, so the receiver must search a wide band to acquire the carrier and then track it continuously. A geostationary satellite does not move relative to a fixed dish, so the shift is effectively zero and the receiver can be tuned and left.

munotes.in820

LEO and MEO

6. Compare Iridium and Globalstar. Both are big LEO voice systems, but they make opposite choices. Iridium flies 66 satellites in near-polar orbits at about 780 kilometres and carries traffic from satellite to satellite, so a call can be routed across the constellation and delivered far from where it started, including over oceans and the poles where no gateway exists; the price is complex satellites and a routing problem in the sky. Globalstar flies 48 satellites at about 1,414 kilometres in inclined orbits with no links between them, so every call is relayed straight down to a gateway; the satellites are simpler and cheaper, but a call only works where a gateway shares the footprint, and the high latitudes are not served.

7. The two costs of altitude cross at about 2,000 km. Why does nothing fly there? Because the number of satellites needed falls with altitude while the delay rises with it, and around 2,000 kilometres the two are equal, which would look like the natural compromise. But that is also where the inner radiation belt begins, and a satellite that spends its life inside it needs heavy shielding and radiation-hardened parts. The belt pushes designs to either side of the crossing, which is why the orbits in use form three separate bands rather than a smooth range.

Contents This chapter on its own page

munotes.in821

Chapter One Hundred Seven

Routing in Satellite Systems

Syllabus topic Module 2, "Satellite Systems: Routing"

In one line

Either the satellite is a mirror that must always have a ground station in sight, which is simple and fails over oceans, or the satellites form a network in the sky, which reaches everywhere and turns routing into a problem of a graph whose edges move.

The two designs

A bent pipe, also called a transparent transponder, receives on the uplink, shifts and amplifies, and retransmits on the downlink. It performs no switching and holds no routing table. Everything it receives goes straight back down inside its own footprint, so the sender and a gateway must be under the same satellite at the same time. All the intelligence sits on the ground, where it is cheap to build, easy to repair and easy to upgrade.

The other design gives each satellite links to its neighbours, called inter-satellite links, and a switch on board. Traffic can then travel across the constellation and come down wherever it is wanted, and the satellite has become a router.

Every other difference follows from that one choice.

The case a bent pipe cannot serve

The program puts a ship in the middle of the Pacific and tries to call Mumbai.

With satellites at 780 km and a ten degree minimum elevation, the satellite above the ship can reach the ground only within 18.7 degrees of its sub-satellite point. The program measures the angle from that satellite to six large gateways, and every one is out of reach: Sydney is 32.0 degrees away, Honolulu 42.9, Los Angeles 76.6, Mumbai 112.0. Gateways in view: none.

So with a bent-pipe constellation the call cannot be placed at all, no matter how good the handset is. The satellite is overhead and working; there is simply nowhere for it to put the traffic. That is the practical meaning of Globalstar needing a gateway in the same footprint as the user, and it is why large parts of the oceans and the high latitudes are not served by that design.

With links between the satellites the same call goes through. The program routes it: up 2,315 km to the satellite overhead, across eight satellites in seven hops, and down 2,096 km to Mumbai, for a total of 25,728 km and 85.8 ms one way. A signal that could travel the surface distance of 12,860 km at the speed of light would take 42.9 ms, so the constellation costs twice the theoretical minimum and delivers a call that otherwise would not exist.

The constellation as a grid

A lattice of six columns and eleven rows of small circles, with faint lines joining each circle to the one above, below and beside it. One route is drawn in heavy line and filled dots: it runs along the bottom row from the left column to the fifth column, then straight up that column to the top. Labels mark the start at the bottom left and the end at the top. Text at the right says the seam, these planes run opposite ways, so nothing crosses; that a link along a plane is 4,033 km always; and that a link across planes is 575 to 3,705 km. A note says that to a router the constellation is a grid, and the grid turns under the traffic, so the same two callers are joined by a different path a few minutes later

Figure 107.1 A polar constellation is a grid, and a route across it is a path in that grid

A near-polar constellation has a shape a router can work with. Each satellite talks to the one ahead of it and behind it in its own plane, and to the one beside it in each neighbouring plane: four links, and the whole constellation is a lattice. Routing across it is then a shortest-path problem, which is what the program solves.

munotes.in822

Routing in Satellite Systems

The two kinds of link are not alike, and the program measures the difference.

Along a plane the satellites keep station with each other, so the distance never changes: 4,033 km, always, and the link is steady. It can be set up once and left.

Across planes the distance depends on where in the orbit the pair happens to be. The program prints it: 3,705 km at the equator, 3,120 km at 33 degrees, 2,433 km at 49 degrees, 1,554 km at 65 degrees, and only 575 km at 81 degrees. The planes converge on the poles, so near them the neighbour is almost overhead and the antenna would have to swing through a large angle very quickly. A real system switches the cross-links off in the polar regions rather than chase the target.

The seam. In a polar constellation the planes are spread over 180 degrees of longitude rather than 360, because each plane covers both sides of the earth. Where the first plane meets the last, the two run in opposite directions, and neighbouring satellites pass each other at up to twice the orbital speed. The program computes what that means for a radio link: a closing speed of 14,924 metres a second, which at 23 GHz is a Doppler shift of up to 1.15 MHz, and at 60 GHz up to 2.99 MHz. A beam cannot be held on a target crossing that fast for long, so the seam is where links are given up and traffic goes around.

Routing here is not ad hoc routing

It is tempting to treat a constellation as a mobile network and reach for the protocols of [Routing in Ad Hoc Networks: Proactive, Reactive and Hybrid]. It is not the same problem, and the difference is worth stating because it makes satellite routing easier, not harder.

In an ad hoc network the topology changes unpredictably: a node moves, a battery dies, a link fades, and no one knew it was coming. In a constellation the topology changes on a timetable. The satellites follow orbits computed years ahead, so where every satellite will be and which links will exist can be known in advance and the routing tables can be precomputed and switched on a schedule. What [Routing Tables and What Happens When the Topology Changes] treats as an emergency is here a diary entry.

munotes.in823

Routing in Satellite Systems

What remains genuinely hard is the traffic, not the topology: load shifts as the earth turns under the constellation, a satellite over a city carries far more than one over an ocean, and a route in use must be replaced while the call continues. That replacement is a handover of the path rather than of the radio link, and [Localization and Handover in Satellite Systems] gives it its name.

Cable, and an honest result

The program sets the constellation against submarine cable on three routes, using the same shortest-path search. The results do not all go the way the textbook slogan expects.

Mumbai to Singapore: 29.3 ms by cable, 36.9 ms through the constellation. The cable wins, and comfortably.

Mumbai to London: 54.0 ms by cable, 37.0 ms through the constellation.

Mumbai to New York: 94.2 ms by cable, 56.2 ms through the constellation.

On the long routes the constellation is faster, and the reason is physical rather than clever. Light in glass travels at about two thirds of its speed in vacuum, and a cable is not laid in a straight line, so a cable route costs roughly twice the time its straight-line distance suggests. A path through space is longer in kilometres but every kilometre is at full speed. This is why low constellations are taken seriously for long-haul traffic where a few tens of milliseconds are worth money, and it is the opposite of the case for [GEO: The Geostationary Orbit], where the distance is so great that the cable always wins.

Two cautions belong with those numbers. They are one instant of the constellation: the hop count and the path change as it turns, and the short Singapore route is penalised by ground legs of over 2,000 km to whichever satellite happens to be nearest. And they are propagation only: switching, queueing and processing at each hop are real and are not counted.

Why the links are hard, and why some systems do without

The program's last section lists the difficulties. Each satellite travels at 7,462 metres a second. Two in the same plane keep station, so that link is steady. Two in neighbouring planes cross at an angle that changes through the orbit, and at the seam they close at up to 14,924 metres a second. The antenna must hold a narrow beam on a target thousands of kilometres away that is moving fast, and the frequency it receives slides by up to a megahertz or more, so the receiver must track as well as point. The mechanism is the one in [Multipath, Fading and the Doppler Effect], applied to a target that is itself in orbit.

munotes.in824

Routing in Satellite Systems

All of that is weight, power, cost and something more to fail, on a satellite that must be replaced when it does. A designer who does not need to serve oceans and poles can leave it all out, keep the satellite simple, and put the complexity in gateways on the ground where an engineer can reach it. That is a defensible choice, and it is the one Globalstar made.

The routing, computed

# Routing a call across a low constellation, and what each way of doing it costs.
import heapq
import math

R, C = 6378137.0, 299792458.0
ALT = 780e3
r = R + ALT
PLANES, PER_PLANE, INCL = 6, 11, 86.4       # a near-polar constellation

def sat_xyz(plane, k, u0=0.0):
    """Position of one satellite: plane sets the ascending node, k the place in it."""
    raan = math.radians(plane * 180.0 / PLANES)
    u = math.radians(k * 360.0 / PER_PLANE + u0)
    i = math.radians(INCL)
    # in the orbital plane, then rotated by inclination and by the node
    x, y, z = r * math.cos(u), r * math.sin(u) * math.cos(i), r * math.sin(u) * math.sin(i)
    return (x * math.cos(raan) - y * math.sin(raan),
            x * math.sin(raan) + y * math.cos(raan), z)

def ground_xyz(lat_deg, lon_deg):
    lat, lon = math.radians(lat_deg), math.radians(lon_deg)
    return (R * math.cos(lat) * math.cos(lon), R * math.cos(lat) * math.sin(lon),
            R * math.sin(lat))

def dist(a, b):
    return math.sqrt(sum((a[i] - b[i]) ** 2 for i in range(3)))

SATS = {(p, k): sat_xyz(p, k) for p in range(PLANES) for k in range(PER_PLANE)}

# 1. How far apart the neighbours are. Along a plane the spacing is fixed;
#    across planes it depends where in the orbit the satellites are.
print("How far apart two satellites that talk to each other are:")
along = dist(SATS[(0, 0)], SATS[(0, 1)])
print("  along one plane, always: %6.0f km" % (along / 1000))
print("  across two planes, and it depends where in the orbit they are:")
rows = []
for k in range(PER_PLANE):
    a, b = SATS[(0, k)], SATS[(1, k)]
    rows.append((abs(math.degrees(math.asin(a[2] / r))), dist(a, b)))
for lat, d in sorted(rows)[::2]:
    print("     %5.1f degrees from the equator: %6.0f km" % (lat, d / 1000))
print("  the cross-links shorten towards the poles, where the planes converge, so a")
print("  real system switches them off there rather than chase a target overhead.")

# 2. Route a call with links between satellites. Dijkstra over the grid: each
#    satellite talks to the one ahead and behind in its plane and to the one
#    beside it in each neighbouring plane.
def neighbours(node):
    p, k = node
    out = [(p, (k + 1) % PER_PLANE), (p, (k - 1) % PER_PLANE)]
    if p > 0:
        out.append((p - 1, k))
    if p < PLANES - 1:
        out.append((p + 1, k))
    return out

def route(src, dst):
    best = {src: 0.0}
    prev = {}
    seen = set()
    q = [(0.0, src)]
    while q:
        d, node = heapq.heappop(q)
        if node in seen:
            continue
        seen.add(node)
        if node == dst:
            break
        for nb in sorted(neighbours(node)):
            nd = d + dist(SATS[node], SATS[nb])
            if nd < best.get(nb, float('inf')) - 1e-6:
                best[nb], prev[nb] = nd, node
                heapq.heappush(q, (nd, nb))
    path, node = [dst], dst
    while node in prev:
        node = prev[node]
        path.append(node)
    return list(reversed(path)), best[dst]

def nearest(point):
    return min(sorted(SATS), key=lambda n: dist(SATS[n], point))

ship = ground_xyz(0.0, -170.0)              # a ship in the mid Pacific
mumbai = ground_xyz(19.076, 72.877)
a, b = nearest(ship), nearest(mumbai)
path, air = route(a, b)
up, down = dist(ship, SATS[a]), dist(SATS[b], mumbai)
total = up + air + down
great_circle = R * math.acos(sum(ship[i] * mumbai[i] for i in range(3)) / (R * R))

print()
print("A call from a ship in the mid Pacific to Mumbai, over links between satellites:")
print("  the two places are %6.0f km apart across the earth's surface" % (great_circle / 1000))
print("  it goes up %5.0f km, across %2d satellites in %2d hops, and down %5.0f km"
      % (up / 1000, len(path), len(path) - 1, down / 1000))
print("  the whole path is %6.0f km, so one way takes %5.1f ms"
      % (total / 1000, total / C * 1000))
print("  against %5.1f ms if a signal could travel the surface distance at light speed"
      % (great_circle / C * 1000))

# 3. The same call without links between satellites. Every hop must find a
#    gateway, and in the middle of an ocean there is none in view.
print()
print("The same call with no links between satellites, so every hop goes to a gateway:")
GATEWAYS = (("Mumbai", 19.076, 72.877), ("Singapore", 1.35, 103.82),
            ("Sydney", -33.87, 151.21), ("Honolulu", 21.31, -157.86),
            ("Los Angeles", 34.05, -118.24), ("Fucino", 41.98, 13.60))
elev_min = math.radians(10.0)
gamma_max = math.acos(R / r * math.cos(elev_min)) - elev_min
print("  the satellite above the ship can reach a gateway only inside %.1f degrees:"
      % math.degrees(gamma_max))
reachable = []
for name, lat, lon in GATEWAYS:
    g = ground_xyz(lat, lon)
    cosang = sum(g[i] * SATS[a][i] for i in range(3)) / (R * r)
    ang = math.degrees(math.acos(max(-1.0, min(1.0, cosang))))
    mark = "in view" if ang <= math.degrees(gamma_max) else "out of reach"
    print("    %-12s %6.1f degrees away, %s" % (name, ang, mark))
    if mark == "in view":
        reachable.append(name)
print("  gateways in view: %s" % (", ".join(reachable) if reachable else "none"))
print("  with no gateway under the satellite the call cannot be placed at all, however")
print("  strong the terminal: that is the price of leaving the links out.")

# 4. Cable against constellation, on the same three routes. Light in glass is
#    two thirds of its vacuum speed and a cable is not laid in a straight line.
print()
print("Cable against a constellation with links between its satellites:")
FIBRE_FACTOR, GLASS = 1.5, 2.0 / 3.0
ROUTES = (("Mumbai to Singapore", (19.076, 72.877), (1.35, 103.82), 3910),
          ("Mumbai to London", (19.076, 72.877), (51.51, -0.13), 7200),
          ("Mumbai to New York", (19.076, 72.877), (40.71, -74.01), 12550))
print("  route                 by cable   through the constellation")
for name, p1, p2, km in ROUTES:
    cable = km * 1000 * FIBRE_FACTOR / (C * GLASS)
    g1, g2 = ground_xyz(*p1), ground_xyz(*p2)
    n1, n2 = nearest(g1), nearest(g2)
    hops, air_m = route(n1, n2)
    space = (dist(g1, SATS[n1]) + air_m + dist(SATS[n2], g2)) / C
    print("  %-22s %5.1f ms   %5.1f ms over %d hops, %6.0f km through space"
          % (name, cable * 1000, space * 1000, len(hops) - 1, space * C / 1000))
print("  on a short route the cable wins easily. On a long one it does not: light in")
print("  glass is two thirds of its speed in vacuum and the cable is not laid straight,")
print("  so a long path through space can arrive first. These are one instant of the")
print("  constellation: the hop count changes as it turns.")

# 5. Why the links are hard: the satellites move relative to each other.
print()
print("Why a link between two satellites is hard to hold:")
v = math.sqrt(3.986004418e14 / r)
print("  each satellite travels at %.0f m/s" % v)
print("  two in the same plane keep station, so that link is steady")
print("  two in neighbouring planes cross at an angle, and at the seam where the")
print("  constellation closes, two planes run opposite ways: up to %.0f m/s apart" % (2 * v))
for f_ghz in (23.0, 60.0):
    print("    at %4.1f GHz that is a Doppler shift of up to %5.2f MHz"
          % (f_ghz, 2 * v / C * f_ghz * 1000))
print("  and the antenna must hold a beam on a target thousands of km away that is")
print("  moving: that is why the seam links are the ones a system gives up first.")
munotes.in825

Routing in Satellite Systems

How far apart two satellites that talk to each other are:
  along one plane, always:   4033 km
  across two planes, and it depends where in the orbit they are:
       0.0 degrees from the equator:   3705 km
      16.3 degrees from the equator:   3556 km
      32.7 degrees from the equator:   3120 km
      49.0 degrees from the equator:   2433 km
      65.2 degrees from the equator:   1554 km
      81.1 degrees from the equator:    575 km
  the cross-links shorten towards the poles, where the planes converge, so a
  real system switches them off there rather than chase a target overhead.

A call from a ship in the mid Pacific to Mumbai, over links between satellites:
  the two places are  12860 km apart across the earth's surface
  it goes up  2315 km, across  8 satellites in  7 hops, and down  2096 km
  the whole path is  25728 km, so one way takes  85.8 ms
  against  42.9 ms if a signal could travel the surface distance at light speed

The same call with no links between satellites, so every hop goes to a gateway:
  the satellite above the ship can reach a gateway only inside 18.7 degrees:
    Mumbai        112.0 degrees away, out of reach
    Singapore      78.2 degrees away, out of reach
    Sydney         32.0 degrees away, out of reach
    Honolulu       42.9 degrees away, out of reach
    Los Angeles    76.6 degrees away, out of reach
    Fucino        152.2 degrees away, out of reach
  gateways in view: none
  with no gateway under the satellite the call cannot be placed at all, however
  strong the terminal: that is the price of leaving the links out.

Cable against a constellation with links between its satellites:
  route                 by cable   through the constellation
  Mumbai to Singapore     29.3 ms    36.9 ms over 2 hops,  11059 km through space
  Mumbai to London        54.0 ms    37.0 ms over 3 hops,  11090 km through space
  Mumbai to New York      94.2 ms    56.2 ms over 5 hops,  16839 km through space
  on a short route the cable wins easily. On a long one it does not: light in
  glass is two thirds of its speed in vacuum and the cable is not laid straight,
  so a long path through space can arrive first. These are one instant of the
  constellation: the hop count changes as it turns.

Why a link between two satellites is hard to hold:
  each satellite travels at 7462 m/s
  two in the same plane keep station, so that link is steady
  two in neighbouring planes cross at an angle, and at the seam where the
  constellation closes, two planes run opposite ways: up to 14924 m/s apart
    at 23.0 GHz that is a Doppler shift of up to  1.15 MHz
    at 60.0 GHz that is a Doppler shift of up to  2.99 MHz
  and the antenna must hold a beam on a target thousands of km away that is
  moving: that is why the seam links are the ones a system gives up first.
munotes.in826

Routing in Satellite Systems

Distinctions

Bent pipe, routing through the groundInter-satellite links
The satellite isA mirror: receive, shift, amplify, retransmitA router with a switch on board
NeedsA gateway in the same footprint as the userNothing under it
Over an oceanThe call cannot be placedThe call is routed across the sky
Satellite complexityLowHigh: switching, pointing, tracking
Where the intelligence isOn the ground, easy to upgradeIn orbit, fixed at launch
DelayDepends on the terrestrial pathComputed across the constellation
ExampleGlobalstarIridium
munotes.in827

Routing in Satellite Systems

Link along a planeLink across planes
Distance4,033 km, unchanging575 to 3,705 km with latitude
GeometryFixed: the pair keep stationChanges continuously
Near the polesUnaffectedPlanes converge, so usually switched off
At the seamNot applicableClosing at up to 14,924 m/s, so given up
DifficultySet up oncePoint, track and retune continuously
munotes.in828

Routing in Satellite Systems

Ad hoc routingConstellation routing
Topology changesUnpredictablyOn a timetable, known years ahead
TablesDiscovered or maintained by protocolCan be precomputed and switched on schedule
The hard partFinding out the topology changedMoving live traffic to the new path

What it does not mean

A bent pipe is not a failure of engineering. It is a deliberate trade of coverage for simplicity, and it puts every part that may need changing on the ground.

Inter-satellite links do not remove ground stations. Traffic still has to enter and leave the network somewhere; the links remove the requirement that it enter and leave under the same satellite.

A constellation is not a mobile ad hoc network. Its motion is known in advance, so its routing is scheduled rather than discovered.

More hops do not automatically mean more delay than a cable. The program's own numbers show a long route arriving sooner through space, because light in glass is slower and cable is not laid straight.

The seam is not a gap in coverage. Both planes cover their ground perfectly well; what is missing is the radio link between them, so traffic goes around rather than across.

The delays computed here are not the delay a user feels. They are propagation only. Switching, queueing, coding and the terrestrial legs at each end all add to them.

Quick revision

  • Two designs: bent pipe, which needs a gateway in the same footprint; and inter-satellite links, which do not.
  • The Pacific test: the satellite reaches the ground only within 18.7 degrees, and no gateway is in view, so a bent-pipe call cannot be placed. With links: up 2,315 km, 7 hops, down 2,096 km, 25,728 km and 85.8 ms one way, against 42.9 ms for the surface distance at light speed.
  • The grid: four links each, two along the plane and two across. Along a plane, 4,033 km always; across planes, 3,705 km at the equator down to 575 km at 81 degrees.
  • The seam: counter-rotating planes closing at 14,924 m/s, a Doppler shift of 1.15 MHz at 23 GHz and 2.99 MHz at 60 GHz, so links there are given up.
  • Cable against constellation: Singapore 29.3 against 36.9 ms, London 54.0 against 37.0, New York 94.2 against 56.2. Long routes are faster through space, because light in glass is two thirds as fast and cable is not straight.
  • Constellation routing is scheduled, not discovered: the topology changes on a timetable, so tables can be precomputed.
  • Iridium has the links, Globalstar does not.
munotes.in829

Routing in Satellite Systems

Test yourself

1. What are the two ways of routing in a satellite system? Either the satellite acts as a bent pipe, receiving on the uplink and retransmitting immediately on the downlink with no switching, so that traffic must be handed to a ground station inside the same footprint and carried onward by the terrestrial network; or the satellites carry links to each other and a switch on board, so that traffic can be routed across the constellation in the sky and brought down wherever it is wanted.

2. Why can a bent-pipe constellation fail to place a call at all? Because the satellite can only reach the ground inside its own footprint. From 780 kilometres with a ten degree minimum elevation that is 18.7 degrees of earth-central angle, and in the middle of an ocean no gateway lies within it: the chapter's test finds the nearest at 32 degrees and the rest much further. The satellite is overhead and working, but there is nowhere to put the traffic, so the call fails however good the handset is.

3. Describe the link structure of a near-polar constellation. Each satellite keeps four links: one ahead and one behind in its own plane, and one to each side in the neighbouring planes, which makes the whole constellation a lattice that a shortest-path algorithm can route across. Links along a plane are a constant 4,033 kilometres because the pair keep station. Links across planes vary from 3,705 kilometres at the equator to 575 kilometres at 81 degrees of latitude, because the planes converge on the poles, so they are usually switched off in the polar regions.

4. What is the seam, and why are links not carried across it? In a polar constellation the orbital planes are spread over 180 degrees rather than 360, so where the first plane meets the last the two run in opposite directions. Satellites there pass each other at up to twice the orbital speed, which the chapter computes as 14,924 metres a second, giving a Doppler shift of up to 1.15 MHz at 23 GHz and 2.99 MHz at 60 GHz, and the relative motion is far too fast to hold a narrow beam on. Links across the seam are therefore not maintained and traffic is routed around it.

munotes.in830

Routing in Satellite Systems

5. How does routing in a constellation differ from routing in an ad hoc network? In an ad hoc network the topology changes without warning, so a protocol has to discover the change and repair the routes after the fact. In a constellation the satellites follow orbits known years in advance, so which links will exist at any moment can be computed beforehand and the routing tables can be precomputed and switched on a schedule. The topology is not the hard part; the hard part is moving live traffic onto a new path while a call is in progress.

6. Can a satellite path be faster than a submarine cable? Yes, on a long route. Light travels in glass at about two thirds of its vacuum speed, and a cable is not laid in a straight line, so a cable costs roughly twice the time its straight-line distance suggests. The chapter computes Mumbai to New York at 94.2 milliseconds by cable and 56.2 through a constellation with inter-satellite links, and Mumbai to London at 54.0 against 37.0. On the short Mumbai to Singapore route the cable still wins, 29.3 against 36.9, because the ground legs up to and down from the satellites dominate a short path.

7. Why might a designer leave inter-satellite links out? Because they are expensive in every currency a satellite has. They need a switch on board, narrow-beam antennas that point and track a target thousands of kilometres away and moving at kilometres a second, and receivers that follow a carrier sliding by up to a megahertz or more. All of that is mass, power and cost, and once launched it cannot be repaired or upgraded. A system that does not need to serve oceans or high latitudes can leave it out, keep the satellites simple and cheap, and put the complexity in gateways on the ground where it can be reached.

Contents This chapter on its own page

munotes.in831

Chapter One Hundred Eight

Localization and Handover in Satellite Systems

Syllabus topic Module 2, "Satellite Systems: Localization, Handover"

In one line

A satellite network has to answer the same question a mobile network answers, which satellite should this call be sent to, but the answer goes stale in minutes because the base stations are in orbit: so it keeps an extra register of where the satellites are, and it hands calls over constantly even when nobody has moved.

Finding a user

The problem is the one [Localization and Calling in GSM] sets out. A call arrives for a number. The network has no idea, from the number alone, where in the world the handset is or which piece of equipment can reach it. It must look that up.

A satellite system keeps the same two registers a mobile network keeps.

The home location register holds each subscriber permanently: the subscription, the services allowed, and a pointer to wherever the subscriber currently is. There is one home register per subscriber and it never changes.

The visitor location register holds the subscribers currently being served in one part of the network, with the detail needed to reach them. A user registers, the visitor register takes them in, and it tells the home register where to send calls.

Then a satellite system needs a third thing the ground has no use for. A terrestrial network's base stations are bolted to the ground: knowing the cell is knowing the place. Here the base stations are in orbit at 7,462 metres a second, so knowing which satellite is serving a user tells you nothing unless you also know where that satellite is now. The system therefore keeps a register of the current positions of all the satellites and of which satellite currently serves each user, sometimes called the satellite user mapping register. It is not a copy of the visitor register: one says who is where, the other says which piece of moving equipment can reach them at this moment.

Registering. A terminal listens, finds a satellite, and registers through it to a gateway. The gateway enters it in the visitor register, tells the home register where it is, and the satellite register records which satellite is carrying it.

Being called. A call for the number reaches the home register, which says which gateway is serving that user. The gateway asks the satellite register which satellite has the user now and where that satellite is, and the call is routed up through it, across the constellation if it has links, and down.

The whole procedure is the terrestrial one plus one lookup, and that lookup has to be kept fresh minute by minute, which is the real cost.

Handover, and why there is so much of it

Four dashed panels, each a small sketch. The first, intra-satellite, shows a satellite above three dashed oval beams laid along a ground line with a user standing in one and an arrow labelled beams move. The second, inter-satellite, shows a user on the ground with a dashed line to a setting satellite on the left and a solid line to a rising satellite on the right. The third, gateway, shows a user with a solid line to one satellite, which has a dashed line to one ground station box and a solid line to another. The fourth, inter-system, shows a user with a dashed line to a satellite and a solid line to a mast. Notes under each panel explain that the beams sweep past a user who has not moved, that one satellite sets and the next takes the call over, that the radio link does not change but the ground station does, and that traffic goes to the ground network where there is one. A line at the top says not one of the four needs the user to have moved, and a line at the bottom says a solid line is the link in use and a dashed line the one being given up

Figure 108.1 The four handovers, and not one of them needs the user to have moved

munotes.in832

Localization and Handover in Satellite Systems

In a terrestrial network handover means the user moved. Here it usually does not. The program makes the point with arithmetic: the satellite's point on the ground runs at 6.65 kilometres every second, while a car on a motorway adds 0.0278 kilometres a second. The user's own movement is a rounding error. Everything that follows is caused by the network moving.

Intra-satellite handover

A satellite's footprint is not one cell. It is divided into spot beams, each a narrow beam from the satellite's antenna, and each beam is a cell in the sense the terrestrial world means: its own frequencies or codes, reused in beams far enough away.

The beams are fixed to the satellite, so they sweep across the ground with it. The program computes how fast. With the whole footprint as one cell a user is inside it for 624.8 seconds. Divide it into 16 beams and each is about 1,039 km across, crossed in 156.2 seconds. Into 48 beams and each is about 600 km across, crossed in 90.2 seconds. Into 96 and it is 63.8 seconds.

So a user sitting perfectly still in a chair is handed from beam to beam about every minute and a half. That is intra-satellite handover, and it is the most frequent thing the system does. It is also the easiest: the same satellite is involved throughout, so the decision and the switch are made on board.

Inter-satellite handover

Eventually the satellite sets. From [LEO and MEO], a satellite at 780 km passing straight overhead is above ten degrees for 10.4 minutes, and the program adds what happens when the pass is not overhead: 9.1 minutes for a user half way to the edge of the swath, and only 4.6 minutes near the edge. Ten minutes is the best case, not the usual one.

At that point the call must move to the next satellite, which is inter-satellite handover. If the constellation has links between its satellites the call may also have to be rerouted across the grid of [Routing in Satellite Systems], so the path changes as well as the radio link.

Gateway handover

A third kind has no counterpart on the ground at all. The user's satellite has not set and the radio link is perfect, but the gateway that connects that satellite to the fixed network is itself passing out of the satellite's footprint, because a gateway is just another point on the ground and leaves the footprint at the same rate a user does: at best 10.4 minutes.

The call must then be moved to another gateway while the radio link stays exactly as it was. The user notices nothing. This is gateway handover, and it exists because in a satellite system the connection to the fixed network is itself made over a radio link that comes and goes.

munotes.in833

Localization and Handover in Satellite Systems

Inter-system handover

The last kind is between the satellite system and a terrestrial one, for a dual-mode handset that can use either.

The program shows which way the preference runs, and why it is not a close call. At 1,600 MHz, reaching the satellite at 780 km costs 154.4 dB. Reaching a mast at the edge of a 10 km cell costs 116.5 dB, a 3 km cell 106.1 dB, and a 500 m cell 90.5 dB. The terrestrial link is 38 to 64 dB cheaper, which is a factor of between six thousand and two and a half million.

So a dual-mode handset uses the ground network wherever there is one, and falls back to the satellite only where there is not. Handover to the satellite happens when the terrestrial signal is lost, and back the moment it returns. Satellite capacity is scarce and expensive, and this is how it is saved for the places that have no alternative.

A ten minute call

Put the four together for a caller sitting still, as the program does. With 48 beams a satellite: about 6.7 beam handovers and 1.0 satellite handovers in ten minutes, plus a gateway handover if the gateway happens to be leaving, plus an inter-system handover if the caller walks indoors near a mast.

Roughly eight handovers in a ten minute call, none of them because anyone moved. A terrestrial network doing that would be considered broken. Here it is normal, and the entire design of the signalling, the registers and the routing exists to make it invisible.

Handover, computed

# Handover in a low constellation: how often, of which kind, and why.
import math

GM, R, C = 3.986004418e14, 6378137.0, 299792458.0
ALT = 780e3
r = R + ALT
ELEV = 10.0

T = 2 * math.pi * math.sqrt(r ** 3 / GM)
v_orbit = math.sqrt(GM / r)
v_track = v_orbit * R / r                    # the sub-satellite point's speed on the ground
gamma = math.acos(R / r * math.cos(math.radians(ELEV))) - math.radians(ELEV)
foot_km = gamma * R / 1000                   # footprint radius on the ground

print("The satellite, and the patch of ground it owns:")
print("  period %.1f min, orbital speed %.0f m/s" % (T / 60, v_orbit))
print("  its point on the ground runs at %.0f m/s, which is %.2f km every second"
      % (v_track, v_track / 1000))
print("  at a %.0f degree minimum elevation the footprint reaches %.0f km from that point"
      % (ELEV, foot_km))

# 1. Intra-satellite handover. The footprint is not one cell: it is divided
#    into spot beams, and a user crosses them one after another.
print()
print("Inside one satellite: crossing its spot beams")
area = math.pi * (foot_km * 1000) ** 2
for beams in (1, 16, 48, 96):
    cell_area = area / beams
    cell_r = math.sqrt(cell_area / math.pi)
    cross_s = 2 * cell_r / v_track
    print("  %3d %-6s each covers %8.0f km2, about %4.0f km across, crossed in %5.1f s"
          % (beams, "beam:" if beams == 1 else "beams:", cell_area / 1e6,
             2 * cell_r / 1000, cross_s))
print("  the beams are fixed to the satellite, so they sweep the ground with it: a user")
print("  standing still is handed from beam to beam, and that is intra-satellite handover.")

# 2. Inter-satellite handover. The satellite itself sets.
print()
print("Between satellites: how long one is usable")
best = T * (2 * math.degrees(gamma) / 360.0)
print("  passing straight overhead, it is above %.0f degrees for %.1f min" % (ELEV, best / 60))
for offset_frac, label in ((0.0, "straight overhead"), (0.5, "half way to the edge"),
                           (0.9, "near the edge of the swath")):
    off = gamma * offset_frac                     # how far off the track the user stands
    half = math.acos(math.cos(gamma) / math.cos(off))
    print("  %-26s usable for %4.1f min" % (label, T * (2 * math.degrees(half) / 360.0) / 60))
print("  a pass that is not overhead is shorter, so the figure above is the best case.")

# 3. What a ten minute call actually costs in handovers.
print()
print("A ten minute call, standing perfectly still:")
CALL_MIN = 10.0
for beams in (16, 48, 96):
    cell_r = math.sqrt(area / beams / math.pi)
    cross_min = (2 * cell_r / v_track) / 60
    print("  with %2d beams a satellite: about %4.1f beam handovers and %.1f satellite handovers"
          % (beams, CALL_MIN / cross_min, CALL_MIN / (best / 60)))
print("  none of them because the caller moved. A car at 100 km/h adds %.4f km/s to a"
      % (100 / 3600.0))
print("  ground track already running at %.2f km/s, which changes almost nothing."
      % (v_track / 1000))

# 4. Gateway handover. The gateway must stay in view of the same satellite.
print()
print("Gateway handover: how long one ground station stays with one satellite")
print("  a gateway enters the footprint and leaves it at the same rate as any point:")
print("  at best %.1f min, the same figure as a user's pass" % (best / 60))
print("  so a call held on one satellite may still have to be moved to another gateway,")
print("  and the user notices nothing: the radio link has not changed at all.")

# 5. Inter-system handover, and why it is preferred in one direction.
print()
print("Inter-system handover: satellite against terrestrial, at 1600 MHz")
def fsl(f_mhz, d_km):
    return 32.44 + 20 * math.log10(f_mhz) + 20 * math.log10(d_km)
sat_db = fsl(1600, ALT / 1000)
for cell_km in (0.5, 3.0, 10.0):
    print("  a terrestrial cell of %4.1f km costs %5.1f dB, against %5.1f dB to the satellite"
          % (cell_km, fsl(1600, cell_km), sat_db))
print("  the terrestrial link is %.0f to %.0f dB cheaper, so a dual-mode handset uses the"
      % (sat_db - fsl(1600, 10.0), sat_db - fsl(1600, 0.5)))
print("  ground network wherever there is one, and the satellite only when there is not.")

# 6. Where the network has to look to find a user.
print()
print("Finding a user: what the registers have to hold")
sats = 66
print("  the home register holds, as on the ground, the user's subscription and where to ask")
print("  the visitor register holds the users currently being served")
print("  a third register is needed that the ground network has no use for: the current")
print("  position of all %d satellites, and which one is serving each user" % sats)
print("  it must be updated every time a user is handed on, which the program above puts at")
print("  once every %.1f min per user in a call, and more often between beams" % (best / 60))
munotes.in834

Localization and Handover in Satellite Systems

The satellite, and the patch of ground it owns:
  period 100.5 min, orbital speed 7462 m/s
  its point on the ground runs at 6649 m/s, which is 6.65 km every second
  at a 10 degree minimum elevation the footprint reaches 2077 km from that point

Inside one satellite: crossing its spot beams
    1 beam:  each covers 13552881 km2, about 4154 km across, crossed in 624.8 s
   16 beams: each covers   847055 km2, about 1039 km across, crossed in 156.2 s
   48 beams: each covers   282352 km2, about  600 km across, crossed in  90.2 s
   96 beams: each covers   141176 km2, about  424 km across, crossed in  63.8 s
  the beams are fixed to the satellite, so they sweep the ground with it: a user
  standing still is handed from beam to beam, and that is intra-satellite handover.

Between satellites: how long one is usable
  passing straight overhead, it is above 10 degrees for 10.4 min
  straight overhead          usable for 10.4 min
  half way to the edge       usable for  9.1 min
  near the edge of the swath usable for  4.6 min
  a pass that is not overhead is shorter, so the figure above is the best case.

A ten minute call, standing perfectly still:
  with 16 beams a satellite: about  3.8 beam handovers and 1.0 satellite handovers
  with 48 beams a satellite: about  6.7 beam handovers and 1.0 satellite handovers
  with 96 beams a satellite: about  9.4 beam handovers and 1.0 satellite handovers
  none of them because the caller moved. A car at 100 km/h adds 0.0278 km/s to a
  ground track already running at 6.65 km/s, which changes almost nothing.

Gateway handover: how long one ground station stays with one satellite
  a gateway enters the footprint and leaves it at the same rate as any point:
  at best 10.4 min, the same figure as a user's pass
  so a call held on one satellite may still have to be moved to another gateway,
  and the user notices nothing: the radio link has not changed at all.

Inter-system handover: satellite against terrestrial, at 1600 MHz
  a terrestrial cell of  0.5 km costs  90.5 dB, against 154.4 dB to the satellite
  a terrestrial cell of  3.0 km costs 106.1 dB, against 154.4 dB to the satellite
  a terrestrial cell of 10.0 km costs 116.5 dB, against 154.4 dB to the satellite
  the terrestrial link is 38 to 64 dB cheaper, so a dual-mode handset uses the
  ground network wherever there is one, and the satellite only when there is not.

Finding a user: what the registers have to hold
  the home register holds, as on the ground, the user's subscription and where to ask
  the visitor register holds the users currently being served
  a third register is needed that the ground network has no use for: the current
  position of all 66 satellites, and which one is serving each user
  it must be updated every time a user is handed on, which the program above puts at
  once every 10.4 min per user in a call, and more often between beams
munotes.in835

Localization and Handover in Satellite Systems

Distinctions

HandoverWhat changesWhat staysHow often
Intra-satelliteThe spot beamThe satellite, the gateway, the routeAbout every 90 s with 48 beams
Inter-satelliteThe satellite, and often the routeThe gateway, if it is still in viewAt best every 10.4 min, less off the track
GatewayThe ground station and the fixed-network pathThe satellite and the radio linkAt best every 10.4 min
Inter-systemThe whole network, satellite or terrestrialThe callWhenever terrestrial coverage begins or ends
munotes.in836

Localization and Handover in Satellite Systems

Terrestrial handoverSatellite handover
Caused byThe user movingThe network moving
A stationary userNever hands overHands over about every 90 s
Speeds involvedA car at 0.03 km/sA footprint at 6.65 km/s
PredictableNoYes: the orbits are known in advance
RegisterHoldsWhy
HomeThe subscription, and where to ask for this userOne per subscriber, permanent
VisitorThe users currently served in this part of the networkDetail needed to reach them now
Satellite user mappingWhere every satellite is, and which serves each userThe base stations move, so the answer goes stale in minutes
munotes.in837

Localization and Handover in Satellite Systems

What it does not mean

Handover here does not mean the user moved. Almost none of it is caused by the user; the beams and the satellites sweep past a stationary caller.

A spot beam is not a smaller satellite. It is one beam of one satellite's antenna, and a beam handover is handled entirely on board.

Gateway handover does not interrupt the radio link. The link to the satellite is untouched; only the ground station and the path beyond it change.

Inter-system handover is not symmetric. Terrestrial is preferred in both directions, because it is tens of decibels cheaper and its capacity is not scarce.

The satellite register is not the visitor register under another name. The visitor register says which users are here; the satellite register says where the moving equipment is and which piece of it currently reaches each user.

Frequent handover is not a fault. It is the unavoidable consequence of putting the base station in an orbit that crosses the sky in ten minutes.

Quick revision

  • Registers: home (subscription and where to ask), visitor (users served here now), and a satellite user mapping register holding the satellites' current positions and each user's serving satellite. The third exists because the base stations move.
  • Calling: home register gives the gateway, the satellite register gives the satellite and its position, the call goes up, across if there are links, and down.
  • Four handovers: intra-satellite (beam to beam), inter-satellite (satellite sets), gateway (ground station leaves the footprint), inter-system (satellite to terrestrial and back).
  • Rates: ground track 6.65 km/s; with 48 beams a cell is about 600 km across and crossed in 90.2 s; a satellite lasts 10.4 min overhead, 9.1 half way out, 4.6 near the edge; a gateway also 10.4 min.
  • A 10 minute call: about 6.7 beam handovers and 1.0 satellite handovers, none caused by the caller.
  • Inter-system preference: satellite 154.4 dB against 116.5 dB for a 10 km cell and 90.5 dB for a 500 m cell, so terrestrial is 38 to 64 dB cheaper and is always preferred.
  • A car adds 0.0278 km/s to a footprint moving at 6.65 km/s: the user's speed is irrelevant.

Test yourself

1. Why does a satellite system need a register that a terrestrial mobile network does not? Because a terrestrial network's base stations are fixed, so knowing which cell serves a user is enough to know how to reach them, and the answer stays true. In a satellite system the base stations are in orbit at over seven kilometres a second, so knowing which satellite serves a user is useless without also knowing where that satellite is at this moment, and both facts change every few minutes. The system therefore keeps a register of the current positions of all the satellites and of which satellite is serving each user, in addition to the home and visitor registers.

munotes.in838

Localization and Handover in Satellite Systems

2. How is a call delivered to a satellite user? The call arrives for the number and reaches the user's home location register, which holds the subscription and a pointer to the gateway currently serving that user. That gateway consults the register of satellite positions and serving satellites to find which satellite has the user now and where it is, and the call is then routed up to that satellite, across the constellation if it has links between its satellites, and down to the terminal.

3. Name and explain the four handovers. Intra-satellite handover moves the call from one spot beam of a satellite to the next, because the beams are fixed to the satellite and sweep across the ground with it. Inter-satellite handover moves the call to the next satellite when the current one sets. Gateway handover moves the connection to a different ground station because the gateway has left the satellite's footprint, while the radio link to the user is untouched. Inter-system handover moves the call between the satellite network and a terrestrial one, in either direction, as terrestrial coverage is lost or regained.

4. Why does a stationary user hand over so often? Because the movement that matters is the network's, not the user's. The satellite's point on the ground travels at 6.65 kilometres every second, so with 48 spot beams a cell about 600 kilometres across is crossed in 90.2 seconds, and the satellite itself is usable for at most 10.4 minutes. A car on a motorway adds 0.0278 kilometres a second to that, which changes nothing. A ten minute call from an armchair therefore takes about 6.7 beam handovers and one satellite handover.

5. What is gateway handover, and why has it no terrestrial counterpart? It is moving a call from one ground station to another because the ground station has passed out of the serving satellite's footprint, even though the user's radio link is unchanged and perfectly good. It has no terrestrial counterpart because on the ground the base station's connection to the fixed network is a cable that does not move. In a satellite system that connection is itself a radio link to a point on a turning earth, and a gateway leaves the footprint at the same rate a user does, at best every 10.4 minutes.

6. In inter-system handover, which network is preferred, and why? The terrestrial network, in both directions. At 1,600 MHz a satellite at 780 kilometres costs 154.4 dB of free space loss, while a mast at the edge of a ten kilometre cell costs 116.5 dB and one at the edge of a 500 metre cell costs 90.5 dB, so the terrestrial link is between 38 and 64 dB cheaper, a factor of thousands to millions. The handset therefore uses the ground network wherever one exists and falls back to the satellite only where none does, which also saves scarce and expensive satellite capacity for the places that have no alternative.

Contents This chapter on its own page

munotes.in839

Chapter One Hundred Nine

Broadcast Systems: Cyclic Repetition, DAB and DVB

Syllabus topic Module 2, "Broadcast Systems"

In one line

A broadcaster who cannot be asked for anything must send everything again and again, so the whole art of broadcast data is deciding how often each item comes round; and the whole art of broadcast radio and television is spending bits in advance on error correction and on a guard interval, because neither a retransmission nor a question is possible.

The asymmetry, and what it forbids

In every other system in this book the receiver can talk back. It acknowledges, it asks again, it complains about congestion, it registers, it hands over. A broadcast system has none of that. One transmitter sends; an unknown number of receivers listen; nothing comes back.

Four consequences follow at once.

A receiver cannot ask. If it wants the item that went out a minute ago, its only option is to wait for the next time that item is sent.

A receiver cannot acknowledge, so nothing can be retransmitted on request. Error control must be forward: enough redundancy is added in advance to fix errors that have not happened yet, because there will be no second chance. This is why broadcast standards spend so much of their capacity on coding.

Receivers arrive at random. Somebody switches on in the middle of everything. A system whose data only makes sense from the beginning would be useless.

The transmitter cannot adapt to one receiver. It is sending to a city, not a terminal, so it must set its modulation and coding for the worst receiver it intends to serve, not the best.

Cyclic repetition

The answer to the first and third of those is the same: send everything over and over, on a cycle, so that a receiver that wants an item, or arrives late, or loses one to interference, only has to wait for it to come round. The arrangement is called a carousel or a broadcast disk, and it turns a one-way channel into something a receiver can use as if it were storage.

The cost is waiting. If an item goes out every g slots and a receiver arrives at a random moment, it waits g/2 on average and g at worst.

How often should each item be sent?

This is the real design question, and the program answers it for a carousel of 20 items with the usual shape of popularity, a few items wanted often and a long tail wanted rarely.

Two rows of lettered boxes representing slots in a broadcast cycle. The upper row, labelled every item equally often, repeats A B C D E over and over, with every E shaded. The lower row, labelled by the square root of popularity, has more A's and fewer E's, again with E shaded. An arrow points down into both rows at the same slot and is labelled a listener tunes in here. Under the upper row an arrow spans three slots to the next E and is labelled waits 3 slots for E; under the lower row an arrow spans seven slots and is labelled waits 7 slots for E. A note says E is the least wanted item, that the square root schedule gives the shorter mean wait over all listeners, and the longer wait to whoever wanted E

Figure 109.1 Two schedules for the same carousel: the second is better on average and worse for the rare item

Flat, every item equally often, gives a mean wait of 10 slots, which is half the number of items.

In proportion to popularity, so that an item wanted twice as often is sent twice as often, gives a mean wait of 10 slots. Exactly the same.

munotes.in840

Broadcast Systems: Cyclic Repetition, DAB and DVB

That is worth stopping on, because it is the opposite of what intuition says. Sending a popular item twice as often halves its wait, but the slots came from somewhere, and everyone else's wait grows by exactly as much as the popular listeners' wait shrinks. Proportional scheduling buys nothing.

By the square root of popularity gives a mean wait of 8.02 slots, which is 20 per cent better than either. This is the schedule that actually wins, and it wins by being less aggressive than proportional: a popular item gets more slots, but by the square root of the ratio rather than the ratio itself.

The price is paid by the tail. Under the flat schedule the rarest item waits 10 slots; under the square root schedule it waits 16.98, which is 70 per cent longer. Under proportional scheduling it waits 35.98, which is why proportional is the worst of the three in practice even though its mean matches flat.

What that is in seconds

The program converts it. For eight kilobyte items, a 16 kbit/s data channel makes each slot 4.10 seconds, so the mean wait is 32.8 seconds. At 128 kbit/s it is 4.1 seconds, and at 1 Mbit/s half a second.

Thirty seconds to see a page is the difference between a service people use and one they do not, and no amount of protocol design removes it: the wait is set by the rate and the number of items. The only real choices are to send fewer items, to send faster, or to schedule better.

DAB

EN 300 401 states its purpose plainly. It "establishes a broadcasting standard for the Digital Audio Broadcasting (DAB) system", designed for delivery "for mobile, portable and fixed reception from terrestrial transmitters in the Very High Frequency (VHF) frequency bands as well as for distribution through cable networks", and designed to provide "spectrum and power efficient techniques in terrestrial transmitter network planning, known as the Single Frequency Network (SFN) and the gap-filling technique".

What is broadcast. Not a station but an ensemble: one multiplex carrying several services, each made of service components, with the configuration itself broadcast so a receiver can find its way. A receiver tunes to the ensemble, not to a programme.

Two channels inside it. The Fast Information Channel carries the information a receiver needs quickly and repeatedly, above all the multiplex configuration: which service is in which part of the multiplex. The Main Service Channel carries the audio and data.

The frame. The standard fixes every time as a multiple of an elementary period of one divided by 2,048,000 seconds. The transmission frame is 196,608 of them, which the program confirms is 96 ms, and it carries 76 OFDM symbols of 1,536 carriers in transmission mode I. The standard says the frame "consists of consecutive Orthogonal Frequency Division Multiplex (OFDM) symbols", generated through "Differential Quadrature Phase Shift Keying (D-QPSK), frequency interleaving" and multiplexing, which is the arrangement [Advanced Modulation: MSK, GMSK, QPSK, QAM and OFDM] describes.

munotes.in841

Broadcast Systems: Cyclic Repetition, DAB and DVB

The capacity. The standard states that a Common Interleaved Frame "consists of 55 296 bits, grouped into 864 Capacity Units (CU) and is transmitted every 24 ms". The program divides those out: 2.304 Mbit/s gross for the main service channel, and one capacity unit is 64 bits, or 2.667 kbit/s. Services are sold in capacity units, which is why a DAB multiplex is described as so many units rather than so many stations.

The guard interval, and the single frequency network. The useful part of a symbol is 2,048 elementary periods, exactly 1 ms, and the guard interval is 504 of them, 246.09 microseconds. Any echo arriving within the guard interval is absorbed rather than causing interference, and the program turns that time into distance: 73.8 km.

That number is what a single frequency network is built on. Two transmitters up to about that far apart can radiate the same signal on the same frequency, and a receiver between them treats the second as a harmless echo of the first rather than as interference. A whole country can then run on one frequency, which is a spectrum saving no analogue system could approach, and it is why a car radio does not have to retune as it drives.

DVB

EN 300 744 "describes a baseline transmission system for digital terrestrial TeleVision (TV) broadcasting". The family extends the same ideas to satellite and cable, and the satellite member is where this chapter's module meets it.

The payload. DVB carries an MPEG transport stream, and the standard fixes the unit: "The total packet length of the MPEG-2 transport multiplex (MUX) packet is 188 bytes." Everything downstream is built on that 188 byte packet.

Forward error correction, paid in advance. An outer Reed-Solomon code is applied first. The standard's own note explains it: the code "has length 204 bytes, dimension 188 bytes and allows to correct up to 8 random erroneous bytes in a received word of 204 bytes". The program prices it: 16 bytes added to 188, which is 8.5 per cent of the payload, spent before the inner convolutional code has spent anything. An interleaver is placed between them so that a burst of errors is spread across many code words rather than destroying one.

munotes.in842

Broadcast Systems: Cyclic Repetition, DAB and DVB

That is what a system with no return channel has to do. A two-way protocol can send data bare and retransmit the rare packet that fails; a broadcast must buy the insurance for every packet, whether it needs it or not.

Two modes and a choice of guard. DVB-T defines a 2K mode and an 8K mode, and the standard says "a flexible guard interval is specified" which "will enable the system to support different network configurations". The program lays out what each combination buys in an 8 MHz channel. The 8K mode's useful symbol is 896 microseconds, so a quarter guard is 224 microseconds and tolerates transmitters 67.2 km apart; an eighth gives 33.6 km, a sixteenth 16.8 km, a thirty-second 8.4 km. The 2K mode's useful symbol is 224 microseconds, so even its widest guard reaches only 16.8 km.

The trade is visible in the table: the guard interval is capacity given away to buy distance. A quarter guard spends a fifth of the airtime on nothing but silence, and gets a national single frequency network in return.

On satellite. EN 302 307-1 is the second generation satellite standard. It replaces the first generation's convolutional and Reed-Solomon coding with LDPC inner coding and BCH outer coding, and adds 8PSK, 16APSK and 32APSK to QPSK. The standard states the gain: "The result is a capacity gain in the order of 30 % at a given transponder bandwidth and transmitted EIRP".

It also does something a terrestrial broadcast cannot. Because a satellite service can have a return channel for interactive and point-to-point use, "Variable Coding and Modulation (VCM) may be applied to provide different levels of error protection to different service components", and combined with a return channel this becomes adaptive coding and modulation, of which the standard says: "ACM systems promise satellite capacity gains of up to 100 % to 200 %". A terminal in heavy rain is given a robust modulation; one in clear sky is given a fast one; the transponder is no longer set for the worst receiver in the footprint.

The program shows what the modulation is worth. On a 36 MHz transponder at code rate three quarters, with a roll-off of 0.20 the symbol rate is 30 Mbaud, and the gross rate runs from 45.0 Mbit/s with QPSK to 67.5 with 8PSK, 90.0 with 16APSK and 112.5 with 32APSK. Two and a half times the rate through the same transponder, for the same bandwidth, paid for entirely in the signal to noise ratio the receiver must achieve, which is to say in dish size and in weather.

The systems, computed

# Broadcasting with no way back: the carousel, and the numbers DAB and DVB fix.
import math

C = 299792458.0

# 1. Cyclic repetition. A receiver that missed an item cannot ask for it: it
#    waits for the next time round. If an item appears every g slots and the
#    receiver tunes in at a random moment, it waits g/2 on average.
N = 20
pop = [1.0 / (i + 1) for i in range(N)]          # a Zipf popularity, the usual shape
total = sum(pop)
p = [x / total for x in pop]

def mean_wait(share):
    """share[i] is item i's fraction of the slots; the gap between its turns is
    the reciprocal, in slots, and a random arrival waits half of it."""
    return sum(p[i] / (2 * share[i]) for i in range(N))

flat = [1.0 / N] * N
proportional = list(p)
root = [math.sqrt(p[i]) / sum(math.sqrt(x) for x in p) for i in range(N)]

print("A carousel of %d items, in slots, with a receiver arriving at random:" % N)
for name, share in (("flat, every item equally often", flat),
                    ("in proportion to popularity", proportional),
                    ("by the square root of popularity", root)):
    waits = [1.0 / (2 * share[i]) for i in range(N)]
    print("  %-34s mean wait %5.2f, the rarest item %5.2f"
          % (name, mean_wait(share), waits[-1]))
print("  flat and proportional give exactly the same mean, which is half the number of")
print("  items: sending a popular item twice as often halves its wait and doubles")
print("  everyone else's in exactly compensating measure.")
print("  the square root schedule is the one that wins, and it wins by %.0f per cent,"
      % (100 * (mean_wait(flat) - mean_wait(root)) / mean_wait(flat)))
print("  at the price of making the rarest item wait %.0f per cent longer."
      % (100 * (1 / (2 * root[-1]) - 1 / (2 * flat[-1])) / (1 / (2 * flat[-1]))))

# 2. What that costs in real time, for a real carousel.
print()
print("The same carousel in seconds, at three broadcast rates:")
ITEM_BITS = 8 * 1024 * 8                          # an 8 kB page
for rate_kbps, label in ((16, "a data channel of 16 kbit/s"),
                         (128, "a channel of 128 kbit/s"),
                         (1024, "a channel of 1 Mbit/s")):
    slot_s = ITEM_BITS / (rate_kbps * 1000.0)
    print("  %-28s one slot is %5.2f s, so the mean wait is %5.1f s"
          % (label, slot_s, mean_wait(root) * slot_s))
print("  that wait is the whole reason a broadcast carousel is a design and not just a")
print("  loop: an item nobody can request has to come round often enough to be useful.")

# 3. DAB, transmission mode I, worked from the standard's own parameters. The
#    standard gives every time as a multiple of the elementary period.
print()
print("DAB, transmission mode I, worked from the standard's own parameters:")
T_ELEM = 1.0 / 2048000                            # the elementary period, seconds
TF, TU, DELTA = 196608 * T_ELEM, 2048 * T_ELEM, 504 * T_ELEM
K, L = 1536, 76
CIF_BITS, CU, CIF_MS = 55296, 864, 24e-3
print("  the elementary period is 1/2 048 000 s, and every time is a multiple of it")
print("  a transmission frame is 196 608 of them, which is %.0f ms, and carries %d"
      % (TF * 1000, L))
print("  OFDM symbols of %d carriers each" % K)
print("  the useful part of a symbol is 2 048 of them, %.0f ms, and the guard interval"
      % (TU * 1000))
print("  is 504 of them, %.2f us" % (DELTA * 1e6))
print("  a Common Interleaved Frame is %d bits every %.0f ms, so the main service"
      % (CIF_BITS, CIF_MS * 1000))
print("  channel carries %.3f Mbit/s gross" % (CIF_BITS / CIF_MS / 1e6))
print("  it is cut into %d capacity units, so one unit is %d bits, or %.3f kbit/s"
      % (CU, CIF_BITS // CU, (CIF_BITS / CU) / CIF_MS / 1000))
print("  and the guard interval is what makes a single frequency network possible:")
print("  an echo arriving up to %.2f us late still falls inside it, which is a path"
      % (DELTA * 1e6))
print("  difference of %.1f km, so two transmitters that far apart may share a frequency."
      % (DELTA * C / 1000))

# 4. DVB-T, the same idea at television rates, for 8 MHz channels.
print()
print("DVB-T in an 8 MHz channel: what the guard interval buys")
print("  mode    useful symbol   guard fraction   guard      transmitters may be")
for mode, tu in (("8K", 896e-6), ("2K", 224e-6)):
    for frac, name in ((0.25, "1/4"), (0.125, "1/8"), (0.0625, "1/16"), (0.03125, "1/32")):
        d = tu * frac
        print("  %-6s %8.0f us %12s %10.0f us %12.1f km apart"
              % (mode, tu * 1e6, name, d * 1e6, d * C / 1000))
print("  the 8K mode with a quarter guard tolerates the largest echo, so it is the one")
print("  a wide single frequency network uses; the 2K mode is for a single transmitter.")

# 5. DVB-T's outer code, and what protection costs.
print()
print("DVB-T's outer code, on MPEG transport packets of 188 bytes:")
K_RS, N_RS, T_RS = 188, 204, 8
print("  Reed-Solomon (%d, %d) adds %d bytes and corrects up to %d wrong bytes in %d"
      % (N_RS, K_RS, N_RS - K_RS, T_RS, N_RS))
print("  that is %.1f per cent of the payload spent, and it is spent before the inner"
      % (100.0 * (N_RS - K_RS) / K_RS))
print("  convolutional code has spent anything at all: a broadcast cannot ask again,")
print("  so it pays in advance.")

# 6. DVB-S2 on a satellite transponder: what the modulation is worth.
print()
print("DVB-S2 on a 36 MHz transponder, gross rates at code rate 3/4,")
print("before framing and the outer code:")
MODS = (("QPSK", 2), ("8PSK", 3), ("16APSK", 4), ("32APSK", 5))
print("  roll-off   symbol rate   " + "".join("%9s" % m for m, _ in MODS))
for roll in (0.35, 0.20):
    sym = 36e6 / (1 + roll)
    print("  %8.2f %9.1f Mbaud" % (roll, sym / 1e6)
          + "".join("%9.1f" % (sym * bits * 0.75 / 1e6) for _, bits in MODS)
          + "  Mbit/s")
print("  the same transponder and the same bandwidth, and %.1f times the rate between"
      % (MODS[-1][1] / float(MODS[0][1])))
print("  the easiest modulation and the hardest. What changes is the signal to noise")
print("  ratio the receiver must have, so the dish and the weather decide which is used.")
munotes.in843

Broadcast Systems: Cyclic Repetition, DAB and DVB

A carousel of 20 items, in slots, with a receiver arriving at random:
  flat, every item equally often     mean wait 10.00, the rarest item 10.00
  in proportion to popularity        mean wait 10.00, the rarest item 35.98
  by the square root of popularity   mean wait  8.02, the rarest item 16.98
  flat and proportional give exactly the same mean, which is half the number of
  items: sending a popular item twice as often halves its wait and doubles
  everyone else's in exactly compensating measure.
  the square root schedule is the one that wins, and it wins by 20 per cent,
  at the price of making the rarest item wait 70 per cent longer.

The same carousel in seconds, at three broadcast rates:
  a data channel of 16 kbit/s  one slot is  4.10 s, so the mean wait is  32.8 s
  a channel of 128 kbit/s      one slot is  0.51 s, so the mean wait is   4.1 s
  a channel of 1 Mbit/s        one slot is  0.06 s, so the mean wait is   0.5 s
  that wait is the whole reason a broadcast carousel is a design and not just a
  loop: an item nobody can request has to come round often enough to be useful.

DAB, transmission mode I, worked from the standard's own parameters:
  the elementary period is 1/2 048 000 s, and every time is a multiple of it
  a transmission frame is 196 608 of them, which is 96 ms, and carries 76
  OFDM symbols of 1536 carriers each
  the useful part of a symbol is 2 048 of them, 1 ms, and the guard interval
  is 504 of them, 246.09 us
  a Common Interleaved Frame is 55296 bits every 24 ms, so the main service
  channel carries 2.304 Mbit/s gross
  it is cut into 864 capacity units, so one unit is 64 bits, or 2.667 kbit/s
  and the guard interval is what makes a single frequency network possible:
  an echo arriving up to 246.09 us late still falls inside it, which is a path
  difference of 73.8 km, so two transmitters that far apart may share a frequency.

DVB-T in an 8 MHz channel: what the guard interval buys
  mode    useful symbol   guard fraction   guard      transmitters may be
  8K          896 us          1/4        224 us         67.2 km apart
  8K          896 us          1/8        112 us         33.6 km apart
  8K          896 us         1/16         56 us         16.8 km apart
  8K          896 us         1/32         28 us          8.4 km apart
  2K          224 us          1/4         56 us         16.8 km apart
  2K          224 us          1/8         28 us          8.4 km apart
  2K          224 us         1/16         14 us          4.2 km apart
  2K          224 us         1/32          7 us          2.1 km apart
  the 8K mode with a quarter guard tolerates the largest echo, so it is the one
  a wide single frequency network uses; the 2K mode is for a single transmitter.

DVB-T's outer code, on MPEG transport packets of 188 bytes:
  Reed-Solomon (204, 188) adds 16 bytes and corrects up to 8 wrong bytes in 204
  that is 8.5 per cent of the payload spent, and it is spent before the inner
  convolutional code has spent anything at all: a broadcast cannot ask again,
  so it pays in advance.

DVB-S2 on a 36 MHz transponder, gross rates at code rate 3/4,
before framing and the outer code:
  roll-off   symbol rate        QPSK     8PSK   16APSK   32APSK
      0.35      26.7 Mbaud     40.0     60.0     80.0    100.0  Mbit/s
      0.20      30.0 Mbaud     45.0     67.5     90.0    112.5  Mbit/s
  the same transponder and the same bandwidth, and 2.5 times the rate between
  the easiest modulation and the hardest. What changes is the signal to noise
  ratio the receiver must have, so the dish and the weather decide which is used.
munotes.in844

Broadcast Systems: Cyclic Repetition, DAB and DVB

Distinctions

A two-way linkA broadcast
The receiver can askYesNo
Errors handled byRetransmission when one occursRedundancy sent in advance, always
Late arrivalConnect and requestWait for the carousel
Adapts toThis receiverThe worst receiver intended to be served
Cost grows withThe number of usersNothing: the same signal serves everyone
munotes.in845

Broadcast Systems: Cyclic Repetition, DAB and DVB

Carousel scheduleMean waitThe rarest item waitsVerdict
Flat, all equal10.00 slots10.00The baseline
Proportional to popularity10.00 slots35.98No better on average, far worse in the tail
Square root of popularity8.02 slots16.9820 per cent better on average
munotes.in846

Broadcast Systems: Cyclic Repetition, DAB and DVB

DABDVB-TDVB-S2
StandardEN 300 401EN 300 744EN 302 307-1
CarriesAudio and data ensemblesMPEG transport streamMPEG and generic streams
ModulationD-QPSK on OFDMQPSK to 64-QAM on OFDMQPSK to 32APSK, single carrier
CodingConvolutionalReed-Solomon plus convolutionalBCH plus LDPC
Frame96 ms, 76 symbols, 1,536 carriers, mode I2K or 8K carriersFrames of fixed coded length
Guard interval246.09 microseconds, 73.8 km224 to 7 microseconds, 67.2 to 2.1 kmNot applicable, single carrier
Adapts per receiverNoNoYes, with VCM and ACM
munotes.in847

Broadcast Systems: Cyclic Repetition, DAB and DVB

What it does not mean

Cyclic repetition is not a loop played at random. It is a schedule, and the schedule decides how long every listener waits.

Sending popular items more often does not, by itself, help. Proportional scheduling has exactly the same mean wait as flat scheduling; only the square root rule improves it.

The guard interval is not protection against noise. It absorbs echoes, so that a delayed copy of the signal falls inside it and does no harm, which is what makes a single frequency network possible.

A single frequency network is not one transmitter. It is many transmitters radiating the same signal on the same frequency, close enough together that their signals reach a receiver inside its guard interval.

Forward error correction is not wasted when the channel is good. It is the price of having no way to ask again, and it is paid on every packet whether or not that packet needed it.

DVB-S2's adaptive modulation is not available to terrestrial broadcast. It works because a satellite service can carry a return channel from each terminal; a broadcast to a city has nothing to adapt to.

Quick revision

  • A broadcast has no return channel: no requests, no acknowledgements, no retransmission, random arrival, and one setting for everyone.
  • Cyclic repetition answers it. An item every g slots costs a random arrival g/2 on average.
  • Schedules: flat 10.00, proportional 10.00, square root 8.02 slots mean wait; the rarest item waits 10.00, 35.98 and 16.98. The square root of popularity is the rule.
  • In seconds, 20 items of 8 kB: 32.8 s at 16 kbit/s, 4.1 s at 128 kbit/s, 0.5 s at 1 Mbit/s.
  • DAB (EN 300 401): ensembles, Fast Information Channel and Main Service Channel, OFDM with D-QPSK. Mode I: frame 96 ms, 76 symbols, 1,536 carriers, guard 246.09 microseconds which is 73.8 km, CIF 55 296 bits every 24 ms giving 2.304 Mbit/s in 864 capacity units of 64 bits.
  • DVB-T (EN 300 744): MPEG packets of 188 bytes, Reed-Solomon (204, 188) correcting 8 bytes in 204 at a cost of 8.5 per cent, then convolutional coding, on OFDM in 2K or 8K mode. 8K with a quarter guard reaches 67.2 km between transmitters.
  • DVB-S2 (EN 302 307-1): BCH plus LDPC, QPSK to 32APSK, a stated capacity gain "in the order of 30 %" over the first generation, and VCM and ACM with gains stated as "up to 100 % to 200 %". On a 36 MHz transponder at rate 3/4 and roll-off 0.20: 45.0 to 112.5 Mbit/s gross.
  • Single frequency network: the guard interval turned into distance.
munotes.in848

Broadcast Systems: Cyclic Repetition, DAB and DVB

Test yourself

1. Why does a broadcast system need cyclic repetition? Because there is no return channel. A receiver cannot request an item, cannot acknowledge one and cannot ask for a retransmission, and receivers switch on at unpredictable moments. The only way a receiver can obtain something it missed, lost to interference or arrived too late for is to wait for it to be sent again, so everything has to be sent repeatedly on a cycle. That turns a one-way channel into something the receiver can treat as storage, at the cost of waiting.

2. If an item is broadcast every g slots, how long does a receiver wait for it? Up to g slots in the worst case, when it arrives just after the item has gone, and g divided by two on average, since the moment of arrival is unrelated to the schedule.

3. Why does sending popular items in proportion to their popularity not reduce the average wait? Because the extra slots have to come from the other items. Sending an item twice as often halves the wait of everyone who wants it, but it takes slots from the rest and lengthens their waits by exactly the compensating amount. The two effects cancel, and the mean over all listeners stays at half the number of items, exactly what a flat schedule gives. What does help is to increase each item's share by the square root of its popularity rather than in proportion, which in the chapter's example lowers the mean wait from 10.00 slots to 8.02, a 20 per cent improvement, at the cost of making the rarest item wait 70 per cent longer.

4. What is a single frequency network, and what makes it possible? It is a network of transmitters all radiating the same signal on the same frequency, so that a whole region needs only one channel and a receiver never has to retune. It is made possible by the guard interval: a copy of the signal arriving late from a distant transmitter falls inside the guard interval and is absorbed instead of interfering. The guard interval therefore converts directly into a distance. DAB's mode I guard of 246.09 microseconds corresponds to 73.8 kilometres, and DVB-T's 8K mode with a quarter guard, 224 microseconds, corresponds to 67.2 kilometres.

5. Describe DAB's transmission mode I. The standard fixes all times as multiples of an elementary period of one divided by 2,048,000 seconds. A transmission frame is 196,608 of them, which is 96 milliseconds, and consists of 76 OFDM symbols with 1,536 carriers, modulated with differential QPSK. The useful part of a symbol is 2,048 periods, exactly 1 millisecond, and the guard interval is 504, about 246.09 microseconds. The frame begins with synchronization symbols, then the Fast Information Channel, then the Main Service Channel. A Common Interleaved Frame of 55,296 bits is sent every 24 milliseconds, which is 2.304 Mbit/s gross, divided into 864 capacity units of 64 bits each.

munotes.in849

Broadcast Systems: Cyclic Repetition, DAB and DVB

6. Why does DVB spend so much capacity on error correction? Because it cannot ask for anything again. A two-way protocol can send data with little protection and retransmit the occasional packet that fails, paying only for the failures. A broadcast has no acknowledgements and no requests, so the protection has to be in the signal before it leaves, on every packet, whether that packet meets interference or not. DVB-T therefore applies a Reed-Solomon code of length 204 bytes to each 188 byte transport packet, which corrects up to 8 wrong bytes at a cost of 8.5 per cent of the payload, adds an interleaver so that bursts are spread across code words, and then applies an inner convolutional code on top of that.

7. What does DVB-S2 add over the first generation satellite standard, and why can it do something terrestrial broadcasting cannot? It replaces convolutional and Reed-Solomon coding with LDPC and BCH, and adds 8PSK, 16APSK and 32APSK to QPSK, which the standard says gives a capacity gain in the order of 30 per cent at the same transponder bandwidth and transmitted power. On a 36 MHz transponder at code rate three quarters and roll-off 0.20, the chapter computes 45.0 Mbit/s with QPSK rising to 112.5 with 32APSK. It also allows variable and adaptive coding and modulation, so that different components, or different terminals, are given different protection, with stated gains of up to 100 to 200 per cent. That last part depends on terminals reporting their channel conditions over a return channel, which a terrestrial broadcast to an unknown audience does not have, so a terrestrial transmitter must always be set for the worst receiver it intends to serve.

Contents This chapter on its own page

munotes.in850

Chapter One Hundred Ten

Across the Two Modules: The Comparisons Question 3 Draws On

Syllabus topic Both modules, for the paper's Question 3

In one line

The two modules ask the same questions of wildly different systems, and almost every disagreement between them comes down to one thing: in Module I the device that transmits is the device that pays, and in Module II it is not.

Why this chapter exists

The examination's first question is set on Module 1 and the second on Module 2. The third is set on both, and it is the one where a candidate who has learned each module as a separate subject comes unstuck.

The way through it is not to memorise more. It is to notice that the two modules keep asking the same five questions, and to be able to say what each system answers and why:

  1. Is there infrastructure, or is there not?
  2. Who pays for the transmission?
  3. What is being addressed, a device or a fact?
  4. What moves, and how often does the link have to change?
  5. Is there a way back?

Every comparison below is one of those five, asked of both modules.

Six systems, one table

# The two modules side by side: the same arithmetic, answering opposite questions.
import math

C = 299792458.0

def fsl(f_mhz, d_km):
    """Free space loss, Recommendation ITU-R P.525 in its logarithmic form."""
    return 32.44 + 20 * math.log10(f_mhz) + 20 * math.log10(d_km)

# The six systems the two modules between them describe.
SYSTEMS = (
    #  name                     MHz     link km   rate bit/s   module
    ("a sensor node, 802.15.4", 2450.0,    0.03,     250e3,   "I"),
    ("Bluetooth Low Energy",    2450.0,    0.01,       1e6,   "I"),
    ("GSM",                      900.0,    3.0,      9.6e3,  "II"),
    ("UMTS",                    2000.0,    1.0,     384e3,   "II"),
    ("a LEO satellite",         1600.0, 2339.0,       2.4e3, "II"),
    ("a GEO satellite",        12000.0, 36305.0,      45e6,  "II"),
)

def show_range(d_km):
    return "%6.0f m " % (d_km * 1000) if d_km < 1 else "%6.0f km" % d_km

def show_rate(bits):
    for div, unit in ((1e6, "Mbit/s"), (1e3, "kbit/s")):
        if bits >= div:
            return "%.4g %s" % (bits / div, unit)
    return "%.0f bit/s" % bits

print("One link of each, at its own frequency and its own reach:")
print("  system                  mod frequency    range       rate      loss   one way")
for name, f, d, rate, mod in SYSTEMS:
    print("  %-23s %3s %6.0f MHz %9s %11s %6.1f dB %7.2f ms"
          % (name, mod, f, show_range(d), show_rate(rate), fsl(f, d), d * 1000 / C * 1000))
losses = [fsl(f, d) for _, f, d, _, _ in SYSTEMS]
ranges = [d for _, _, d, _, _ in SYSTEMS]
print("  the reach spans a factor of %.1f million between the shortest and the longest,"
      % (max(ranges) / min(ranges) / 1e6))
print("  and the loss only %.0f dB between them: loss grows with the logarithm of"
      % (max(losses) - min(losses)))
print("  distance, which is the reason the far systems are possible at all.")

# 1. The same ten kilometres, crossed by each.
print()
print("Crossing ten kilometres, and what it takes:")
for name, f, d, rate, mod in SYSTEMS:
    hops = max(1, int(math.ceil(10.0 / d)))
    if hops == 1:
        print("  %-25s one hop, and %.1f km of reach to spare" % (name, d - 10))
    else:
        print("  %-25s %5d hops of %.2f km" % (name, hops, d))
print("  the sensor network's answer is a path; the cellular answer is a mast; the")
print("  satellite's answer is that ten kilometres is not a distance at all.")

# 2. The energy argument, and why it points opposite ways.
print()
print("One long hop or many short ones, with loss growing as the fourth power:")
N_EXP = 4.0
E_ELEC = 50e-9          # joules a bit, to run the electronics at each end of a hop
E_AMP = 1.3e-15         # joules a bit a metre to the fourth, for the amplifier
D = 300.0
print("  sending one bit %.0f m, in k equal hops, with %.0f nJ of electronics a hop:" % (D, E_ELEC * 1e9))
best = None
for k in (1, 2, 3, 4, 5, 6, 8, 12):
    total = k * 2 * E_ELEC + k * E_AMP * (D / k) ** N_EXP
    if best is None or total < best[1]:
        best = (k, total)
    print("    %2d hop%s %9.2f uJ" % (k, "s:" if k > 1 else ": ", total * 1e6))
print("  the best is %d hops, and the reason is that the amplifier term falls as k to the"
      % best[0])
print("  power %.0f while the electronics term only rises as k." % (N_EXP - 1))
print("  a sensor node runs this arithmetic because it pays for its own transmission.")
print("  a handset does not: the mast has mains power and a tall antenna, so the network")
print("  buys the long hop and the handset keeps its battery. Same formula, opposite answer.")

# 3. How often the link changes, and why. The two modules give almost the same
#    number for entirely opposite reasons.
print()
print("How long one link lasts, and what moves:")
CASES = (("a sensor node to its parent", 0.0, None, "nothing moves, so the link lasts"),
         ("a walker in a 500 m cell", 5.0 / 3.6, 500.0, "the user moves"),
         ("a car in a 3 km cell", 100.0 / 3.6, 3000.0, "the user moves"),
         ("a train in a 10 km cell", 160.0 / 3.6, 10000.0, "the user moves"),
         ("a still user, LEO spot beam", 6649.0, 600000.0, "the network moves"),
         ("a still user, LEO satellite", 6649.0, 4154000.0, "the network moves"),
         ("a dish, GEO satellite", 0.0, None, "neither moves, so the link lasts"))
for name, speed, size, why in CASES:
    if size is None:
        print("  %-28s %-24s until the equipment fails" % (name, why))
    else:
        print("  %-28s %-24s %6.1f s, %5.1f min"
              % (name, why, size / speed, size / speed / 60))
print("  a car in a city cell and a motionless caller under a satellite beam change link")
print("  at almost the same rate. In one the user is moving and in the other the cell is.")

# 4. The same word, two different problems.
print()
print("Localisation, the word both modules use:")
print("  in module I it means: where is this node, in metres, when nobody told it")
print("  in module II it means: which cell or satellite is serving this subscriber")
print("  the first is answered by ranging and geometry and is wrong by metres;")
print("  the second is answered by a database lookup and is either right or wrong.")
SPEED_SOUND = 343.0
print("  and the accuracy each needs is the reverse of what the speed suggests:")
for metres in (0.1, 1.0, 10.0):
    print("    to place something within %5.1f m, a clock must be right to %8.2f ns by"
          % (metres, metres / C * 1e9))
    print("      radio, but only to %8.2f ms by sound" % (metres / SPEED_SOUND * 1e3))
print("  which is why a sensor network that must find itself to the metre listens for")
print("  sound, and a satellite system that must do the same carries atomic clocks.")

# 5. Spread spectrum, used by both modules for opposite reasons.
print()
print("Spreading, and what each module wants from it:")
for name, chips, rate, mod in (("802.15.4, module I", 32, 250e3, "I"),
                               ("UMTS, module II", 256, 15e3, "II")):
    gain = 10 * math.log10(chips)
    print("  %-20s %3d chips a symbol, processing gain %4.1f dB" % (name, chips, gain))
print("  module I spreads to survive interference from everything else in the band.")
print("  module II spreads so that many callers can share one band at once, each")
print("  recovered by its own code. The mechanism is one; the purpose is two.")
munotes.in851

Across the Two Modules: The Comparisons Question 3 Draws On

One link of each, at its own frequency and its own reach:
  system                  mod frequency    range       rate      loss   one way
  a sensor node, 802.15.4   I   2450 MHz     30 m   250 kbit/s   69.8 dB    0.00 ms
  Bluetooth Low Energy      I   2450 MHz     10 m     1 Mbit/s   60.2 dB    0.00 ms
  GSM                      II    900 MHz      3 km  9.6 kbit/s  101.1 dB    0.01 ms
  UMTS                     II   2000 MHz      1 km  384 kbit/s   98.5 dB    0.00 ms
  a LEO satellite          II   1600 MHz   2339 km  2.4 kbit/s  163.9 dB    7.80 ms
  a GEO satellite          II  12000 MHz  36305 km   45 Mbit/s  205.2 dB  121.10 ms
  the reach spans a factor of 3.6 million between the shortest and the longest,
  and the loss only 145 dB between them: loss grows with the logarithm of
  distance, which is the reason the far systems are possible at all.

Crossing ten kilometres, and what it takes:
  a sensor node, 802.15.4     334 hops of 0.03 km
  Bluetooth Low Energy       1000 hops of 0.01 km
  GSM                           4 hops of 3.00 km
  UMTS                         10 hops of 1.00 km
  a LEO satellite           one hop, and 2329.0 km of reach to spare
  a GEO satellite           one hop, and 36295.0 km of reach to spare
  the sensor network's answer is a path; the cellular answer is a mast; the
  satellite's answer is that ten kilometres is not a distance at all.

One long hop or many short ones, with loss growing as the fourth power:
  sending one bit 300 m, in k equal hops, with 50 nJ of electronics a hop:
     1 hop:      10.63 uJ
     2 hops:      1.52 uJ
     3 hops:      0.69 uJ
     4 hops:      0.56 uJ
     5 hops:      0.58 uJ
     6 hops:      0.65 uJ
     8 hops:      0.82 uJ
    12 hops:      1.21 uJ
  the best is 4 hops, and the reason is that the amplifier term falls as k to the
  power 3 while the electronics term only rises as k.
  a sensor node runs this arithmetic because it pays for its own transmission.
  a handset does not: the mast has mains power and a tall antenna, so the network
  buys the long hop and the handset keeps its battery. Same formula, opposite answer.

How long one link lasts, and what moves:
  a sensor node to its parent  nothing moves, so the link lasts until the equipment fails
  a walker in a 500 m cell     the user moves            360.0 s,   6.0 min
  a car in a 3 km cell         the user moves            108.0 s,   1.8 min
  a train in a 10 km cell      the user moves            225.0 s,   3.8 min
  a still user, LEO spot beam  the network moves          90.2 s,   1.5 min
  a still user, LEO satellite  the network moves         624.8 s,  10.4 min
  a dish, GEO satellite        neither moves, so the link lasts until the equipment fails
  a car in a city cell and a motionless caller under a satellite beam change link
  at almost the same rate. In one the user is moving and in the other the cell is.

Localisation, the word both modules use:
  in module I it means: where is this node, in metres, when nobody told it
  in module II it means: which cell or satellite is serving this subscriber
  the first is answered by ranging and geometry and is wrong by metres;
  the second is answered by a database lookup and is either right or wrong.
  and the accuracy each needs is the reverse of what the speed suggests:
    to place something within   0.1 m, a clock must be right to     0.33 ns by
      radio, but only to     0.29 ms by sound
    to place something within   1.0 m, a clock must be right to     3.34 ns by
      radio, but only to     2.92 ms by sound
    to place something within  10.0 m, a clock must be right to    33.36 ns by
      radio, but only to    29.15 ms by sound
  which is why a sensor network that must find itself to the metre listens for
  sound, and a satellite system that must do the same carries atomic clocks.

Spreading, and what each module wants from it:
  802.15.4, module I    32 chips a symbol, processing gain 15.1 dB
  UMTS, module II      256 chips a symbol, processing gain 24.1 dB
  module I spreads to survive interference from everything else in the band.
  module II spreads so that many callers can share one band at once, each
  recovered by its own code. The mechanism is one; the purpose is two.
munotes.in852

Across the Two Modules: The Comparisons Question 3 Draws On

The first table is worth reading slowly. The reach runs from 10 metres to 36,305 km, a factor of 3.6 million, and yet the loss runs only from 60.2 dB to 205.2 dB, a span of 145 dB. That is the whole reason a geostationary link exists at all: loss grows with the logarithm of distance, so multiplying the distance by a million costs about 120 dB rather than a million times the power. [Signal Propagation: Ranges, Path Loss and How a Signal Travels] works out the mechanism; here it is the bridge between the two modules.

munotes.in853

Across the Two Modules: The Comparisons Question 3 Draws On

Infrastructure, or none

Module I: a sensor networkModule II: cellular, satellite, broadcast
Is there a base stationUsually not: a sink, reached over many hopsAlways: a mast, a gateway, a transmitter
Who is in chargeNobody, or a cluster head chosen for a roundThe network, definitively
A node failsThe network reroutesA mast fails and its cell goes dark
Deployed byScattering, often at randomPlanning, surveying and licensing
Taught in[What a Wireless Sensor Network Is], [Ad Hoc Networks: MANETs, and How a Sensor Network Differs][Cellular Systems: Cells, Clusters and Frequency Reuse], [The GSM System Architecture]
munotes.in854

Across the Two Modules: The Comparisons Question 3 Draws On

The program's second section makes the difference concrete. To cross ten kilometres a sensor network at 30 metres a hop needs 334 hops; GSM needs 4; a satellite needs one, with thousands of kilometres to spare. Three completely different answers to one distance, and each is right for its own system.

Who pays for the transmission

This is the deepest difference in the paper, and most of the others follow from it.

munotes.in855

Across the Two Modules: The Comparisons Question 3 Draws On

In a sensor network the node that transmits is running on a battery nobody will change, so the transmitter pays, and the whole design of Module I bends around that fact. In a cellular network the handset has a battery too, but the mast has mains power and a tall antenna, so the network can afford to reach further than the handset can and the handset is spared.

The program shows the same formula giving opposite answers. Sending one bit 300 metres with loss growing as the fourth power costs 10.63 microjoules in one hop and 0.56 microjoules in four, because the amplifier term falls as the cube of the hop count while the electronics term only rises linearly. A sensor node therefore relays. A handset does not relay, because it is not the one paying for the long hop.

Module IModule II
Who pays for reachThe node itselfThe base station, gateway or transmitter
ConsequenceMany short hopsOne long hop to the mast
Battery replacedNeverBy the user, nightly
The design goalLifetime in months or yearsCapacity and quality
Taught in[Single Hop or Multiple Hops: The Energy Argument Worked Out], [How Long a Node Lasts: The Energy Budget Worked Out][Cellular Systems: Cells, Clusters and Frequency Reuse], [Channel Allocation, Cell Splitting, Sectorisation and Cell Breathing]

What is addressed: a device, or a fact

Module IModule II
The address isAn attribute: the temperature in the north fieldAn identity: this subscriber, this number
Duplicate answersUseful: aggregate themWrong: there is one subscriber
The network mayCombine, drop and summarise in flightOnly carry, unchanged
CalledData-centricAddress-centric
Taught in[Design Principles: Data Centricity, Location, Activity and Heterogeneity], [Directed Diffusion and Rumour Routing][Localization and Calling in GSM], [The GSM System Architecture]

The consequence for an answer script: in Module I the network is allowed to change the data, and in Module II it is not. In-network processing is a virtue in one module and a fault in the other.

Sharing the medium

Module IModule II
The scarce thingEnergyCapacity, and licensed spectrum
So the MAC optimisesTime spent with the radio switched onUsers per megahertz per cell
Typical answerDuty cycling, listen and sleep, preamble samplingFDMA, TDMA, CDMA, scheduled and assigned
Idle listening isThe largest waste there isNot a consideration
Collisions costEnergy, twiceCapacity
Taught in[MAC Protocols for Sensor Networks: The Job and Where the Energy Goes], [S-MAC: Periodic Listen and Sleep, and Keeping Neighbours in Step], [Duty Cycling: Preamble Sampling, B-MAC and X-MAC][Multiplexing: Space, Frequency, Time and Code], [The GSM Radio Interface: Carriers, the TDMA Frame and Bursts]
munotes.in856

Across the Two Modules: The Comparisons Question 3 Draws On

One mechanism, two purposes: spreading

Both modules spread a signal over far more bandwidth than the data needs, and they do it for different reasons. The program computes the processing gain of each: 15.1 dB from 802.15.4's 32 chips a symbol, and 24.1 dB from UMTS at 256 chips.

Module I spreads to survive: the 2.4 GHz band is unlicensed and full of other people's transmissions, so the gain is bought as immunity. Module II spreads to share: every caller gets a different code and all of them use the same band at the same time, and the gain is what lets the receiver pull one out of the sum. The mathematics is the same and the purpose is opposite.

Taught in [Spread Spectrum and Direct Sequence] and [The 802.15.4 Physical Layer] for the first, and [The UMTS Radio Interface: W-CDMA, Codes, Power Control and Soft Handover] for the second.

Routing: discovered, scheduled, or refused

Module I: a sensor networkModule II: cellularModule II: satellite
Topology changesUnpredictably, as nodes sleep and dieHardly at all: masts do not moveOn a timetable, known years ahead
So routes areDiscovered and repaired by protocolConfigured, and rarely changedPrecomputed and switched on schedule
What is optimisedEnergy, and the life of the weakest nodeLoad and qualityDelay and hop count
Taught in[Routing Strategies in WSNs: A Map], [Energy-aware Routing][The GSM System Architecture][Routing in Satellite Systems]

What moves, and how often the link changes

The program's fourth section is the one to memorise, because it contains the paper's neatest cross-module observation.

A car at 100 km/h in a 3 km cell holds its link for 108 seconds. A motionless caller under a low satellite's spot beam holds its link for 90.2 seconds. Almost the same number, and for exactly opposite reasons: in one the user crosses a stationary cell, and in the other a cell moving at 6.65 kilometres a second crosses a stationary user.

At the two ends of the range, a sensor node's link to its parent and a fixed dish's link to a geostationary satellite both last until the equipment fails, because in one nothing moves and in the other everything moves together.

What movesLink lastsHandover caused by
A sensor node to its parentNothingUntil it failsA node dying, not motion
A walker in a 500 m cellThe user360 sThe user
A car in a 3 km cellThe user108 sThe user
A still user, LEO spot beamThe network90.2 sThe network
A still user, LEO satelliteThe network624.8 sThe network
A dish to a GEO satelliteNeither, relativelyUntil it failsNothing
Taught in[Handover in GSM], [Localization and Handover in Satellite Systems]
munotes.in857

Across the Two Modules: The Comparisons Question 3 Draws On

One word, two problems: localisation

Both modules use the word, and they do not mean the same thing.

In Module I it means: where is this node, in metres, when nobody told it and it has no receiver for a navigation signal. It is answered by ranging and geometry, and the answer is wrong by metres.

In Module II it means: which cell, or which satellite, is serving this subscriber. It is answered by a database lookup in a home or visitor register, and the answer is either right or wrong.

The program prices the first. To place something within one metre a clock must be right to 3.34 nanoseconds by radio, but only to 2.92 milliseconds by sound, a difference of nearly a million. That is why a sensor network that must find itself to the metre listens for a sound pulse alongside the radio one, and why a satellite navigation system that must do the same carries atomic clocks and corrects for relativity.

Taught in [Time Synchronisation and Localisation] for the first, and [Localization and Calling in GSM] and [Localization and Handover in Satellite Systems] for the second.

Is there a way back?

Two-way: sensor network, cellular, satellite telephonyOne-way: broadcast
Receiver can askYesNo
Errors handled byRetransmission when one happensRedundancy sent in advance, always
Arriving lateConnect and requestWait for the carousel
Adapted toThis receiverThe worst receiver intended to be served
Cost grows withThe number of usersNothing
Taught in[Traditional Transport Control Protocols: TCP and UDP], [Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST][Broadcast Systems: Cyclic Repetition, DAB and DVB]

There is a second, subtler crossing here. A sensor network floods when it must reach everyone, and a broadcaster broadcasts, and both are one-to-many. But flooding in a dense network causes the broadcast storm of [Flooding, Gossiping and the Broadcast Storm], where every node rebroadcasts and the medium collapses, while a broadcast transmitter reaches a million receivers for the cost of one transmission. One-to-many is cheap when one transmitter owns the channel and ruinous when every receiver is also a transmitter.

Transport, and what reliability is worth

Module IModule II
Reliability wantedSometimes per packet, usually per eventPer byte, in order
AcknowledgementHop by hop, where it is cheapEnd to end
Congestion meansA funnel near the sinkToo many callers, handled by admission
An acknowledgement costsEnergy comparable to the dataAlmost nothing
Taught in[Why a Sensor Network Cannot Simply Run TCP], [Transport Protocols Built for Sensor Networks: PSFQ, ESRT, CODA and RMST][Traditional Transport Control Protocols: TCP and UDP], [GEO: The Geostationary Orbit]
munotes.in858

Across the Two Modules: The Comparisons Question 3 Draws On

The satellite chapter belongs on the right-hand side for a reason worth stating: over a geostationary link the same protocol that works on a cable delivers 1.93 Mbit/s with a 64 kB window, whatever the link's own speed. Module I's objection to that protocol is energy; Module II's is delay; the protocol is the same one.

Security: what is being protected

Module IModule II
The threatNodes physically captured, keys extracted, false data injectedEavesdropping, cloning, unauthorised use of the network
Trust anchorKeys placed before deploymentA tamper-resistant card issued by the operator
Public-key cryptographyExpensive, often avoidedAffordable
A compromised deviceIs inside the network, and speaks with authorityIs one subscriber, and is barred
Taught in[Security in Ad Hoc and Sensor Networks: Goals, Constraints and Attacks], [Keys and Link Security: Key Predistribution, SPINS and 802.15.4][GSM Security]

Scale, cost and lifetime

Module IModule II
Number of devicesHundreds to thousands, in one deploymentMillions, across a country
Cost of one deviceAs low as possible: it will be abandonedSubstantial, and paid for by a subscription
LifetimeLimited by the battery: months or yearsLimited by the contract or the satellite's fuel
Failure of oneExpected, and designed forAn incident
IdentityOften none: a node is interchangeableCentral: the subscriber is the product
Taught in[The Challenges of Wireless Sensor Networks], [Figures of Merit: Scalability, Robustness and Measuring a Network][GSM and Its Mobile Services], [UMTS and IMT-2000]

How to answer Question 3

A cross-module comparison is not a list of facts about two systems. It is an argument, and it has a shape that earns marks in the eight minutes the question allows.

  1. Name the difference that causes the rest. Usually: who pays for the transmission, or whether there is infrastructure, or whether there is a way back.
  2. Give the number that shows it. One will do: 334 hops against 4; 10.63 microjoules against 0.56; 108 seconds against 90.2; 60.2 dB against 205.2.
  3. Draw the consequence for each side. The sensor network relays and sleeps; the cellular network builds a mast and sells a contract.
  4. Say what is the same. Spreading, OFDM, TDMA and free-space loss appear in both modules. Saying so shows the comparison was understood rather than memorised.
  5. Finish with the trade. Every difference is a purchase: energy bought with hops, capacity bought with reuse, coverage bought with delay, reliability bought with redundancy.

Quick revision

  • Q.3 is set on both modules. The five questions that cross: infrastructure or not, who pays, device or fact addressed, what moves, and is there a way back.
  • Reach 10 m to 36,305 km, a factor of 3.6 million, for a loss span of only 145 dB: loss grows with the logarithm of distance.
  • Ten kilometres: 334 hops at 30 m, 4 GSM cells, one satellite hop with thousands of kilometres spare.
  • Energy: one bit 300 m costs 10.63 microjoules in one hop and 0.56 in four. The node relays because it pays; the handset does not because the mast pays.
  • What moves: a car in a 3 km cell holds a link 108 s; a still user under a LEO spot beam 90.2 s. Same number, opposite cause.
  • Spreading: 15.1 dB at 32 chips to survive interference in Module I; 24.1 dB at 256 chips to share a band in Module II.
  • Localisation: metres by ranging in Module I, a register lookup in Module II. One metre needs 3.34 ns by radio or 2.92 ms by sound.
  • One-to-many is cheap when one transmitter owns the channel and ruinous when every receiver is also a transmitter.
munotes.in859

Across the Two Modules: The Comparisons Question 3 Draws On

Test yourself

1. Compare a wireless sensor network with a cellular network. A sensor network has no infrastructure: nodes are scattered, often at random, and reach a sink over many short hops, with no node permanently in charge. A cellular network is planned, licensed and built around masts, and every handset talks to one directly. The difference that causes most of the rest is who pays for the transmission. A sensor node runs on a battery nobody will replace, so it relays over short hops, sleeps as much as it can and lets the network combine and summarise data in flight; the chapter's arithmetic puts one bit over 300 metres at 10.63 microjoules in a single hop and 0.56 in four. A handset's mast has mains power and a tall antenna, so the network buys the long hop and the handset keeps its battery. The sensor network addresses facts and tolerates failures as normal; the cellular network addresses subscribers, must not alter their data, and treats a failure as an incident.

2. In what sense do a sensor network and a cellular network both use spread spectrum, and in what sense not? Both spread a signal over far more bandwidth than the data needs, and both gain by the ratio of chip rate to symbol rate: 32 chips a symbol in 802.15.4 gives 15.1 dB, and 256 in UMTS gives 24.1 dB. The purposes are opposite. The sensor network spreads to survive, because it works in an unlicensed band full of other people's transmissions and buys the gain as immunity. The cellular system spreads to share, because every caller is given a different code and all of them transmit in the same band at the same time, and the gain is what lets a receiver recover one signal from the sum.

munotes.in860

Across the Two Modules: The Comparisons Question 3 Draws On

3. Compare handover in a cellular network, in a low satellite constellation, and in a sensor network. In a cellular network handover happens because the user moves out of a cell: a car at 100 km/h in a 3 km cell holds its link for 108 seconds. In a low satellite constellation the user need not move at all, because the spot beams and the satellites sweep past at 6.65 kilometres a second: a motionless caller changes beam about every 90.2 seconds and satellite every 10.4 minutes, and a call may also be moved to a different gateway while the radio link is untouched. The two rates are almost equal for opposite reasons. In a sensor network there is no handover in this sense: nothing moves, and a node changes parent only when a node fails, a battery dies or a link degrades.

4. Why does a sensor network relay over many hops while a handset does not? Because transmission energy grows far faster than distance, roughly as its fourth power in the model used here, while the electronics at each end of a hop cost a fixed amount per bit. Splitting a distance into k hops divides the amplifier term by k to the power three and multiplies the electronics term by k, so there is an optimum: over 300 metres it is four hops, which costs 0.56 microjoules against 10.63 for one. A sensor node runs that arithmetic because it is the device paying. A handset is not: the mast has mains power, a tall antenna and no battery to protect, so the network absorbs the cost of the long hop and the handset spends nothing on relaying anyone else's traffic.

5. The word localisation appears in both modules. What does each mean by it? Module I means finding where a node physically is, in metres, when nobody configured it and it carries no navigation receiver; it is answered by ranging and geometry and the answer carries an error. Module II means finding which cell or satellite currently serves a subscriber; it is answered by a lookup in the home and visitor registers, with a satellite system adding a register of where its satellites are, and the answer is simply right or wrong. The accuracies required are very different: to place something within a metre a clock must be right to 3.34 nanoseconds by radio but only 2.92 milliseconds by sound, which is why sensor networks range with sound and satellite navigation carries atomic clocks.

munotes.in861

Across the Two Modules: The Comparisons Question 3 Draws On

6. Why is one-to-many cheap in a broadcast system and ruinous in a sensor network? Because of who owns the channel. A broadcast transmitter owns its frequency, transmits once, and reaches every receiver in its footprint at no extra cost per listener, which is why a satellite reaches a continent from one uplink. In a sensor network every receiver is also a transmitter sharing one medium, so flooding makes every node rebroadcast, the rebroadcasts collide with each other, and the medium collapses in what is called the broadcast storm. The same one-to-many intention costs nothing in one case and destroys the network in the other.

7. What do both modules have in common? More than a first reading suggests. Free-space loss governs every link in the book, from ten metres to 36,305 kilometres, and grows with the logarithm of distance in both. TDMA appears as a schedule that saves energy in Module I and as a way to fit callers into a carrier in Module II. Spreading appears in both, for opposite purposes. OFDM carries DAB and DVB and appears in the newer wireless standards. Routing, congestion, security and localisation are asked of both. What differs is never the physics; it is which resource is scarce, and therefore which of the available answers is the right one.

Contents This chapter on its own page

munotes.in862

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself for the same subject.

Issue
Done!