Part II Collecting and analysing data
• Question: What is the mean average of your dataset?
• Answer: Add up the incomes you have for all households (1120 + 241 + 876 + 201 + 112 + 345
+ 567 + 156 + 154 + 1345 = 5117). Then divide that number by the number of households you
have (5117/10). Your answer is 511.7.
In a spreadsheet, you can calculate this using the formula =AVERAGE
The mean can give you a good estimate of what is “normal” when the rest of your data is distributed
evenly above and below it. However, if the rest of your data is “skewed” to either side, a different
measure of the average may be more appropriate.
The median is the numerical value separating the higher half of values in the dataset from the lower
half. It is useful when the rest of your data are not evenly distributed on either side of the mean. In
the example of average household income above, the mean income value is 511.7. However, this is
quite a large number in relation to most of the values in the dataset. It results because there are a few
households with very high incomes, which skew the data. In this case, the median may provide a better
estimate of what is “normal” than the mean.
So how is the median calculated? Firstly, sort the data (ascending or descending, it does not matter) and
the value in the middle of the dataset is the median. If there is an even number of values in the dataset,
take the average of the two middle-values.
• Question: What is the median household income?
• Answer: First, sort the data: 112, 154, 156, 201, 241, 345, 567, 876, 1120, 1345. There are ten
values and the middle two values are 241 and 345. Find the average between these two numbers
(241 + 345 / 2 = 293). The median is 293.
In a spreadsheet, you can calculate this using the formula =MEDIAN
The mode is the value that appears most often in a set of data. Sometimes neither the mean nor
median really tells us what we want to know. For example, if you would like to know the average number
of children per household enrolled in school, you might have the following dataset:
0, 1, 1, 1, 1, 2, 2, 2, 3, 5
The mean number of children enrolled in school per household is 1.8, and the median is 1.5. But what
you really want to find out is how many children do the majority of households have enrolled in school.
You can see that “1” child is the most frequent answer. This is the mode.
In a spreadsheet, you can calculate this using the formula =MODE
In the case that more than one value is the most frequent, the dataset can be bi-modal (two average
values) or multi-modal (more than two average values).
8.3.4. Variation
The next important piece of information you might want to identify is the size of the variation in the
dataset. This is crucial when looking at aggregated and disaggregated data. For instance, say you are
investigating the realization of the right to work. You want to find out the average unemployment rate.
However, you might also want to test how representative the average is of different municipalities within
your country. There are two common measures for doing this.
The standard deviation is a measure of by how much, on average, data values are off the mean. The
following three steps show how to calculate it:
1. Sum the square of the differences between the values and the mean.
2. Divide that sum by the number of values minus one.
3. Take the square root.
Chapter 8: Analysing data: A short introduction to working with spreadsheets | 91
Select target paragraph3
Connect to a paragraph
Connect to an entity
Disable highlights
Add to table of contents