education, ethnicity, occupation, and home ownership. Moreover, the Census is population-wide,
which promotes both aggregate and disaggregate analysis.
Table 2 below illustrates the process undertaken in creating our population spine from the
2013 Census. We begin with the full sample of just over 4.3 million individuals in the 2013 Census,
before dropping 325,338 individuals without a household identification number (ID). We are then left
with a sample of just over 4 million individuals, equating to 1.5 million households. One explanation
for an individual lacking a household ID is that the individual was not at home during Census night.
This means that for some households, we underestimate household size. The impact of dropping
these individuals on the net equivalised household income is not straightforward: the lacking
individual might have contributed to the overall household income but may also have elevated the
poverty risk by enlarging the respective OECD scale. 10
The next exclusion is to drop households with no income information (neither from Inland Revenue,
WfF tax credit or AS) for the entire year. 11 This trims our dataset to 1.4 million households (3.7 million
individuals), as shown in row 3 of Table 2. This is the population that is used to calculate the poverty
threshold. This involves linking all individuals (and the associated 1.4 million households) in the
population spine with their income information. Using the unique individual identifier that is available
in all used datasets, we match individuals with their records from Inland Revenue, WfF tax credits and
AS. The Inland Revenue records contain income information at the monthly level from April 1999
across seven different sources: wages and salaries, benefits, pensions, paid parental leave,
withholding payments, ACC claims and student loans (as well as the relevant tax deductions for each
income source). WfF data also contain information at the monthly level, though this might result in
missing households that receive the WfF tax credit at an annual level. Using these income data
sources, we are able to identify an individual’s income (as well as the income of their associated
household unit). Based on the household structure, the equivalised household income distribution is
then used to calculate the poverty threshold.
In results not shown here (for brevity sake), we compared individuals dropped due to lack of household ID to those in the
reduced sample. In general, we find these individuals are older, have lower levels of income, and are more likely to be New
Zealand European. Given these differences, this sub-sample is not a random component of the population, and therefore
the loss of these individuals from the sample can potentially bias resulting estimates.
11 We assume that the three main reasons for the lack of income information are (listed in no specific order): i) household
members received income from sources that are not covered by the data being used (e.g., self-employment income,
overseas income and investment income are filed in IR3 and not captured in the EMS), ii) linking issues between household
members in the 2013 Census and their income records, and iii) members have not received any income. Excluding these
households from the sample might bias the income distribution which is used to derive the poverty threshold. However, the
direction of the bias is unclear as the listed cases include potentially low- and high-income households.
10
Page 16