education, ethnicity, occupation, and home ownership. Moreover, the Census is population-wide, which promotes both aggregate and disaggregate analysis. Table 2 below illustrates the process undertaken in creating our population spine from the 2013 Census. We begin with the full sample of just over 4.3 million individuals in the 2013 Census, before dropping 325,338 individuals without a household identification number (ID). We are then left with a sample of just over 4 million individuals, equating to 1.5 million households. One explanation for an individual lacking a household ID is that the individual was not at home during Census night. This means that for some households, we underestimate household size. The impact of dropping these individuals on the net equivalised household income is not straightforward: the lacking individual might have contributed to the overall household income but may also have elevated the poverty risk by enlarging the respective OECD scale. 10 The next exclusion is to drop households with no income information (neither from Inland Revenue, WfF tax credit or AS) for the entire year. 11 This trims our dataset to 1.4 million households (3.7 million individuals), as shown in row 3 of Table 2. This is the population that is used to calculate the poverty threshold. This involves linking all individuals (and the associated 1.4 million households) in the population spine with their income information. Using the unique individual identifier that is available in all used datasets, we match individuals with their records from Inland Revenue, WfF tax credits and AS. The Inland Revenue records contain income information at the monthly level from April 1999 across seven different sources: wages and salaries, benefits, pensions, paid parental leave, withholding payments, ACC claims and student loans (as well as the relevant tax deductions for each income source). WfF data also contain information at the monthly level, though this might result in missing households that receive the WfF tax credit at an annual level. Using these income data sources, we are able to identify an individual’s income (as well as the income of their associated household unit). Based on the household structure, the equivalised household income distribution is then used to calculate the poverty threshold. In results not shown here (for brevity sake), we compared individuals dropped due to lack of household ID to those in the reduced sample. In general, we find these individuals are older, have lower levels of income, and are more likely to be New Zealand European. Given these differences, this sub-sample is not a random component of the population, and therefore the loss of these individuals from the sample can potentially bias resulting estimates. 11 We assume that the three main reasons for the lack of income information are (listed in no specific order): i) household members received income from sources that are not covered by the data being used (e.g., self-employment income, overseas income and investment income are filed in IR3 and not captured in the EMS), ii) linking issues between household members in the 2013 Census and their income records, and iii) members have not received any income. Excluding these households from the sample might bias the income distribution which is used to derive the poverty threshold. However, the direction of the bias is unclear as the listed cases include potentially low- and high-income households. 10 Page 16

Select target paragraph3