The source of data is critical, and must be relevant for the model and purpose being considered. As a result, local data would usually be more likely to be applicable than international data. However, international data may be relied on provided it remains reasonably applicable to Australia.66 Consideration should be given to how international data would need to be modified for applicability before use. Data sets may be inaccurate if affected by selection bias, such that the data is not representative of a population. Notably, individuals or groups that have faced systemic discrimination may be inaccurately represented, or under-represented in data sets. This may also be true of model outputs used as input data within a separate AI system – those model outputs may be biased in some manner. For example, they may have higher error rates for marginalised groups. Data sets that are outdated or incomplete may also lead to inaccurate or inappropriate outcomes in an AI system. A data set may be incomplete if it contains insufficient data points or insufficient characteristics or details about individuals. A way to address these issues may be to obtain additional data. For example, an insurer may seek to acquire data that is up to date, more complete or more representative. Insurers should consider the costs, risks and benefits of obtaining additional data sets, as well as any alternative options. Aside from additional financial costs, collecting extensive personal information about customers might give rise to other reputational or legal risks, such as obligations under the Privacy Act 1988 (Cth). In order to rely on the data exemption under the ADA, DDA, or SDA, it may be necessary for the insurer to obtain any data that is available or reasonably obtainable, where such data is reasonable to rely on. If an insurer does not have data from internal claims, they may not have sufficient data to rely on the data exemption,67 and may need data from external sources where it is reasonable to rely on such data. If such additional data cannot be reasonably obtained, then the no data exemption available under the ADA and DDA may apply. (ii) Preprocess the data Data preprocessing is an important and common step to address missing, incomplete or inaccurate data points. Raw data usually requires preprocessing before it can be effectively used to train an AI system. Preprocessing techniques help to ensure more reliable results. Provided they are used appropriately, they may also help to ensure less discriminatory results. Some preprocessing techniques include: • smoothing, which removes ‘noise’, such as outliers, corrupt or meaningless information, from data • grouping data points, such as banding of age groups, or clustering detailed medical conditions under wider terms • transformation, which turns the data into the proper format needed for analysis • imputation, which replaces missing data with suitably substituted data. Guidance Resource: Artificial intelligence and discrimination in insurance pricing and underwriting • 2022 • 23

Select target paragraph3