The source of data is critical, and must be relevant
for the model and purpose being considered. As
a result, local data would usually be more likely to
be applicable than international data. However,
international data may be relied on provided
it remains reasonably applicable to Australia.66
Consideration should be given to how international
data would need to be modified for applicability
before use.
Data sets may be inaccurate if affected by selection
bias, such that the data is not representative of a
population. Notably, individuals or groups that have
faced systemic discrimination may be inaccurately
represented, or under-represented in data sets.
This may also be true of model outputs used as
input data within a separate AI system – those
model outputs may be biased in some manner.
For example, they may have higher error rates for
marginalised groups.
Data sets that are outdated or incomplete may
also lead to inaccurate or inappropriate outcomes
in an AI system. A data set may be incomplete if
it contains insufficient data points or insufficient
characteristics or details about individuals.
A way to address these issues may be to obtain
additional data. For example, an insurer may seek
to acquire data that is up to date, more complete or
more representative.
Insurers should consider the costs, risks and
benefits of obtaining additional data sets, as well
as any alternative options. Aside from additional
financial costs, collecting extensive personal
information about customers might give rise to
other reputational or legal risks, such as obligations
under the Privacy Act 1988 (Cth).
In order to rely on the data exemption under
the ADA, DDA, or SDA, it may be necessary for
the insurer to obtain any data that is available
or reasonably obtainable, where such data is
reasonable to rely on. If an insurer does not have
data from internal claims, they may not have
sufficient data to rely on the data exemption,67 and
may need data from external sources where it is
reasonable to rely on such data. If such additional
data cannot be reasonably obtained, then the no
data exemption available under the ADA and DDA
may apply.
(ii) Preprocess the data
Data preprocessing is an important and common
step to address missing, incomplete or inaccurate
data points.
Raw data usually requires preprocessing before
it can be effectively used to train an AI system.
Preprocessing techniques help to ensure
more reliable results. Provided they are used
appropriately, they may also help to ensure
less discriminatory results. Some preprocessing
techniques include:
•
smoothing, which removes ‘noise’, such as
outliers, corrupt or meaningless information,
from data
•
grouping data points, such as banding of age
groups, or clustering detailed medical conditions
under wider terms
•
transformation, which turns the data into the
proper format needed for analysis
•
imputation, which replaces missing data with
suitably substituted data.
Guidance Resource: Artificial intelligence and discrimination in insurance pricing and underwriting • 2022 • 23