Refers to having confidence that data meets quality standards and is ready to act on (e.g., analyze, understand, make decisions about). Data trust cannot be taken for granted; there should be evidence of the data’s quality standards.
The Data Management Association of the UK defines six dimensions of data quality required for trustworthy data:
- Accuracy —degree to which data correctly describes the real-world object or event in question.
- Completeness —proportion (percentage) of data stored versus 100% complete (e.g., are there blank values indicating certain data has not been populated).
- Consistency —absence of difference when comparing two or more representations of an item against a definition (e.g., do dates and names match).
- Timeliness — extent to which data is current enough to represent reality for realistic use.
- Uniqueness — no item or entity instance is recorded more than once based upon how that item is identified (nonduplication).
- Validity or conformity — extent to which data conforms to the syntax (format, type, range) of its definition.
There is important context behind data trust: (1) the multiplicity of data sources, and (2) potential errors in data management:
- Manual data quality management cannot handle immense volumes of data that come from multiple sources of data; e.g., from SaaS and web applications, direct data entry like web forms, unstructured sources like social media posts, machines such as smartphones and “Internet of Things” devices.
- Errors in data management are a factor in quality data. Errors are introduced by people who make mistakes, machines that are not foolproof, and errors can be introduced by data that often passes through complex information systems coded by multiple developers.