IIUM Repository

An assessment of open data sets completeness

Ali, Abdulrazzak and Asmai, Siti and Emran, Nurul Akmar and Ismail, Amelia Ritahani (2019) An assessment of open data sets completeness. International Journal of Advanced Computer Science and Applications, 10 (6). pp. 557-562. ISSN 2158-107X E-ISSN 2156-5570

[img] PDF (pdf)
Restricted to Registered users only

Download (229kB) | Request a copy
[img] PDF (scopus)
Restricted to Registered users only

Download (377kB) | Request a copy

Abstract

The rapid growth of open data sources is driven by free-of-charge contents and ease of accessibility. While it is convenient for public data consumers to use data sets extracted from open data sources, the decision to use these data sets should be based on data sets’ quality. Several data quality dimensions such as completeness, accuracy, and timeliness are common requirements to make data fit for use. More importantly, in many cases, high-quality data sets are desirable in ensuring reliable outcomes of reports and analytics. Even though many open data sources provide data quality guidelines, the responsibility to ensure data of high quality requires commitment from data contributors. In this paper, an initial investigation on the quality of open data sets in terms of completeness dimension was conducted. In particular, the results of the missing values in 20 open data sets measurement were extracted from the open data sources. The analysis covered all the missing values representations which are not limited to nulls or blank spaces. The results exhibited a range of missing values ratios that indicated the level of the data sets completeness. The limited coverage of this analysis does not hinder understanding of the current level of data completeness of open data sets. The findings may motivate open data providers to design initiatives that will empower data quality policy and guidelines for data contributors. In addition, this analysis may assist public data users to decide on the acceptability of open data sets by applying the simple methods proposed in this paper or performing data cleaning actions to improve the completeness of the data sets concerned.

Item Type: Article (Journal)
Additional Information: 4296/77296
Uncontrolled Keywords: Data completeness; missing values; open data; open data sources; data collection
Subjects: Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Kulliyyahs/Centres/Divisions/Institutes (Can select more than one option. Press CONTROL button): Kulliyyah of Information and Communication Technology > Department of Computer Science
Kulliyyah of Information and Communication Technology > Department of Computer Science
Depositing User: Amelia Ritahani Ismail
Date Deposited: 08 Jan 2020 16:18
Last Modified: 28 Feb 2020 11:56
URI: http://irep.iium.edu.my/id/eprint/77296

Actions (login required)

View Item View Item

Downloads

Downloads per month over past year