
==== Front
Data Brief
Data Brief
Data in Brief
2352-3409
Elsevier

S2352-3409(24)00780-7
10.1016/j.dib.2024.110816
110816
Data Article
StatMetaQA: A dataset for closed domain question answering in Indonesian statistical metadata
Rachmawati Nur
Yulianti Evi evi.y@cs.ui.ac.id
⁎
Universitas Indonesia Computer Science Faculty of Computer Science, Universiity of Indonesia Campus in Depok West Java Indonesia Depok, West Java 16431, Indonesia
⁎ Corresponding author. evi.y@cs.ui.ac.id
14 8 2024
12 2024
14 8 2024
57 11081618 3 2024
4 8 2024
5 8 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
A closed domain question answering (QA) dataset in statistical metadata is important to build an effective QA system about statistic. This dataset can be utilized to train or fine-tune the QA models in statistic. Further, it can also be exploited to evaluate the effectiveness of any QA methods in statistical domain. In this research, we build a new dataset of statistical metadata documents and question-answer pairs annotations of these documents in Indonesian language, called StatMetaQA (Statistical Metadata Question Answering). The collection of statistical metadata documents is used as the knowledge base of a QA system, while the collection of question-answer pairs annotations is used to train or fine-tune the QA models in statistic. The collection of statistical metadata documents, consisting of 861 statistical activity metadata documents and 1,231 statistical indicator metadata documents, was obtained from a website managed by the Statistics Indonesia (http://sirusa.bps.go.id). Next, the collection of question-answer pairs about statistical metadata, consisting of 28,863 question-answer pairs from 1,000 statistical metadata documents, was obtained using two strategies: human and automatic annotation. Here, 7353 question-answer pairs were manually annotated by human, and 21,510 question-answer pairs were automatically generated by machine using our predefined templates that were applied on some document fields of statistical metadata.

Keywords

Statistical metadata
Question answering
Dataset
Indonesia
==== Body
pmcSpecifications TableSubject	Natural Language Processing, Question Answering	
Specific subject area	Closed Domain Question Answering in Statistical Metadata	
Data format	Raw	
Type of data	Text/String	
Data collection	The collection of 2,092 statistical metadata documents was scraped from a website managed by the Statistics Indonesia (http://sirusa.bps.go.id). The collection of 28,863 question-answer pairs was annotated from 1,000 samples of statistical metadata documents in our collection using two strategies: human and automatic annotations. Human annotation was performed by six university students in statistic major who have been familiar with statistical terms. Next, automatic annotation was performed by machine using templates that are defined on some fields in the statistical metadata.	
Data source location	University of Indonesia
Zone: West Java
Country : Indonesia	
Data accessibility	Repository name: StatMetaQA
Data identification number: https://doi.org/10.5281/zenodo.11517629
Direct URL to data: https://github.com/wawatsmart/StatMetaQA
Instructions for accessing these data: displayed in the URL above	
Related research article	-	

1 Value of the Data

• This dataset is valuable because it is the first QA dataset in the statistical metadata domain. None of previous work has created this dataset, which consists of a collection of statistical metadata documents and question-answer pairs annotations about statistical metadata. This dataset is useful to build an Indonesian QA system in statistical domain.

• The collection of statistical metadata documents in this dataset consists of 861 statistical activity metadata documents and 1,231 statistical indicator metadata documents. It can be used as a knowledge base (i.e., source of information to find the answers) for a QA system in statistical domain.

• The collection of question-answer pairs annotations about statistical metadata in this dataset consists of 28,863 question-answer pairs that are extracted from 1,000 metadata documents. Here, 7,353 question-answer pairs were manually annotated by humans, while 21,510 question-answer pairs were automatically annotated by machine using our predefined templates. The big amount of this question-answer pairs annotations will benefit the learning process in the machine reading comprehension tasks in order to find the answers based on the given context. More specifically, it can be used to train deep-learning models or fine-tune transformer-based models for a specific question answering task in statistical domain.

• The collection of question-answer pairs annotations in this dataset uses a universal format in machine reading comprehension tasks. This format is also used in the well-known dataset in question answering, i.e., SQuAD [1]. Therefore, researchers can easily adopt this dataset for their research purposes.

2 Background

Many people need a clear information about statistic, for example, “how many Indonesian citizens are there based on the 2023 census data?” or “What is the sampling method used in the statistical activity of the 2022 producer price survey?”. This information can be obtained using a question answering (QA) system in statistical metadata, i.e., a document which contains important information about statistic. This QA system is useful when people want to know directly the answers to their questions in statistical domain. To build a closed domain QA system in statistic, we need a QA dataset in statistical domain. However, such dataset still does not exist yet because none of the previous work has focused on question answering in statistical metadata. Therefore, in this study, we build a new question answering dataset in statistical metadata, called StatMetaQA, which consists of a collection of statistical metadata documents and their question-answer pair annotations in Indonesian language. This dataset can be used to train deep-learning models or fine-tune the transformer-based models to find answer spans from the given context.

3 Data Description

This dataset consists of two elements: document collection and question-answer pair annotations.A. The Collection of Statistical Metadata Documents

The collection of statistical metadata documents in StatMetaQA was obtained from a website managed by the Statistics Indonesia (http://sirusa.bps.go.id). This website archives a collection of statistical metadata documents consisting of statistical activity metadata, statistical indicator metadata, and statistical variable metadata.

The statistical activity metadata describes general information, planning, design, data collection & processing, variables, and indicators of a particular statistical activity. It comprises three categories: basic, sectoral and special statistics. Based on the Law of the Republic of Indonesia No. 16 of 1997 on Statistics, basic statistics are statistics intended for broad purposes, both for the government and society, which have cross-sectoral characteristics, national scale, macro scale, and the implementation of which is the responsibility of the agency. Examples of basic statistical activities are population censuses, agricultural censuses, village censuses, etc. Then, sectoral statistics are statistics intended to meet the needs of certain agencies in the context of carrying out government and development tasks which are the main tasks of the agency concerned. Examples of sectoral statistics activities metadata are disaster data compilation metadata by BNPB (National Disaster Management Agency) and health profile metadata provided by the health service. Next, special statistics are statistics aimed at meeting the specific needs of the world of business, education, social culture and other interests in community life, which are carried out by institutions, organizations, individuals and/or elements in other communities. Examples of special statistics activities are research activities from the community.

The statistical indicator metadata contains information about the indicators (i.e., control variables that can be used to measure changes in an event or activity) used in a basic statistical activity. Meanwhile, the statistical variable metadata contains the variables used in a statistical activity. It presents in every statistical activity metadata document.

To build the StatMetaQA dataset, this study utilizes the statistical metadata documents that come from basic statistical activity metadata and statistical indicator metadata. Only these two types of metadata that were taken into account because they have many fields containing a long text, which would potentially be a good source for question-answer annotations by human. This is different from sectoral statistical activity metadata, special statistical activity metadata, and statistical variable metadata which do not contain such long text elements. The attributes extracted from basic statistical activity metadata in our dataset together with the contents for each attribute are presented in Table 1.Table 1 The attributes of basic statistical activity metadata in StatMetaQA dataset.

Table 1Attribute	Example of Text	
Id	3293	
metadata_type	kegiatan dasar (basic activity)	
activity_code	3572	
year	2020	
activity_name	Pemutakhiran Data Perkembangan Desa (Updating Village Development Data)	
producer	Subdit. Stat. Ketahanan Wilayah (Subdirectorate of Regional Resilience Statistics)	
sector	-	
source_of_funds	APBN (State Budget)	
organizer	Subdit. Stat. Ketahanan Wilayah (Subdirectorate of Regional Resilience Statistics)	
general_explanation	Pada tahun 2014, Pemerintah pertama kali melalui Kementerian Keuangan dan BPS telah menyusun Indeks Kesulitan Geografis (IKG) yang dihitung dari data Potensi Desa (Podes) 2014. Angka ini kemudian dijadikan salah satu input formulasi besaran dana desa pada tahun 2015 - 2020. Selain itu data Podes juga digunakan untuk … (In 2014, the Government, through the Ministry of Finance and BPS, first compiled a Geographical Difficulty Index (IKG) which was calculated from 2014 Village Potential (Podes) data. This figure was then used as an input for the formulation of the amount of village funds in 2015 - 2020. Apart from that Podes data is also used to ….)	
objectives_and_benefits_of_activity	- Menyediakan data dasar untuk menghitung Indeks Kesulitan Geografis (IKG) yang nantinya akan dipergunakan sebagai salah satu variabel dalam pengalokasian Dana Desa (DD).
- Menyediakan …
(- Provide basic data for calculating the Geographic Difficulty Index (IKG) which will later be used as a variable in allocating Village Funds (DD).
- Provides …)	
activity_frequency	Tahunan (Annual)	
activity_history	Kegiatan ini merupakan lanjutan kegiatan Updating Podes yang mulai dilaksanakan tahun 2019. (This activity is a continuation of the Updating Podes activity which began to be implemented in 2019.)	
changes_that_occurred_from_previous_activities	Perubahan yang terjadi pada kegiatan Updating Podes 2020 dibandingkan dengan kegiatan Updating Podes Tahun 2019 adalah adanya penambahan variabel yang dimuat. (The change that occurred in the 2020 Updating Podes activity compared to the 2019 Updating Podes activity was the addition of variables that were loaded.)	
data_collection_frequency	Harian. (Daily).	
type_of_data_collection	Longitudinal	
reference_used	Undang-Undang Nomor 6 Tahun 2014 tentang Desa, Peraturan Menteri Keuangan No 199/PMK.07/2017, Instrumen Podes tahun-tahun sebelumnya. (Law Number 6 of 2014 concerning Villages, Minister of Finance Regulation No. 199/PMK.07/2017, Podes Instruments of previous years.)	
classification_used	MFD Semester II 2019	
variable_content	Variabel utama dan konsep yang digunakan
1. Nama variabel : Umur.
Konsep definisi : Umur dihitung dalam tahun dengan …
(Main variables and concepts used
1. Variable name: Age.
Definition concept: Age is calculated in. …)	
how_to_collect_data	Sensus (Census)	
area_coverage	Seluruh kabupaten/kota di Indonesia (All districts/cities in Indonesia)	
observation_unit	Desa/kelurahan dan wilayah dengan sebutan lain yang setingkat desa/kelurahan. (Villages/subdistricts and areas with other designations at village/subdistrict level.)	
scope_of_respondent	Kepala Desa/Lurah; Kepala Unit Pemukiman Transmigrasi (UPT); Kepala Satuan Permukiman Transmigrasi (SPT) atau narasumber lain yang relevan; (Village Head/Lurah; Head of the Transmigration Settlement Unit (UPT); Head of the Transmigration Settlement Unit (SPT) or other relevant sources; .)	
using_secondary_data_from_other_work_units	Ya (Yes)	
method_of_collecting_data	Wawancara langsung,Lainnya. (Live interviews, Others.)	
pilot_study	Ya (Yes)	
instrument	PODES2020.UPDATING dengan menggunakan CAPI (PODES2020.UPDATING using CAPI (computer assisted personal interviewing).	
data_collection_officer	Staff,KSK,Mitra. (Staff, KSK, Partners.)	
supervisor_coordinator	2307	
enumerator	9316	
conduct_officer_training	Tidak (No)	
method_to_know_data_collection_performance	Revisit,Task Force.	
adjustment_non_response	Tidak ada penggantian sampel (There is no sample replacement)	
work_unit_that_performs_processing	Sendiri, Integrasi pengolahan. (Own, processing integration.)	
processing_method	Batching,Editing,Coding,Data Entri/Scan. (Batching, Editing, Coding, Data Entry/Scan.)	
the_application_technology_used	aplikasi CAPI pada android,web monitoring, dan SPSS. (CAPI application on Android, web monitoring, and SPSS.)	
analysis_method	Tabulasi silang (Cross tabulation)	
analysis_unit	Desa (Village)	
there_are_other_work_units_that_use_this_data	Tidak (No)	
treatment_of_outliers	Lainnya (Other)	
data reliability	The example of the document doesn't contain information about this attribute.	
data_quality_improvement	1. Berbagai metode melalui sistem monitoring dilakukan pada Podes 2020 untuk meningkatkan kualitas data yang dihasilkan.
2. Adanya pengecekan tabulasi …
(1. Various methods through monitoring systems were carried out in Podes 2020 to improve the quality of the data produced.
2. There is a tabulation check ...)	
data_comparability	Antar wilayah dan antar waktu (Between regions and between times)	
data_revision_method	Dalam Updating Podes 2020, secara eksplisit sebenarnya tidak ada metode revisi data. Apabila terdapat data yang diduga oleh pemeriksa/pengawas masih belum sesuai dengan kondisi lapangan, maka. (In the 2020 Updating Podes, there is actually no data revision method explicitly. If there is data that the examiner/supervisor suspects is still not in accordance with field conditions, then …)	
conduct_evaluation_study	Ya (Yes)	
recommendations_for_upcoming_future	1. Perbaikan aplikasi CAPI agar tidak lagi terjadi error pada saat pencacahan.
2. Perbaikan pada web monitoring agar ….
(1. Improve the CAPI application so that errors no longer occur during enumeration.
2. Improvements to web monitoring so that …)	
publication content	Diseminasi Publikasi 1. Judul publikasi : Statistik Infrastruktur Indonesia 2020. (Publication Dissemination 1. Publication title: Indonesian Infrastructure Statistics 2020.)	
microdata content	Diseminasi Data Mikro 1. Nama data mikro : Hasil Updating Podes 2020. (Dissemination of Micro Data 1. Name of micro data: 2020 Podes Updating Results.)	
availability_year_data	dari tahun 2014 (from 2014)	
questionnaire_documentation_content	Kuesioner 1. Kuesioner : Kuesioner Updating PODES 2020 (pendataan di tingkat desa);. (Questionnaire 1. Questionnaire: 2020 PODES Updating Questionnaire (data collection at village level);)	
guideline_content	Pedoman
1. Pedoman : Pedoman Pencacah Pemutakhiran Data Perkembangan Desa 2020.
2. Pedoman : Pedoman Instalasi dan …
(Guidelines
1. Guidelines: Guidelines for Enumerators for Updating Village Development Data 2020.
2. Guidelines: Guidelines for Installing and …)	
full_content	The combination of text of all attributes above	

We can see from the above table that there are a lot of attributes describing various information contained in the basic statistical activity metadata. The average length of statistical activity metadata documents in our dataset is 744.9 words. This differs from the statistical indicator metadata, which in general is relatively short. The average length of statistical indicator metadata documents in our dataset is 101.3 words. Table 2 describes the attributes extracted from statistical indicator metadata in our dataset together with the contents for each attribute.Table 2 The attributes of statistical indicator metadata in StatMetaQA dataset.

Table 2No	Attribute	Example of Text	
1	id	117	
2	metadata_type	indikator (indicator)	
3	indicator_code	124	
4	indicator_name	Indeks Berantai Produksi Perikanan Budidaya (Aquaculture Production Chain Index)	
5	definition_concept	Indeks berantai produksi perikanan budidaya adalah angka yang menunjukkan perbandingan produksi perikanan budidaya pada tahun tertentu terhadap periode tahun sebelumnya. Produksi perikanan budidaya mencakup semua hasil budidaya ikan/binatang air lainnya/tanaman air yang dipanen dari sumber perikanan alami atau dari tempat pemeliharaan, …
(The aquaculture production chain index is a number that shows the comparison of aquaculture production in a particular year compared to the previous year. Aquaculture production includes all cultivated fish/other aquatic animals/aquatic plants harvested from natural fisheries sources or from rearing areas, whether cultivated by fishing companies or fishing households.)	
6	Benefit	• Indeks ini dapat memberikan informasi tentang perkembangan produksi suatu jenis perikanan budidaya setiap tahun berjalan dibandingkan dengan tahun sebelumnya.

• Untuk melihat besarnya perubahan produksi suatu jenis perikanan budidaya setiap tahun berjalan dibangkan tahun sebelumnya

(• This index can provide information about the development of production of a type of aquaculture each year compared to the previous year. • To see the magnitude of changes in production of a type of aquaculture each year compared to the previous year)

	
7	additional_information	Sumber Data : Statistik Perikanan Budidaya Indonesia, Direktorat Jenderal Perikanan Budidaya Kementerian Kelautan dan Perikanan
Level Penyajian : Nasional
Publikasi: Indikator Pertanian, Statistik Perikanan Budidaya Indonesia. …
(Data Source: Indonesian Aquaculture Statistics, Directorate General of Aquaculture, Ministry of Maritime Affairs and Fisheries. Level of Presentation: National Publication: Agricultural Indicators, Indonesian Aquaculture Statistics. Information)	
8	interpretation	Iit > 100, berarti produksi suatu jenis perikanan budidaya mengalami peningkatan dari periode tahun sebelumnya Iit = 100, berarti berarti produksi suatu jenis perikanan budidaya tidak mengalami perubahan Iit < 100, berarti produksi suatu jenis perikanan budidaya mengalami penurunan dari periode tahun sebelumnya
(Iit > 100, meaning the production of a type of aquaculture has increased from the previous year's period. Iit = 100, meaning the production of a type of aquaculture has not changed. Iit < 100, meaning the production of a type of aquaculture has decreased from the previous year's period)	
9	full_concept	The combination of text of all attributes above	

Overall, the collection of Indonesian statistical metadata documents in our dataset consists of 861 basic statistical activity metadata documents and 1,231 indicator metadata documents. Both of them have CSV and pickle format. There are 64 attributes for statistical activity metadata documents and 9 attributes for statistical indicator metadata in our dataset, as described earlier in Table 1, Table 2. This collection serve as the knowledge base to find answers to the users’ questions in statistic.B. The Collection of Question-Answer Pairs Annotations

The collection of question-answer pairs was annotated from metadata documents in our dataset. This collection is a machine reading comprehension dataset that can be used to train deep learning models or fine-tune transformer-based models to develop a QA system in statistic. This dataset contains 28,863 question-answer pairs that were obtained by annotating 1,000 metadata documents in our dataset (500 basic statistical activity metadata documents and 500 statistical indicator documents). Out of 28,863 question-answer pairs, 7,353 question-answer pairs were manually annotated by human and 21,510 question-answer pairs were manually annotated by machine. In the human annotation process, human annotators create natural language questions that could be answered by a text passage in the metadata documents, and annotate such text passages (span of text) in the metadata documents as the answers. Six annotators from undergraduate students in statistic major were recruited to perform the annotation. On the other hand, in the automatic annotation process, some predefined templates based on the metadata document's fields were used to automatically generate questions, and the text belongs to the fields were automatically used as the answers.

Since there are manual and automatic annotations, this dataset has two types of questions: human and automatic. Human questions were produced naturally by humans, while automatic questions were built automatically using templates based on some fields in the metadata document. Human questions are more natural than automatic question in which it may use different vocabularies (or have small word overlaps) with the text identifying answers. See Table 5 for few examples of our human questions. On the other hand, automatic questions are generated using templates. Therefore, it potentially has similar vocabularies (or have high word overlaps) with the text identifying answers. See Table 6 for the question templates used to generate our automatic questions. Based on this intuition, human questions are in general harder than automatic questions. Therefore, a QA system will be more difficult to find accurate answers for human questions compared to automatic questions.

Our question-answer pairs dataset is randomly split into training, validation, and testing data. The splitting is performed individually for each type of questions and as a whole. The statistics of data in each split are described in Table 3. For human questions, the proportion for training, validation, and testing are 80:10:10, with the splitting are conducted based on documents. Here, out of 1,000 documents annotated, the question-answer pairs created from 800 documents were used as training data, those created from 100 documents were used as validation data, and those created from the other 100 documents were used as testing data. Here, we have 5,763 questions in the training data, 810 questions in the validation data, and 780 questions in the testing data. For the automatic questions, the exact numbers of validation and testing data follows the numbers in the human questions. We justify that this enables us to compare the performance of a QA system in answering human questions and automatic questions in fair, since they have the same number of testing data. At last, we also provide a split of merging data, which simply combines together the data for human and automatic questions for each split.Table 3 Question-answer split.

Table 3Question Type	Train	Eval	Test	Total	
Human Question	5,763	810	780	7,353	
Automatic Question	19,920	810	780	21,510	
Merge (Human-Automatic Question)	25,683	1,620	1,560	28,863	

Each data in the collection of question-answer pairs annotations in our dataset are stored in JSON format using the universal JSON format that has been used in a well-known QA dataset, SQuAD [1]. Therefore, it can be easily utilized by other researchers for their research purposes. The attributes for each data are presented in Table 4. In total, there are 15 attributes for each question.Table 4 Attributes of question-answer pair annotation in StatMetaQA dataset.

Table 4No	Attributes	Descriptions	
1	document_id	the identifier of the document.	
2	title	the tittle of the document.	
3	paragraphs	Paragraphs is the paragraphs of document. But in this dataset, one document text collects in one paragraphs only. Every paragraphs consists of document_id, context and qas.	
4	context	the context of the document.	
5	qas	It consists of question, id, and answers.	
6	question	the question.	
7	id	the identifier of the question.	
8	answers	It consists of answer_id, document_id, question_id, text, answer_start, answer_end, answer_category.	
9	answer_id	the identifier of the answer.	
10	question_id	the identifier of the question.	
11	text	the answer text.	
12	answer_start	the start token of the answer in the context	
13	answer_end	the end token of the answer in the context	
14	answer_category	answer_category is the category of the answer. The category of the answer consist of short, long, yes, no and other. For this dataset we set answer_category with null.	
15	is_impossible	It explain about the question is answerable or not. It set ‘True” if the question is answerable. It sets “False” if the question is not answerable	

4 Experimental Design, Materials and Methods

A. The Flow of the StatMetaQA Dataset Creation Process

In general, there are two flows of creating the StatMetaQA dataset. The first one is the flow of creating a collection of statistical metadata documents. The second one is the flow of creating a collection of question-answer pairs from the statistical metadata documents. All of these flows are illustrated in Fig. 1.Fig. 1 The flow of creating StatMetaQA dataset.

Fig 1

The flow of creating the collection of statistical metadata documents in the StatMetaQA dataset are as follows:• Document Collection

We collect all statistical metadata documents that are available at Sirusa website (http://sirusa.bps.go.id) by implementing a Python program. A total of 2,460 basic statistical activity metadata documents and 1,669 statistical indicator metadata documents were obtained.

• Document Selection

Some criteria were applied to select documents that contain long text, so they can be used later in the human annotation process to create question-answer pairs data. A statistical indicator metadata document is selected if it has contents for the definition concept, additional information, or interpretation fields. Next, a basic statistical activity metadata document is selected if it has contents for the general explanation, history activity, or main variables fields. The reason of choosing these fields as our criteria is because we analyze that there are a lot of human questions that can potentially be created from the text in this field, since they contain long text. This document selection process results in 1,844 basic statistical activity metadata documents and 1,467 statistical indicator metadata documents that satisfy our selection criteria.

• Document Preprocessing

Some preprocessing steps that are performed to the metadata documents include deleting the NaN values, decoding characters that cannot be read in CSV to UTF-8, removing noisy terms (content with html tags), and translating the metadata document entry codes.

• Duplication Checking Based on Cosine Similarity

The duplication of documents is checked using cosine similarity. We use 90% similarity score as the threshold of the cosine similarity value. If cosine similarity value between two documents is more than 90%, we assume that they are duplicate and one of them will be filtered out. This process is conducted in few iterations until there is no more duplication cases found in our data. This process results in 2,092 unique documents consisting of 861 statistical activity metadata documents and 1,231 indicator metadata documents.

The flow of creating the collection of question-answer annotations regarding statistical metadata in the StatMetaQA dataset is initiated with the Document Sampling process. A random sampling was performed to select 1,000 documents that would be annotated (from a total of 2,092 documents in our dataset mentioned above). We used stratified random sampling with same size allocation: 500 samples from basic statistical activity metadata documents and 500 samples from statistical indicator metadata documents. More detailed process in the human annotation and manual annotation are described in the following text.

The specific flow of creating human questions data is described as follows:• Creating Annotation Guideline

The annotation guideline contains information about the task descriptions, the detailed instruction to perform annotation, examples of questions & answers annotation, and the step-by-step procedures to use the Haystack annotation tool [2]. The guideline is expected to give a complete overview on how the human annotation should be performed.

• Annotator Recruitment

The annotators in our study are six diploma students in statistic major. Therefore, they have good knowledge about statistic and are already familiar with many statistical terms. We personally contacted the annotators to ask for their interest in the annotation process.

• Annotator Training

The 2-hours training was given to the annotators on how the annotation should be conducted using the annotation tool that we built using Haystack tool. During the training, the annotation guideline was explained to the annotators. Then at the end of the day, a practice session was performed to check whether the annotator understood how to perform the annotation process.

• Pilot Annotation

Pilot test was carried out to annotate a small number of documents by annotators. Then, the annotation results will be evaluated by researchers to identify the issues encountered in the annotation, and the evaluation results will be discussed together with the annotators in the discussion forum. The purpose of this test is to make the annotators more familiar with the annotation task as well as to discover early the annotation mistakes that may be done by the annotator. This is important to prevent the mistakes to occur again in the actual annotation process. Our pilot test includes 10 documents (5 basic statistical activity metadata documents and 5 statistical indicator metadata documents) that were annotated by all annotators. The purpose of providing the same document was to see the similarity of understanding between each annotator after training was carried out.

• Actual Annotation

Each annotator performs the actual annotation on 150 disjoint documents, including 75 indicator metadata documents and 75 basic statistical activity metadata documents. In total, 900 documents were annotated in this step. As was done in the research by Bojic et al. [3], the annotator would make annotations by reading the document, creating the questions and annotating the answers in the context. This annotation process was carried out within 7 days. At the end of each day, researchers created a discussion forum to evaluate the annotators' results for that day and discussed to them the errors found by researchers. This aims to avoid the same errors to happen in the next day annotation. Examples of the question-answer pair that has been created from a given context text are shown in Table 5.Table 5 Examples of question-answer pairs in human questions.

Table 5No	Question	Answer	Context	
1	Sejak tahun berapa Pendidikan Non Formal (Paket A, Paket B, dan Paket C) turut diperhitungkan pada Indikator APM?
(Since what year has Non-Formal Education (Package A, Package B and Package C) been taken into account in the APM Indicator?)	2007
(2007)	Nama indikator : Angka Partisipasi Murni (APM). Konsep definisi : Proporsi penduduk pada kelompok umur jenjang pendidikan tertentu yang masih bersekolah terhadap penduduk pada kelompok umur tersebut. Sejak tahun 2007, Pendidikan Non Formal (Paket A, Paket B, dan Paket C) turut diperhitungkan. Manfaat : Untuk mengukur daya serap sistem pendidikan terhadap penduduk usia sekolah. Keterangan tambahan : . Interpretasi : APM menunjukkan seberapa banyak penduduk usia sekolah yang sudah dapat memanfaatkan fasilitas pendidikan sesuai pada jenjang pendidikannya. Jika APM = 100, berarti seluruh anak usia sekolah dapat bersekolah tepat waktu.
(Indicator name: Pure Participation Rate (APM). Definition concept: The proportion of the population in an age group with a certain level of education who is still in school compared to the population in that age group. Since2007, Non-Formal Education (Package A, Package B and Package C) has been taken into account. Benefits: To measure the educational system's absorption capacity for the school age population. Additional information : . Interpretation: APM shows how much of the school age population is able to utilize educational facilities according to their educational level. If APM = 100, it means thatall school age children can go to school on time.)	
2	Apa artinya jika APM = 100?
(What does it mean if APM = 100?)	seluruh anak usia sekolah dapat bersekolah tepat waktu
(all school age children can go to school on time)	
3	Apa saja jenis-jenis Pendidikan non formal yang diperhitungkan dalam proporsi penduduk pada kelompok umur jenjang pendidikan tertentu sejak tahun 2007?
(What types of non-formal education are taken into account in the proportion of the population in certain age groups with educational levels since 2007?)	Paket A, Paket B, dan Paket C
(Package A, Package B, and Package C)	

• Overlap Document Annotation

Annotation of overlap documents was carried out to enable the calculation of inter-annotator agreement. This annotation was performed on 90 documents (45 basic statistical activity metadata documents and 45 indicator metadata documents). There are two processes in this case: reference question-answer annotation, and alternative answer annotation. In the reference question-answer annotation, each annotator was asked to write questions and answers to 15 different documents (7 basic activity metadata documents and 8 indicator metadata documents). These answers are considered as the ground truth answers for the corresponding questions. Next, in the alternative answer annotation, the annotators are asked to annotate the answers to the questions that are created by other annotators. For example, if Annotator 1 created questions and reference answers for Doc 1-15, then the rest annotators (Annotators 2-6) will annotate alternative answers for those questions.

• Inter-rater Agreement Score Calculation

This inter-annotator agreement calculation is performed to understand the level of agreement of annotation among annotators in annotating answers for human questions. Our calculation refers to the researches conducted on SleepQA [3], RadQA [4], FriendQA [5], CMQA [6], and SQuAD [7], where consensus will be calculated using exact matching (EM) and F1-score. The calculation then refers to the research conducted by Bojic et al. (2022) [3]. For each question obtained in the overlap document annotation above, researchers compute the EM and F1-score between its reference answer and each of its alternative answers, then we take the maximum score as the agreement score for that question. The scores for all questions are then averaged as the final inter-rater agreement score.

Next, the specific flow of creating automatic questions data are as follows:• Question Templates Creation

In this step, templates are used to create automatic questions based on statistical metadata document fields. The templates used to create questions are “WH Question + field name + activity name + year” for basic statistical activity metadata and “WH Question + field name + indicator name” for statistical indicator metadata. The list of automatic question template is displayed in Table 6 below.Table 6 The automatic question templates.

Table 6No	The Automatic Question Template	The Translated Automatic Question Template in English Version	
1	Siapa produsen dari kegiatan [nama kegiatan][tahun]?	What are the benefits of [indicator_name] indicator?	
2	Apa sektor kegiatan dari [nama kegiatan][tahun]?	What is the sector of activity of [indicator_name] indicator?	
3	Apa sumber dana dari kegiatan [nama kegiatan][tahun]?	What is the source of funds for [activity_name] activity?	
4	Siapa penyelenggara dari kegiatan [nama kegiatan][tahun]?	Who is the organizer of [name of activity][year]?	
5	Apa tujuan dan manfaat dari kegiatan [nama kegiatan][tahun]?	What is the objectives and benefit of activity of [name of activity][year]?	
6	Berapa frekuensi kegiatan dari [nama kegiatan][tahun]?	What is the frequency of activities from [name of activity][year]?	
7	Apa perubahan yang terjadi pada kegiatan [nama kegiatan][tahun] dibandingkan kegiatan sebelumnya?	What changes have occurred in the activity [name of activity][year] compared to the previous activity?	
8	Berapa frekuensi pengumpulan data dari kegiatan [nama kegiatan][tahun]?	What is the frequency of data collection from [activity name][year] activity?	
9	Apa tipe pengumpulan data dari kegiatan [nama kegiatan][tahun]?	What is the type of data collection from [activity name][year] activity?	
10	Apa referensi yang digunakan pada kegiatan [nama kegiatan][tahun]?	What references are used in the activity [name of activity][year]?	
11	Apa saja klasifikasi yang digunakan pada kegiatan [nama kegiatan][tahun]?	What classifications are used for [name of activity][year] activity?	
12	Bagaimana cara pengumpulan data pada kegiatan [nama kegiatan][tahun]?	How is data collected on [name of activity][year] activity?	
13	Apa jenis rancangan sampel pada kegiatan [nama kegiatan][tahun]?	What is the type of sample design for [name of activity][year] activity?	
14	Apa metode pemilihan sampel stage terakhir pada kegiatan [nama kegiatan][tahun]?	What is the method for selecting the final stage sample for activity [name of activity][year]?	
15	Apa metode pemilihan sampel probabilitas pada kegiatan [nama kegiatan][tahun]?	What is the probability sample selection method for [name of activity][year] activity?	
16	Apa kerangka sampel yang digunakan pada kegiatan [nama kegiatan][tahun]?	What sampling frame was used for [name of activity][year] activity?	
17	Berapa keseluruhan fraksi sampel/overal sampling fraction pada kegiatan [nama kegiatan][tahun]?	What is the overall sampling fraction for [name of activity][year] activity?	
18	Apakah kegiatan [nama kegiatan][tahun] menggunakan perkiraan sampling error?	Does [name of activity][year] activity use sampling error estimates?	
19	Berapa perkiraan sampling error dari kegiatan [nama kegiatan][tahun]?	What is the estimated sampling error for [name of activity][year] activity?	
20	Apa unit sampel pada kegiatan [nama kegiatan][tahun]?	What is the sample unit in [name of activity][year] activity?	
21	Bagaimana alokasi sampel pada kegiatan [nama kegiatan][tahun]?	How is the sample allocated to [name of activity][year] activity?	
22	Apa cakupan wilayah pada kegiatan [nama kegiatan][tahun]?	What is the area coverage of [name of activity][year] activity?	
23	Apa unit observasi pada kegiatan [nama kegiatan][tahun]?	What is the unit of observation for [name of activity][year] activity?	
24	Apa cakupan responden pada kegiatan [nama kegiatan][tahun]?	What is the scope of respondent in [name of activity][year] activity?	
25	Apakah kegiatan [nama kegiatan][tahun] menggunakan data sekunder dari unit kerja instansi lain?	Does [name of activity][year] activity use secondary data from other agency work units?	
26	Apa metode pengumpulan data pada kegiatan [nama kegiatan][tahun]?	What is the method for collecting data on [name of activity][year] activity?	
27	Apakah kegiatan [nama kegiatan][tahun] melakukan pilot study?	Does [name of activity][year] activity conduct a pilot study?	
28	Apa instrumen yang digunakan pada kegiatan [nama kegiatan][tahun]?	What instruments are used in [name of activity][year] activity?	
29	Siapa saja petugas pengumpul data pada kegiatan [nama kegiatan][tahun]?	Who are the data collection officers for [name of activity][year] activity?	
30	Berapa jumlah pengawas/kortim pada kegiatan [nama kegiatan][tahun]?	How many supervisors/coordinators are there for [name of activity][year] activity?	
31	Berapa jumlah pencacah pada kegiatan [nama kegiatan][tahun]?	How many enumerators were there for [name of activity][year] activity?	
32	Apakah kegiatan [nama kegiatan][tahun] mengadakan pelatihan petugas?	Does [name of activity][year] activity provide officer training?	
33	Apa metode untuk mengetahui kinerja pengumpulan data pada kegiatan [nama kegiatan][tahun]?	What is the method to determine the performance of data collection on [name of activity][year] activity?	
34	Apa penyesuaian non respon yang diterapkan pada kegiatan [nama kegiatan][tahun]?	What non-response adjustments apply to [name of activity][year] activity?	
35	Unit kerja apa yang melakukan pengolahan pada kegiatan [nama kegiatan][tahun]?	What work unit carries out processing on [name of activity][year] activity?	
36	Apa metode pengolahan data pada kegiatan [nama kegiatan][tahun]?	What is the data processing method for [name of activity][year] activity?	
37	Apa saja teknologi aplikasi yang digunakan pada kegiatan [nama kegiatan][tahun]?	What application technologies are used in [name of activity][year] activity?	
38	Apa metode estimasi yang digunakan pada kegiatan [nama kegiatan][tahun]?	What estimation method is used for [name of activity][year] activity?	
39	Apa komposisi penimbang pada kegiatan [nama kegiatan][tahun]?	What is the weighing composition for activity [name of activity][year]?	
40	Apa metode analisis pada kegiatan [nama kegiatan][tahun]?	Apa metode analisis pada kegiatan [nama kegiatan][tahun]?	
41	Apa unit analisis pada kegiatan [nama kegiatan][tahun]?	What is the unit of analysis for [name of activity][year] activity?	
42	Apakah ada unit kerja lain yang menggunakan data ini pada kegiatan [nama kegiatan][tahun]?	Are there other work units that use this data for [name of activity][year] activity?	
43	Bagaimana perlakuan terhadap outlier pada kegiatan [nama kegiatan][tahun]?	How are outliers treated in [name of activity][year] activity?	
44	Bagaimana reliabilitas data pada kegiatan [nama kegiatan][tahun]?	What is the reliability of the data on [name of activity][year] activity?	
45	Bagaimana peningkatan kualitas data pada kegiatan [nama kegiatan][tahun]?	How can data quality be improved for [name of activity][year] activity?	
46	Bagaimana keterbandingan data pada kegiatan [nama kegiatan][tahun]?	How is the data comparable for [name of activity][year] activity?	
47	Apa metode revisi data pada kegiatan [nama kegiatan][tahun]?	What is the data revision method for [name of activity][year] activity?	
48	Apa informasi tentang kualitas data pada kegiatan [nama kegiatan][tahun]?	What information about the quality of data on [name of activity][year] activity?	
49	Apakah kegiatan [nama kegiatan][tahun] melakukan studi evaluasi?	Does [name of activity][year] activity carry out evaluation studies?	
50	Apa rekomendasi untuk kegiatan [nama kegiatan][tahun]?	What are the recommendations for [name of activity][year] activity?	
51	Apa saja publikasi yang dihasilkan pada kegiatan [nama kegiatan][tahun]?	What publications were produced during [name of activity][year] activity?	
52	Apa saja data mikro yang dihasilkan pada kegiatan [nama kegiatan][tahun]?	What micro data is generated for [name of activity][year] activity?	
53	Kapan data mikro dari kegiatan [nama kegiatan][tahun] mulai tersedia?	When will micro data from [name of activity][year] activity become available?	
54	Apa saja data/variabel yang tidak bisa diberikan kepada pengguna data pada kegiatan [nama kegiatan][tahun]?	What data/variables cannot be provided to data users for [name of activity][year] activity?	
55	Apa saja kuesioner yang digunakan pada kegiatan [nama kegiatan][tahun]?	What questionnaires are used in [name of activity][year] activity?	
56	Apa saja pedoman yang digunakan pada kegiatan [nama kegiatan][tahun]?	What are the guidelines used for [name of activity][year] activity?	
57	Apa definisi dari variabel [nama variabel utama] pada kegiatan [nama kegiatan][tahun]?	What is the definition of the variable [name of main variable] in [name of activity][year] activity?	
58	Apa periode enumerasi dari variabel [nama variabel utama] pada kegiatan [nama kegiatan][tahun]?	What is the enumeration period of the variable [name of main variable] in [name of activity][year] activity?	

• Answer Creation

The answer to each automatic question was taken directly from the complete contents of the field. For example, the answer to the question with the template of “What are the producer of [indicator_name] indicator?” is obtained from the text in the field “producer”.

After the human and automatic questions data are generated, the data validation process is performed to examine the validity of the data. For human questions dataset, validation is carried out on the entire dataset.For automatic questions dataset, researchers validated the test set only. In the validation process, this research analyzes several factors, such as question grammatical errors, question-answer validity/coherence errors, answer span annotation errors, question ambiguity errors, and duplication errors. The duplication error happens when both the question and answer have exactly the same wording or when the question has exactly the same wording but the answer is different. If question grammatical error/answer span annotation error is found, the researcher will make repair. If a question-answer validity/coherence error is found, the researcher will discard the question. If question ambiguity error is found, the researcher will hold a discussion forum with annotators to determine what the correct question is. If the duplication error where questions and answers with exactly the same wording is found, the researcher will delete the question. If duplication error occurs where the question is worded exactly the same but the answer is different, the researcher will make repair.B. Examining the Quality of Datasets

As described earlier, we calculate F1-score and exact match scores to measure the level of agreement of annotation among annotators to annotate answers for human questions. The results are shown in Table 7. We obtain high agreement scores in both metrics which shows that the annotators tend to agree in the annotation results of the given documents. This indicates consistent annotation are performed by annotators, and therefore we argue that this contributes positively to the quality of our dataset.Table 7 The inter-annotator agreement calculation.

Table 7Metric	Score (%)	
F1-Score	97.47	
Exact Match	88.10	

To examine the quality of automatic questions in our dataset, we performed a user study to test whether there is a difference between questions created by machines and those created by humans. We hypothesize if there is no difference between human questions and automatic questions, then it means that the quality of automatic questions is comparable to those of human questions. There are 47 users who participated in this study, and all of them are the staffs at the Statistics Indonesia who are used to work with statistical data in their job. In this user study, 5 human questions and 5 automatic questions were chosen randomly to be evaluated by users. These questions are mixed together, and the type of questions (i.e., human or automatic) are not informed to the users. Given a question, the users are asked two questions, such as “Does this question make sense?” and “Is this a human question (a question made by a human) or an automatic question (an automatic question made by a machine)?”.

The results of the user study are described in Table 8. We can see that the average number of logical human questions is 76.17%, while the average number of logical automatic questions is 80.00%. These two results show that the human questions and automatic questions are logically comparable. Then, human questions are more often detected as human questions, which indicates that it flows well as natural questions. On the other hand, the automatic questions are more often detected as human questions, which indicates that they are comparable with or look like human questions and therefore also flow like natural questions. We obtain insight from this study that the quality of automatic questions is good as it is comparable to human questions.Table 8 The user study results.

Table 8Measurement	Value (%)	
Average human questions that make sense	76,17	
Average automatic questions that make sense	80,00	
The human questions are detected as human questions on average	57,87	
The human questions are detected as automatic questions on average	42,13	
The automatic questions are detected as automatic questions on average	41,28	
The automatic questions are detected as human questions on average	58,72	

Now, we want to compare the quantity of our StatMetaQA dataset and other QA datasets that were built in previous work. This comparison is displayed in Table 9. It appears that the size of StatMetaQA dataset is bigger compared to the dataset created by Bondarenko et al. [8], Bojic et al. [3], and Soni et al. [4]; and almost similar to the dataset created by Lal et al. [9]. Therefore, we can conclude that in general, the quantity of StatMetaQA dataset is comparable to that of other QA datasets in previous work.Table 9 The comparison of the quantity of StatMetaQA against other QA datasets.

Table 9Author	Domain	Dataset size	
Rajpukar et al. [7]	general / open domain (Wikipedia)	More than 100,000 questions	
Lal et al. [9]	short narrative (focus on “Why” questions)	30,000 question-answer pair	
Bojic et al. [3]	sleeping training case	7,005 passage which consists of 4,250 single annotations and 750 overlap annotations.	
Soni et al. [4]	radiology report	6,148 question-answer pairs which come from 1,009 report documents.	
Bondarenko et al. [8]	causal question	1,000 random question which is come from the collection of QA dataset (taken 100 question from each dataset)	
Lu et al. [10]	school article	100,000 human-annotated context question–answer triples that are collected from 1,825 articles.	
Ours	statistical metadata	28,863 question-answer pairs (7,353 human questions and 21,510 automatic questions) that are collected from 1,000 documents.	

Limitations

We have some limitations on the validation part, especially in automatic question. In automatic question, only the test set has been validated completely. It happens because of time and effort issues.

Ethics Statement

The authors have read and follow the ethical requirements.

CRediT Author Statement

Nur Rachmawati: Conceptualization, Data Curation, Formal analysis, Investigation, Methodology, Resources, Software, Validation, Writing - original draft Evi Yulianti: Conceptualization, Project Administration, Funding Acquisition, Supervision, Formal analysis, Methodology, Writing – review & editing.

Data Availability

StatMetaQA (Original data) (Github).

StatMetaQA (Original data) (StatMetaQA).

Acknowledgments

This research was funded by the 10.13039/501100014952 Directorate of Research and Development , 10.13039/501100006378 Universitas Indonesia , under Hibah PUTI Pascasarjana 2023 (Grant No. NKB-020/UN2.RST/HKP.05. 00/2023).

Declaration of Competing Interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
==== Refs
References

1 P. Rajpurkar, R. Jia, and P. Liang, “Know what you don't know: unanswerable questions for SQuAD.” arXiv, Jun. 11, 2018. Accessed: Feb. 05, 2024. [Online]. Available: http://arxiv.org/abs/1806.03822.
2 Pietsch M. Möller T. Kostic B. Risch J. Pippi M. Jobanputra M. Zanzottera S. Cerza S. Blagojevic V. Stadelmann T. Soni T. Lee S. Haystack: the end-to-end NLP framework for pragmatic builders [Computer software] https://github.com/deepset-ai/haystack
3 Bojic I. SleepQA: a health coaching dataset on sleep for extractive question answering Proceedings of the 2nd Machine Learning for Health symposium Nov. 2022 PMLR 199 217 Accessed: Feb. 05, 2024. [Online]. Available: https://proceedings.mlr.press/v193/bojic22a.html
4 Soni S. Gudala M. Pajouhi A. Roberts K. RadQA: a question answering dataset to improve comprehension of radiology reports Calzolari N. Béchet F. Blache P. Choukri K. Cieri C. Declerck T. Goggi S. Isahara H. Maegaard B. Mariani J. Mazo H. Odijk J. Piperidis S. Proceedings of the Thirteenth Language Resources and Evaluation Conference Jun. 2022 European Language Resources Association Marseille, France 6250 6259 Accessed: Feb. 05, 2024. [Online]. Available: https://aclanthology.org/2022.lrec-1.672
5 Yang Z. Choi J.D. FriendsQA: open-domain question answering on TV show transcripts Proceedings of the 20th Annual SIGdial Meeting on Discourse and Dialogue 2019 Association for Computational Linguistics Stockholm, Sweden 188 197 10.18653/v1/W19-5923
6 Ju Y. Wang W. Zhang Y. Zheng S. Liu K. Zhao J. CMQA: a dataset of conditional question answering with multiple-span answers Calzolari N. Huang C.-R. Kim H. Pustejovsky J. Wanner L. Choi K.-S. Ryu P.-M. Chen H.-H. Donatelli L. Ji H. Kurohashi S. Paggio P. Xue N. Kim S. Hahm Y. He Z. Lee T.K. Santus E. Bond F. Na S.-H. Proceedings of the 29th International Conference on Computational Linguistics Oct. 2022 International Committee on Computational Linguistics Gyeongju, Republic of Korea 1697 1707 Accessed: Jul. 01, 2024. [Online]. Available: https://aclanthology.org/2022.coling-1.146
7 Rajpurkar P. Zhang J. Lopyrev K. Liang P. SQuAD: 100,000+ questions for machine comprehension of text Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing 2016 Association for Computational Linguistics Austin, Texas 2383 2392 10.18653/v1/D16-1264
8 Bondarenko A. CausalQA: a benchmark for causal question answering Calzolari N. Huang C.-R. Kim H. Pustejovsky J. Wanner L. Choi K.-S. Ryu P.-M. Chen H.-H. Donatelli L. Ji H. Kurohashi S. Paggio P. Xue N. Kim S. Hahm Y. He Z. Lee T.K. Santus E. Bond F. Na S.-H. Proceedings of the 29th International Conference on Computational Linguistics Oct. 2022 International Committee on Computational Linguistics Gyeongju, Republic of Korea 3296 3308 Accessed: Feb. 05, 2024. [Online]. Available: https://aclanthology.org/2022.coling-1.291
9 Lal Y.K. Chambers N. Mooney R. Balasubramanian N. TellMeWhy: a dataset for answering why-questions in narratives Findings of the Association for Computational Linguistics: ACL-IJCNLP 2021 2021 Association for Computational Linguistics 596 610 10.18653/v1/2021.findings-acl.53 Online:
10 P. Lu et al., “Learn to explain: multimodal reasoning via thought chains for science question answering.” arXiv, Oct. 17, 2022. Accessed: Feb. 05, 2024. [Online]. Available: http://arxiv.org/abs/2209.09513.
