
==== Front
Data Brief
Data Brief
Data in Brief
2352-3409
Elsevier

S2352-3409(24)00797-2
10.1016/j.dib.2024.110833
110833
Data Article
Acquisition and processing of Motor Imagery and Motor Execution Dataset (MIMED) for six movement activities
Wirawan I Made Agus imade.aguswirawan@undiksha.ac.id
a⁎
Maneetham Dechrit b
Darmawiguna I Gede Mahendra a
Niyomphol Arnon c
Sawetmethikul Pakornkiat d
Crisnapati Padma Nyoman b
Thwe Yamin b
Agustini Ni Nyoman Mestri e
a Data Science Lab, Engineering and Vocational Faculty, Universitas Pendidikan Ganesha, Udayana Street, No. 11 Singaraja, Bali, 81116, Indonesia
b Mechatronics Engineering lab, Faculty of Technical Education, Rajamangala University of Technology Thanyabur, 39 หมู่ที่ 1 ถนน รังสิต - นครนายก Tambon Khlong Hok, Amphoe Khlong Luang, Chang Wat Pathum Thani 12110, Thailand
c Electrical Engineering Lab, Faculty of Technical Education, Rajamangala University of Technology Thanyabur, 39 หมู่ที่ 1 ถนน รังสิต - นครนายก Tambon Khlong Hok, Amphoe Khlong Luang, Chang Wat Pathum Thani 12110, Thailand
d Electrical and Automation Systems Engineeeing, Faculty of Technical Education, Rajamangala University of Technology Thanyabur, 39 หมู่ที่ 1 ถนน รังสิต - นครนายก Tambon Khlong Hok, Amphoe Khlong Luang, Chang Wat Pathum Thani 12110, Thailand
e Medicine Faculty, Universitas Pendidikan Ganesha, Udayana Street, No. 11 Singaraja, Bali, 81116, Indonesia
⁎ Corresponding author. imade.aguswirawan@undiksha.ac.id
14 8 2024
10 2024
14 8 2024
56 11083311 7 2024
5 8 2024
6 8 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
The MIMED dataset is a dataset that provides raw electroencephalogram signal data for activities: raising the right-hand, lowering the right-hand, raising the left-hand, lowering the left-hand, standing, and sitting. In addition to raw data, this dataset provides feature data that undergoes a baseline reduction process. The baseline reduction process is a process to increase the value of EEG signal features. The feature values ​​of the enhanced EEG signal can be easily recognized in the classification process. The device used is Emotiv Epoc X, which consists of 14 channels. Participants involved in this experiment were 30 students from the Bali region in Indonesia. Four recording scenarios were carried out on the first day and four further scenarios on the second day. Two datasets were obtained based on the recording scenario: the motor movement and image datasets. The duration of motor execution is 40 minutes, while motor imagery is 8 minutes for each scenario.

Keywords

Motor imagery
Motor execution
Electroencephalogram
Baseline reduction
==== Body
pmcSpecifications TableSubject	Computer Science, Biological sciences, Neuroscience, Engineering, Health, and medical sciences.	
Specific subject area	Artificial Intelligence, Human-Computer Interaction, Signal Processing, Bioinformatics, Biotechnology, Neuroscience: Sensory Systems, Biomedical Engineering, Orthopaedics, Physical Therapy and Rehabilitation.	
Data format	Raw data (.edf, and .mat), Analysis and processing data (.mat)	
Type of data	Electroencephalogram (EEG)	
Data collection	1) The number of participants was 30 (16 men and 14 women) from several regions in Bali, Indonesia.

2) Each participant will undergo a recording process in two stages, namely morning and evening, on two different days.

3) Each stage consists of 2 scenarios, each scenario being 12 minutes long. Participants are asked to perform six activities (movement execution) repeated randomly for the first 10 minutes in each scenario. Then, in the last 2 minutes, each participant was asked to imagine six activities (motor imagery). This scenario is repeated the next day.

4) The six activities are raising the right-hand, lowering the right-hand, raising the left-hand, lowering the left-hand, standing, and sitting.

5) The EEG recording tool uses Emotiv Epoc x.

	
Data source location	Data was collected at the Data Science Laboratories, Engineering and Vocational Faculty, Universitas Pendidikan Ganesha. Location coordinates -8.116087176153865, 115.0877732131544.	
Data accessibility	Repository name: Mendeley Data
Data identification number: 10.17632/zs25xxjkm9.3
Direct URL to data: https://data.mendeley.com/datasets/zs25xxjkm9/3	

1 Value of the Data

Motor Imagery and Motor Execution Dataset (MIMED) is a data collection related to six motor activities based on electroencephalogram signals. This dataset is important to study because:1. The dataset provided consists of the motor execution (ME) and motor imagery (MI) datasets. Providing ME and MI datasets is essential to ensure the EEG signal pattern for movements performed internally matches the actual movement. Therefore, it is necessary to provide this dataset [1].

2. Most of the ME and MI datasets that are publicly available are only for arm movements and a few leg movements, so it is essential to provide a dataset that can represent the movements of raising the right-hand, lowering the right-hand, raising the left-hand, lowering the left-hand, standing, and sitting. These six movements are essential to study because they are related to the patient's rehabilitation process [2]

3. EEG signals contain five frequency bands: delta, theta, alpha, beta, and gamma. Four of these are closely related to human movement activities. The four frequency bands are theta, alpha, beta, and gamma. Some currently available datasets do not provide four frequency bands of EEG signals. This research will provide four frequency bands of EEG signal features [1].

4. EEG signals are greatly affected by interference. This interference causes the EEG signal to be unable to represent patterns according to the activity being carried out. Several researchers have studied baseline reduction processes to overcome this problem. The baseline reduction approach can represent EEG signal patterns based on the activities carried out, thus having an impact on increasing accuracy [3]. This approach uses the average of the baseline signal feature values ​​to reduce the trial signal features. Therefore, this data set provides average data of feature values ​​in the baseline signal and the trial EEG signal. Providing these feature data constitutes a contribution to the proposed dataset.

2 Data Description

The participants sampled in this study were 30 healthy people, 16 men, and 14 women, from students in Universitas Pendidikan Ganesha. The determination of 30 participants was based on the normal distribution of the data [1,4]. Recording participant EEG signals using Emotiv Epoc X. Emotiv Epoc X is a light and easy-to-use EEG signal recording tool. This tool is safe for use by the public. The description of this tool has been presented on the web page https://www.emotiv.com/epoc-x/. Fig. 1 shows the Emotiv Epoc X tool.Fig 1 Emotiv Epoc X device.

Fig 1

Channel placement uses International 10–20 system. The following is the position of EEG channel placement, as in Fig. 2.Fig 2 The 14 EEG channels' position follows the system 10-20 international standard.

Fig 2

In Fig. 2, NASION represents the front of the skull, precisely in the upper center of the forehead, while INION shows the lower back. Meanwhile, F, T, P, and O represent the head's frontal, temporal, parietal, occipital, and central parts. In addition, channel AF is a channel placed between Fp and F, while FC is a channel placed between F and C. The channel placement position in this experiment uses the International 10–20 system. These values represent percentages (%) of the distance between the NASION and INION channels. Table 1 presents the sequence of EEG channels stored in the data set. Table 1 presents the sequence of EEG channels stored in the data set.Table 1 Sequence of recording data based on EEG channel.

Table 1No	Channels	No	Channels	
1	AF3	8	O2	
2	F7	9	P8	
3	F3	10	T8	
4	FC5	11	FC6	
5	T7	12	F4	
6	P7	13	F8	
7	01	14	AF4	

Based on the proposed contribution, three data types are produced: raw data, analysis data, and processing data. The overall data hierarchy can be presented in Fig. 3.Fig 3 The hierarchy of MIMED datasets.

Fig 3

The three data types produced will be described in detail in the following explanation.

2.1 Raw data

Raw data from the Emotiv Epoc X tool is of type .edf (European Data Format). This data format (year-month-day-participant_no-scenario_no-day_no.edf) is then converted into PX-Y-Z.mat format. Data that has been converted to .mat is stored in the Matlab Dataset.zip. The X symbol represents the number of participants, the Y symbol represents the number of scenarios, and the Z symbol represents the number of day recordings. The code for data conversion is attached to the EDF-to-Mat.zip. Raw EEG signal data has been generated for each scenario and each day (see Fig. 4).Fig 4 The snippet of four variables in raw data for the first participant (P01-1-1.mat) in the first scenarios on the first day.

Fig 4

Once converted to .mat, the raw data produces four variables: channel, Fs, and signal. In this variable, there are 58 rows and 97792 columns of data. The length of the column represents the length of the recording duration (±12 minutes), while the size of the row represents the amount of data per channel. However, the data from rows 5th to 18th only represent EEG signals. Furthermore, of the 56 channels available in the channel variables, only channels from columns 5th to 18th represent 14 EEG channels (according to Table 1).

2.2 Analysis data

All movement activities were accumulated per participant from the raw data obtained. Six movements are collected, such as raising and lowering the left-hand, raising and lowering the right-hand, and standing and sitting movements, which will be combined per participant into one motor execution dataset (in folder Motor Execution.zip). Fig. 5 presents an example of the ME dataset for the first participant for all activities and scenarios.Fig 5 Analysis data for motor execution (ME) in the first participant (P01.mat) for all activities and all scenarios.

Fig 5

This dataset contains five variables: channel, Fs, join_data1, labels_selfassessment, and subject. The channel variable represents the number of EEG channels used, which is 14. The Fs variable is a sampling frequency of 128 Hz for one second of an EEG signal. The subject variable is the subject of the participant. The variable of joined_data1 represents EEG signal data for twenty-four scenarios (8 recordings x 3 activities), which consists of several scenarios as in Table 2:Table 2 Description of raising and lowering left-hand activity, raising, and lowering the right-hand, and standing and sitting movements in twenty-four columns.

Table 2No. columns	Activities	Values	
1st	raising left-hand activity for the first scenario on the first day	11904×14	
2nd	lowering left-hand activity for the first scenario on the first day	11904×14	
3rd	raising left-hand activity for the second scenario on the first day	11904×14	
4th	lowering left-hand activity for the second scenario on the first day	11904×14	
5th	raising left-hand activity for the first scenario on the second day	11904×14	
6th	lowering left-hand activity for the first scenario on the second day	11904×14	
7th	raising left-hand activity for the second scenario on the second day	11904×14	
8th	lowering left-hand activity for the second scenario on the second day	11904×14	
9th	raising right-hand activity for the first scenario on the first day	13824×14	
10th	lowering right-hand activity for the first scenario on the first day	13824×14	
11th	raising right-hand activity for the second scenario on the first day	13824×14	
12th	lowering right-hand activity for the second scenario on the first day	13824×14	
13th	raising right-hand activity for the first scenario on the second day	13824×14	
14th	lowering right-hand activity for the first scenario on the second day	13824×14	
15th	raising right-hand activity for the second scenario on the second day	13824×14	
16th	lowering right-hand activity for the second scenario on the second day	13824×14	
17th	standing activity for the first scenario on the first day	13824×14	
18th	sitting activity for the first scenario on the first day	13824×14	
19th	standing activity for the second scenario on the first day	13824×14	
20th	sitting activity for the second scenario on the first day	13824×14	
21th	standing activity for the first scenario on the second day	13824×14	
22th	sitting activity for the first scenario on the second day	13824×14	
23th	standing activity for the second scenario on the second day	13824×14	
24th	sitting activity for the second scenario on the second day	13824×14	

The value in join_data1 is 11904×14 for each scenario's raising and lowering left-hand movement. Meanwhile, for raising and lowering the right-hand and standing and sitting movements, the value is 13824×14 for each scenario. The value 11904 represents the length of the EEG signal data for 93 seconds (93 seconds x 128 Hz). Meanwhile, 13824 means the length of EEG signal data for 108 seconds (108 s x 128 Hz. In the first 3 seconds of raising and lowering left-hand activity, raising and lowering the right-hand, and standing and sitting movements are EEG signals that represent a relaxed condition (baseline signal). Meanwhile, the following four seconds until the last second represent movement activity (trial signal). The value 14 represents 14-channel EEG. Data from each cell was collected by sorting the raw EEG signal data based on the duration of the stimulus media given to participants for 12 minutes (Media Stimulus English Ver.mp4). This stimulus media stimulates all activities in one recorded scenario for each participant. Each participant will be recorded in four scenarios on two days. Each sign of activity that appears on this media lasts 5 seconds. A recap of the duration of all scenarios is presented in the recap_duration.xlsx file. The variable labels_selfassessment represents the labels of the participants' twenty-four movement scenarios. The values [1,1,1] represent raising the left-hand, while the values [0,0,0] represent lowering the left-hand. The value [1,1,0] represents the activity of raising the right-hand, and the value [0,0,1] represents lowering the right-hand. The last value [1,0,0] represents standing activity, and [0,1,1] represents sitting activity.

Furthermore, the variables in the MI dataset are almost identical to the motor movement (ME) dataset. However, the MI dataset does not contain the labels_selfassessment variable. The motor imagery (in the folder Motor Imagery.zip) dataset provides a dataset for testing artificial intelligence models predicting activities based on training data from motor movements. Fig. 6 presents an example of MI data representation for standing and sitting in the first participant (P01.mat).Fig 6 Motor Imagery (MI) analysis data for standing and sitting activities for the first participant (P01.mat) in all scenarios.

Fig 6

The activities of standing and sitting for the imagery dataset have four cells in joined_data. These four cells represent four recording scenarios in two days, which consist of:1. The first cell represents standing and sitting activity for the first scenario on the first day.

2. The second cell represents standing and sitting activity for the second scenario on the first day.

3. The third cell represents standing and sitting activity for the first scenario on the second day.

4. The fourth cell represents standing and sitting activity for the second scenario on the second day.

The same scenario also exists for raising and lowering the right-hand and left-hand motor imagery. Each cell in the rising and lowering of the left-hand and standing and sitting imagery dataset has a joined_data value of 2432×14 while raising and lowering the right-hand imagery has a value of 1152×14. The value 2432 represents the EEG signal data for 19 seconds (19 s x 128 Hz). Meanwhile, 1152 represents the amount of EEG signal data for 9 seconds (9 s x 128 Hz). The value 14 represents a 14-channel EEG.

2.3 Processing data

The data that has been analyzed is then processed to obtain feature values. There are several process stages: Segmentation, Decomposition, Normalization, and Feature Extraction. Each process will be explained in detail in the experimental design and methods. The segmentation process separates baseline signal data and trial signals in the variable joined_data1. The decomposition process separates four frequency bands (theta, alpha, beta, and gamma) in the baseline and trial signals. Next, the normalization process normalizes baseline and trial signal data to avoid outliers. The feature extraction process is a process to extract baseline and trial signal data from 128 values ​​into one feature value (1-second feature value). This process results in data processed for each participant (DE_PX.mat). All processed data is stored in the Dataset_1D.zip folder. The following is a snippet of data in Fig. 7.Fig 7 Preprocessing data for motor execution (ME) in the first participant (DE_P01.mat) for all activities and all scenarios.

Fig 7

The processed data has six variables: base_data, combination1_labels, combination2_labels, combination3_labels, data, and data_seconds_list variables. The base_data variable is a variable that stores feature data from the baseline EEG signal. The baseline signal is a signal that represents neutral conditions. In one participant, the EEG baseline signal feature value in the base_data variable is 24×54. The value 24 represents one EEG baseline signal feature value for 24 experiments. This feature value is the average of three seconds of EEG baseline signals.

Meanwhile, the value 56 represents four frequency bands on each channel (total channels are 14). The data variable represents the feature values ​​of the trial signal. The data size of this variable is 2400×56, where 2400 represents the length of the experimental duration for six movement activities in 24 experiments. A detailed description of data length in data variables is presented in Table 3.Table 3 Length of trial signal feature data for each activity.

Table 3No	Activity	Duration of trial signal	Number of trials	
1	Raising left-hand	90 seconds	4 (two scenarios in two days)	
2	Lower left-hand	90 seconds	4 (two scenarios in two days)	
3	Raising right-hand	105 seconds	4 (two scenarios in two days)	
4	Lower right-hand	105 seconds	4 (two scenarios in two days)	
5	Standing	105 seconds	4 (two scenarios in two days)	
6	Sitting	105 seconds	4 (two scenarios in two days)	

Next, the data_seconds_list variable has a size of 1×24. This variable represents the total duration for each experiment (24 experiments). Experiments 1 to 8 have 90 seconds, while experiments 9 to 24 have a duration of 105 seconds (as in Table 3). Python code for processing motor execution data available in 1D_Dataset-Execution.py file

In the processing data for the MI dataset, there are three types of MI data: MI data processing for right-hand raising and lowering activities, MI data processing for left-hand raising and lowering activities, and MI data processing for standing and sitting activities. Fig. 8 presents a snapshot of data analysis variables for the first participant's standing and sitting. Python code for processing motor imagery data available in 1D_Dataset-Imagery.py file.Fig 8 Motor Imagery (MI) processing data for standing and sitting activities for the first participant (DE_P01.mat) in all scenarios.

Fig 8

3 Experimental Design, Materials and Methods

3.1 Determination of participants

The procedure for determining participants was based on research conducted by Koelstra et al. [5], and the study conducted by Miranda Correa et al.[6], as shown in Fig. 9.Fig 9 Flowchart for determining participants.

Fig 9

In this study, 79 people filled out the research questionnaire age range 19 - 24 who were undergraduate students at the Engineering and Vocational Faculty, Universitas Pendidikan Ganesha. Of the 79 participants, 30 were willing and in good health to be recording subjects. Next, 30 subjects were asked to fill out and sign a consent form involving two witnesses who knew the main participants well.

3.2 Arrangement of EEG signal recording infrastructure

It is essential to arrange the supporting infrastructure, including monitoring screen size, lighting, viewing angle, viewing distance, and keeping the room quiet to provide participants with comfort in the recording process and impact the resulting EEG signal data quality. The arrangement of laboratory infrastructure refers to research conducted by Sarma and Barma [7]. The equation can be used to determine the viewing distance between the participant and the monitor layer to determine the viewing distance between the participant and the monitor layer (Eq. (1)).(1) y^=−14+70x1+2x2−0.0015x22+0.46x32

Where the y^ is predicted corresponding viewing distance (in millimeters), x1 represents the size of the TV monitor (in inches), x2 represents the indoor lighting value (in lux), and x3 represents the viewing angle (in degrees °), as illustrated in Fig. 10.Fig 10 Illustration of room arrangement using formula 1.

Fig 10

In preparing the procedures for recording EEG signals in this experiment, referring to articles from research by Koelstra et al. [5], Katsigiannis and Ramzan [8], and Miranda Correa et al. [6]. The EEG signal recording process also uses the Emotiv Epoc X. EEG signal recording was noninvasive, where the recording device (EEG channel) was attached to the participant's scalp. In several previous studies, the Emotiv Epoc X tool has been widely used for recording EEG signals and is safe [6,8].

3.3 Recording scenarios

The recording process for each participant was carried out for two days. On the first day, it was carried out in the morning, while the second session was in the afternoon. Each participant followed the same two recording scenarios on the first and second days. EEG signal recording on different days was carried out to obtain different EEG signal patterns even in the same participant [9]. This condition is caused by the EEG signal pattern being greatly influenced by individual characteristics, cognition, and other personality traits [3].

Each scenario will produce two types of data: motor execution (ME) data and motor imagery (MI) data. ME data is generated when participants carry out movement activities following the stimulus media displayed on the screen. MI data is produced when participants imagine movements according to the stimulus media on the screen. The total recording duration for each scenario was 12 minutes. The first 10 minutes are ME data, while the last 2 are MI data. Recording ME data is to determine the correctness of data labels related to movement activities. Meanwhile, MI functions as test data for motor movement. So that the resulting EEG signal data pattern can represent movement patterns correctly (ground truth) [1].

Below is a scenario for recording EEG signals in participants who raising and lowering the left-hand, raising and lowering the right-hand, and standing and sitting based on the stimulus media provided. This scenario was carried out in two cycles for each participant, with different recording times and days. In Fig. 11, one scenario of the EEG signal recording process can be presented.Fig 11 EEG signal recording scenario.

Fig 11

Fig. 11 presents the recording of EEG signals for one participant who performs standing, sitting, raising and lowering the left-hand, and raising and lowering the right-hand. The following stages of the recording process can be explained as follows:1. The EEG signal recording process can begin after the sample has been determined and the recording schedule has been set.

2. Check the body temperature of the participant, check the participant's level of fatigue, and provide hand sanitizer. This stage is carried out to ensure the participant's health condition before participating in the experimental process.

3. Fill in the attendance list.

4. In the initial stage, the researcher explains the mechanism of this experiment, such as the conditions and procedures for watching stimulus media.

5. Determine the distance between the participant's seat and the monitor layer (Y). This setting is essential to keep participants comfortable during the EEG signal recording. Determining the distance between participants sitting with the monitor layer is based on research conducted by Sarma and Barma [7].

6. Installation of EEG equipment. The EEG tool used in this research is Emotiv Epoc X with several channels 14.

7. The rules for placing EEG channels use the international 10-20 system following reference journals from researchers Katsigiannis and Ramzan [8] and Miranda Correa et al. [6].

8. EEG recording. After installing the EEG channel, the EEG signal was recorded using Emotiv Epoc X.

9. The following article references research conducted by [5], Katsigiannis and Ramzan [8], Miranda Correa et al. [6], and research conducted by Sarma and Barma[7] it is vital to keep participants in a calm condition. The process to condition calm is carried out by controlling the participant's breathing guided by the researcher, where several stages must be carried out by the participant, namely: inhale for 4 seconds, hold breath for 4 seconds, exhale for 4 seconds, and hold breath for 4 seconds [10]. This process is repeated for approximately 2 minutes. This procedure eliminates the Hawthorne and white coat effects[11]. If the leading participant feels comfortable, then proceed to the next stage.

10. Listen to the stimulus media & participants make movements according to the stimulus media. The stimulus media used is Go/No Go media, which displays signs such as the right-hand, left-hand, feet, and a stop sign. Each sign represents an activity the participant must follow; for example, if the right-hand sign appears, the participant must raise the right-hand, and the same goes for the left-hand sign. If the foot signal appears, participants are required to stand, while at the stop sign, participants are asked to sit. These signs appear randomly and at different time durations on the stimulus media. Through this media, participants can focus on activities such as standing, sitting, raising and lowering their right-hand, and raising and lowering their left-hand according to the signs that appear. Fig. 12 is a screenshot of the stimulus media used.Fig 12 Stimulus media footage (A) Raising the right-hand, (B) Raising the left-hand, (C) Standing, (D) Lowering left-hand, lowering right-hand, or sitting.

Fig 12

One recording was carried out for ± 12 minutes to anticipate participants' tiredness. The duration of this experiment is based on research conducted by [5] and Miranda Correa et al. [6].1. End the EEG signal recording process and remove the EEG device. After the stimulus media has finished watching, the researcher will instruct the participant to end the EEG signal recording process and remove the channel attached to the head.

2. Put participants in a calm condition. Supporting participants were asked to regulate their breath with their eyes closed to condition participants to be calm and relaxed. This process lasted several minutes until the participant stated his condition was calm. This procedure is based on research conducted by Sarma and Barma [7].

3. Check body temperature and check participants' level of fatigue. This stage was carried out to ensure the participant's health condition after participating in the experimental process.

4. The recording is complete. One cycle of the recording process is carried out for ± 12 minutes, where the total recording for this scenario is three recordings, namely morning and afternoon, both on the same day and different days.

This recording process results from raw EEG signal data shown in Fig. 4. After the conversion to .mat, the data analysis process is carried out.

3.4 Data analysis

The converted raw data (.mat) is sorted for each movement activity in different scenarios. This process refers to the recap_duration.xlsx file. The results of this analysis are analytical data for all activities of each participant in Motor Execution (ME) data. A snapshot of the ME data analysis variables is presented in Fig. 5. The same is true for the Motor Imagery (MI) data. Still, this data is presented in three activities: MI data analysis for raising and lowering the right-hand, MI data analysis for raising and lowering the left-hand, and MI data analysis for standing and sitting activities. A snapshot of the ME data analysis variables is presented in Fig. 6.

3.5 Segmentation

Several data processing processes are carried out to produce four frequency bands and reduce interference in the EEG signal, such as segmentation, decomposition, normalization, and feature extraction. Segmentation is carried out to separate baseline and trial EEG signals in data analysis. The 1st to 3rd seconds are used as the baseline signal, while the following 4th second is used as the trial signal. Segmentation will be performed for fourteen channels in twenty-four trials for all participants. An illustration of the baseline and trial signal segmentation process on the AF3 channel is shown in Fig. 13.Fig 13 EEG signal segmentation process on AF3 channel for raising left-hand activity from participant P01.

Fig 13

The green line is the baseline signal, while the red line is the trial signal. Table 4 presents a snippet of Python code for the segmentation process on the ME dataset. This process refers to the predefined representation of analytical data (Fig. 5).Table 4 Code snippet of the segmentation process.

Table 4Line no.	Code	
1	trial_signal = data[:,experiment][0][384:, channel]	
2	base_signal = data[:,experiment][0][:384, channel]	

The baseline signal is in line 1, the 1st – 384th Hz (first 3 seconds). In the second line, the test signal is 385 Hz – until the end.

3.6 Normalization

The normalization process is to overcome outlier values ​​in the baseline and trial signals amplitude. This normalization uses the z-score method [12]. The following is a snippet of the normalization program code in Table 5.Table 5 Code snippet of the z-score normalization process.

Table 5Line no.	Code	
1	def normalization (signal):	
2	return (signal- signal.mean())/signal.std()	

In line 1, the normalization function is defined with the signal parameter. Next, in line 2, the z-score normalization formula is determined. The normalization value will be provided when the normalization function is called (return).

3.7 Decomposition

Based on the frequency band, EEG signals are divided into five types: Delta, Theta, Alpha, Beta, and Gamma. However, in the decomposition process of the baseline and trial signals, only four types of frequencies are used, namely Theta, Alpha, Beta, and Gamma, for each segment and each channel in the EEG baseline signal [13]. Delta frequency (0.5 – 4 Hz) is not used because it does not correlate with motor imagery. After all, in this frequency range, a person is in a deep rest condition [14]. An illustration of the baseline signal decomposition process for each frequency band can be seen in Table 6.Table 6 Code snippet of the baseline signal decomposition for each frequency band.

Table 6Line no.	Code	
1	base_theta = butter_bandpass_filter(base_signal, 4, 8, frequency, order=3)	
2	base_alpha = butter_bandpass_filter(base_signal, 8,14, frequency, order=3)	
3	base_beta = butter_bandpass_filter(base_signal,14,31, frequency, order=3)	
4	base_gamma = butter_bandpass_filter(base_signal,31,45, frequency, order=3)	

3.8 Feature extraction

The feature extraction process is carried out for each second (128 Hz), each frequency band, and each channel. The feature extraction process in this research uses the Differential Entropy (DE) method, shown in Table 7.Table 7 Code snippet of the feature extraction using the DE method.

Table 7Line No.	Code	
1	def compute_DE(signal):	
2	variance = np.var(signal,ddof=1)	
3	Return math.log(2*math.pi*math.e*variance)/2	

In line 1, the compute_DE function with signals is defined as a parameter. The second line is used to calculate the variance value of the signal value using the NumPy (np.var) library. The ddof (Delta Degrees of Freedom) parameter determines the divisor used in the calculation. In line 3, the DE method formula is defined based on the variance value that has been generated. The resulting DE value will be provided when the compute_DE function is called (return).

The entire Python program code for the Segmentation process to Feature Extraction (processing data) for the ME dataset is presented in the 1D Dataset-Execution.py file. In contrast, the MI dataset is presented in the 1D_Dataset-Imagery.py program. A snapshot of the ME dataset variables is presented in Fig.5, while the MI dataset is presented in Fig. 6.

3.9 Baseline reduction

The baseline reduction process was conducted to test the success of reducing interference on the EEG signal (valuable data no. 4). This process uses the Relative Difference method. In the Relative Difference method, a reduction process is carried out to eliminate interference in the test signal data by a division process. This method code can be seen in Table 8.Table 8 Code snippet of the baseline reduction using the Relative Difference method.

Table 8Line no.	Code	
1	def get_relative_difference (vector1, vector2):	
2	return vector1/vector2	

In line 1, the get_relative_difference function is defined with the function parameters vector1 and vector2. The vector1 value is the DE feature value for the trial signal, while vector2 is the average DE feature value for the baseline signal for all channels and four frequency bands in 1 second. In the 2nd line, vector1 is divided by vector2 (Relative Difference). This reduction value will be given when the get_relative_difference function is called (return). The trial signal feature pattern from the baseline reduction process can enhance the EEG trial signal feature. By improving this pattern, the classification method can recognize patterns better. This process is essential because the EEG signal has low signal characteristics [3]. Fig. 14 present a visualization of the trial signal feature patterns without and with the baseline reduction process.Fig 14 visualization of the trial signal feature patterns without and with the baseline reduction process.

Fig 14

3.10 Classification

A classification process is carried out to measure the success of increasing EEG signal data. The resulting accuracy, precision, recall, and F1 Score values ​​are used as a reference for the success of the proposed data processing. This research will use the Decision Tree classification method (file code DecisionTree_6Class.py). The Decision Tree method's node depth is 20. In addition, the Support Vector Machine method is also used to measure the performance of the baseline reduction approach on this dataset (file code SVM_6Class.py). Fig. 15 shows the comparative accuracy values ​​for each participant between with and without baseline reduction on the ME dataset.Fig 15 Accuracy comparison diagram between with and without baseline reduction for each participant in the ME dataset.

Fig 15

Overall, the accuracy, precision, recall, and F1 score values ​​between with and without baseline reduction on the ME dataset can be presented in Table 9.Table 9 Comparison of accuracy, precision, recall, and F1 score between with and without baseline reduction.

Table 9No.	Method	Approach	Accuracy	Precision	Recall	F1 score	
1	Decision Tree	With baseline	75.42 %	75.65 %	75.31 %	75.14 %	
2	Decision Tree	Without baseline	20.08 %	20.11 %	19.94 %	19.45 %	
3	Support Vector Machine	With baseline	83.23 %	84.85 %	83.19 %	83.21 %	
4	Support Vector Machine	Without baseline	23.85 %	23.86 %	23.47 %	21.87 %	

The results of the baseline reduction approach are proven to increase accuracy, precision, recall, and F1 score. These results align with previous studies that applied a baseline reduction approach to trial EEG signals in different case studies [13,[15], [16], [17], [18]]. Recapitulation of the classification results and the ethical clearance are inserted in the Supplementary_file.zip.

Limitations

This dataset was collected to obtain raw data, data analysis, and data processing for electroencephalogram signals related to six motor movement activities and imaging. By providing this dataset, it is hoped that future researchers can study various optimal methods for preprocessing, feature extraction, and classification to recognize six motor execution and motor imagery activities. Even though this data has been proven to help machine learning models increase accuracy, several further studies are essential, such as selecting the right channel to represent motor execution and imagery. This study will produce a dataset that optimally represents motor execution and imagery activity. Apart from that, further research is essential to study more movements, such as movements of fingers, toes, and other activities, so that Motor Imagery and Motor Execution Datasets that represent various human activities will be collected.

Ethics Statement

Before the start of the experiment, all participants provided written informed consent according to the World Association Declaration of Helsinki. The Ethics Committee, Universitas Pendidikan Ganesha, number 038/UN.48.24.11/LT/2023, approved all scenarios in this study (Ethical Clearance.pdf).

CRediT authorship contribution statement

I Made Agus Wirawan: Conceptualization, Methodology, Validation, Formal analysis, Investigation, Resources, Data curation, Writing – original draft, Visualization, Writing – review & editing. Dechrit Maneetham: Methodology, Validation, Formal analysis, Writing – review & editing, Investigation. I Gede Mahendra Darmawiguna: Methodology, Validation, Formal analysis, Writing – review & editing, Investigation. Arnon Niyomphol: Methodology, Validation, Formal analysis, Writing – review & editing, Investigation. Pakornkiat Sawetmethikul: Methodology, Validation, Formal analysis, Writing – review & editing, Investigation. Padma Nyoman Crisnapati: Methodology, Validation, Formal analysis, Writing – review & editing, Investigation. Yamin Thwe: Methodology, Validation, Formal analysis, Writing – review & editing, Investigation. Ni Nyoman Mestri Agustini: Methodology, Investigation, Writing – review & editing, Supervision.

Appendix Supplementary materials

Image, application 1

Data Availability

The MIMED dataset (Original data) (Mendeley Data).

Acknowledgements

This work is partially supported by Funding for the 2023 fiscal year, No: 1175/UN48.16/LT/2023, from the Institute for Research and Community Service at the Universitas Pendidikan Ganesha. The authors also gratefully acknowledge the helpful comments and suggestions of the reviewers, which have improved the presentation.

Declaration of Competing Interest

There is no conflicts of interest.

Supplementary material associated with this article can be found, in the online version, at doi:10.1016/j.dib.2024.110833.
==== Refs
References

1 Gwon D. Won K. Song M. Nam C.S. Jun S.C. Ahn M. Review of public motor imagery and execution datasets in brain-computer interfaces Front. Media SA 2023 10.3389/fnhum.2023.1134869
2 de Sousa D.G. Two weeks of intensive sit-to-stand training in addition to usual care improves sit-to-stand ability in people who are unable to stand up independently after stroke: a randomised trial J. Physiother. 65 3 2019 152 158 10.1016/j.jphys.2019.05.007 31227279
3 Wirawan I.M.A. Wardoyo R. Lelono D. The challenges of emotion recognition methods based on electroencephalogram signals: a literature review Int. J. Electr. Comput. Eng. (IJECE) 12 2 2022 1508 10.11591/ijece.v12i2.pp1508-1519
4 Bekele W.B. Ago F.Y. Sample size for interview in qualitative research in social sciences: a guide to novice researchers Res. Educ. Policy Manage. 4 1 2022 42 50 10.46303/repam.2022.3
5 Koelstra S. DEAP: a database for emotion analysis; using physiological signals IEEE Trans. Affect. Comput. 3 1 2012 18 31 10.1109/T-AFFC.2011.15
6 Miranda Correa J.A. Abadi M.K. Sebe N. Patras I. AMIGOS: a dataset for affect, personality and mood research on individuals and groups IEEE Trans. Affect. Comput. 2018 1 14 10.1109/TAFFC.2018.2884461 no. i, pp
7 Sarma P. Barma S. Review on stimuli presentation for affect analysis based on EEG IEEE Access. 8 2020 51991 52009 10.1109/ACCESS.2020.2980893
8 Katsigiannis S. Ramzan N. Dreamer: a database for emotion recognition through EEG and ECG signals from wireless low-cost off-the-shelf devices IEEE J. Biomed. Health Inform. 22 1 2018 98 107 10.1109/JBHI.2017.2688239 28368836
9 Plucińska R. Jędrzejewski K. Malinowska U. Rogala J. Leveraging multiple distinct EEG training sessions for improvement of spectral-based biometric verification results Sensors 23 4 2023 10.3390/s23042057
10 James Nestor, Breath Cara Bernafas dengan Benar. PT Gramedia Pustaka Utama, 2021.
11 Jimenez-Molina A. Retamal C. Lira H. Using psychophysiological sensors to assess mental workload during web browsing Sensors 18 2 2018 1 26 10.3390/s18020458
12 Aytekin A. Comparative analysis of normalization techniques in the context of MCDM problems Decis. Making: Appl. Manage. Eng. 4 2 2021 1 25 10.31181/dmame210402001a
13 Wirawan I.M.A. Wardoyo R. Lelono D. Kusrohmaniah S. Modified weighted mean filter to improve the baseline reduction approach for emotion recognition Emerg. Sci. J. 6 6 2022 1255 1273 10.28991/ESJ-2022-06-06-03
14 Kumar J.S. Bhuvaneswari P. Analysis of electroencephalography (EEG) signals and its categorization - a study Procedia Eng. 38 2012 2525 2536 10.1016/j.proeng.2012.06.298
15 Wirawan I.M.A. Wardoyo R. Lelono D. Kusrohmaniah S. Asrori S. Comparison of baseline reduction methods for emotion recognition based on electroencephalogram signals 2021 Sixth International Conference on Informatics and Computing (ICIC) Yogyakarta 2021 IEEE 1 7 10.1109/ICIC54025.2021.9632948
16 Yang Y. Wu Q. Qiu M. Wang Y. Chen X. Emotion recognition from multi-channel EEG through parallel convolutional recurrent neural network Proceedings of the International Joint Conference on Neural Networks 2018 1 7 10.1109/IJCNN.2018.8489331 2018-July, no. July
17 Yang Y. Wu Q. Fu Y. Chen X. Continuous convolutional neural network with 3D input for EEG-based emotion recognition International Conference on Neural Information Processing 2018 Springer International Publishing 433 443 10.1007/978-3-030-04239-4_39
18 Liu Y. Multi-channel EEG-based emotion recognition via a multi-level features guided capsule network Comput. Biol. Med. 123 2020 103927 10.1016/j.compbiomed.2020.103927 no. March
