
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)12620-5
10.1016/j.heliyon.2024.e36589
e36589
Review Article
Deep learning-based human body pose estimation in providing feedback for physical movement: A review
Tharatipyakul Atima a1
Srikaewsiew Thanawat b1
Pongnumkul Suporn suporn.pongnumkul@nectec.or.th
a⁎
a National Electronics and Computer Technology Center (NECTEC), Pathumthani 12120, Thailand
b Suranaree University of Technology, Nakhonratchasima 30000, Thailand
⁎ Corresponding author. suporn.pongnumkul@nectec.or.th
1 The first and second author contributed equally.

26 8 2024
15 9 2024
26 8 2024
10 17 e3658916 6 2023
19 8 2024
19 8 2024
© 2024 The Author(s)
2024
https://creativecommons.org/licenses/by-nc/4.0/ This is an open access article under the CC BY-NC license (http://creativecommons.org/licenses/by-nc/4.0/).
Pose estimation has various applications in analyzing human body movement and behavior, including providing feedback to users about their movements so they can adjust and improve their movement skills. To investigate the current research status and possible gaps, we searched Scopus and Web of Science for articles that (1) human ‘body’ pose estimation is used and (2) user movement is assessed and communicated. We used either a bottom-up or top-down approach to analyze 45 articles for methods used to estimate human body pose, assess movement, provide feedback to users, as well as methods to evaluate them. Our review found that pose estimation systems typically used CNNs while movement assessment methods varied from mathematical formulas or models, rule-based approaches, to machine learning. Feedback was primarily presented visually in verbal forms and nonverbal forms. The experiments to evaluate each part ranged from the use of public datasets to human participants. We found that pose estimation libraries play an important role in the advancement of this field. Nevertheless, the effectiveness and factors for choosing movement assessment methods for a new context are still unclear. In the end, we suggest that studies about feedback prioritization and erroneous feedback are needed.

Graphical abstract

Keywords

Pose estimation
Movement assessment
Augmented feedback
Physical movement
Review
==== Body
pmc1 Introduction

Human pose estimation is the process of determining the positions and orientations of specific human body parts, such as the head, shoulders, arms, and legs. Pose estimation has a wide range of applications in fields that involve the analysis and understanding of human movement and behavior. For example, in human-computer interaction (HCI), the estimation of hand pose could enable gesture-based natural interaction [1]. In healthcare, pose estimation could help monitor and analyze the movements and posture of patients in rehabilitation or therapy settings (e.g., [2], [3]). While doing a physical movement, such as rehabilitation or sports training, pose estimation could be used for various purposes [4], [5], [6], [7], and one of them is to provide feedback about the movement to a user.

Feedback is an important aspect of physical movement learning and training, as it allows individuals to assess their performance and adjust to improve their movement skills. Feedback could be task-intrinsic, from the sensory system of a performer, or augmented, from an external source to a performer. Traditionally, a coach or an instructor provides augmented feedback as verbal cues or corrections during a training session. For example, a yoga instructor may tell a learner who does an inverted U-shape downward dog pose to do a V-shape instead. With automated pose estimation, a computer system could assess user movement and provide augmented feedback. Da Gama et al. [8], for example, reviewed 31 articles that used Kinect to assess and provide feedback on motor rehabilitation. The review suggested development possibilities and further studies of using Kinect to rehabilitate at home. While there is a large body of knowledge on using Kinect or other sensors for pose estimation and physical movement applications, recent advances in pose estimation using a web camera open new opportunities in this field.

The utilization of widely available hardware like web cameras signifies the integration of pose estimation techniques into a broader spectrum of physical movement applications, such as in performing arts [9], [10], sports [11], [12], [13], [14], [15], various types of exercise [16], [17], [18], [19], [20], [21], [22], [23], [24], [25], [26], and healthcare [27], [2]. As a result, several reviews have been published (e.g., [5], [6], [7], [28]). While the existing reviews serve the purpose of identifying prevailing methods, challenges, and potential future directions in the realm of pose estimation for physical movement applications, none of them focus on how they are used for assessing a movement and providing feedback to a user, which is an important aspect of physical movement applications and an indication of how pose estimation is adopted in practice.

This paper presents a review of recent articles that use deep learning-based human pose estimation to assess user movement and provide feedback on the user's physical movement. We divided the physical movement applications into three parts or modules: (1) pose estimation that detects keypoints of a human body; (2) movement assessment that uses keypoints or the motion of keypoints to evaluate the quality or the value of the movement; and (3) augmented feedback presentation that communicates the results of other modules to users.

The goal of our paper is threefold:• to investigate methods used for pose estimation, movement assessment, and augmented feedback presentation;

• to investigate how those methods were experimented or evaluated;

• to discuss the current research status of each part, possible gaps, and future research directions.

We organize this article as follows. Section 2 gives an overview of the current topics, followed by Section 3 that presents the review methodology. Section 4 and 5 presents the results and discussions. We conclude this article in Section 6.

2 Background and related work

This section clarifies the terms used in this paper and summarizes related articles. Multiple review and survey papers exist on pose estimation for physical movement applications, smart technology for movement assessment, and augmented feedback and their effectiveness. The review and survey papers are summarized in Table 1.Table 1 Examples of related review and survey papers. No. is the number of papers that the article reviewed. There are cases where articles do not explicitly state the number of papers being reviewed. In such cases, the asterisk (*) indicates that we counted the number of papers from the article's table. The double asterisk (**) indicates that we counted the number of papers from the references, which could be more than the number of articles that actually were reviewed.

Table 1Paper	Year	Domain	No.	Focus	Note	
[29]	'20	General	74**	Pose estimation	2D	
[30]	'21	General	81**	Pose estimation	3D, Markerless	
[31]	'21	General	150**	Pose estimation	3D, Deep learning-based	
[32]	'21	General	124**	Pose estimation	Deep learning-based	
[33]	'23	General	361**	Pose estimation	Deep learning-based	
[34]	'23	General	206**	Pose estimation	Deep learning-based, 2D	
[5]	'21	Training assistance	8	Pose estimation		
[6]	'21	Human Health and Performance	53*	Pose estimation	Assessment as a part of improvement	
[7]	'21	Sports and physical exercise	20	Pose estimation	Camera-based	
[28]	'22	Rehabilitation	62**	Pose estimation	Computer vision and IMU-based	
[35]	'20	Motor learning	72**	Movement assessment	Machine learning-based	
[36]	'20	Sports training	109	Movement assessment	Intelligent data analysis methods	
[37]	'22	Rehabilitation, sports, wellbeing	40*	Movement assessment	Machine learning-based	
[38]	'20	Healthcare	41	Feedback presentation	Real-time feedback	
[39]	'14	Exercise	52**	Feedback presentation		
[40]	'21	Physical education	23	Feedback presentation		
[41]	'21	Physical education	11	Feedback presentation	Video-based visual feedback	

2.1 Pose estimation for physical movement applications

In this paper, we use the term pose estimation for a process for determining the position of key points of a person's body from a given image or video. The methods for pose estimation can be classified into two-dimensional (2D) [29] and three-dimensional (3D) [31], [30], and could be further classified based on the number of people in the image (single-person or multi-person) or approaches (e.g., top-down or bottom-up) [32]. In addition to the aforementioned review articles, Zheng et al. [33] recently reviewed deep learning-based human pose estimation, datasets, evaluation metrics, and applications. Chen et al. [34] covered methodological frameworks, common benchmark datasets, evaluation metrics, and performance comparisons of 2D human pose estimation articles.

Pose estimation has a wide range of applications. For physical movement learning or training, a lot of works have been proposed, and as a result, there exist several articles that review them. For instance, Difini et al. [5] systematically reviewed the usage of human pose estimation for training assistance. They identified 8 articles and investigated the challenges, the technology used, the application context, and the accuracy of human pose estimation. Stenum et al. [6] reviewed the applications of pose estimation in three domains: motor and non-motor development; human performance optimization, injury prevention, and safety; and clinical motor assessment. Their identified application limitations included occlusions, limited training data, capture errors, positional errors, and limitations of recording devices. Their identified application limitations included user-friendliness (i.e., set-up time, delayed results, and programming and training requirements), outcome measure challenges, limited hardware infrastructure, technology challenges, and lack of validation and feasibility data. Badiola-Bengoa and Mendez-Zorrilla [7] reviewed 20 articles on pose estimation in sports and physical exercise. They discussed the available data, methods, performance, opportunities, and challenges. Niu et al. [28] surveyed articles that used inertial measurement units (IMU) and computer vision for human pose estimation in rehabilitation applications. They summarized the research status and challenges as well as suggested that the two methods can be combined to get better outcomes.

These reviews are useful for identifying existing methods, challenges, and future directions of using pose estimation in physical movement applications. However, they had not reviewed how those applications assessed movement and provided feedback to users, which is a gap in the research literature. This paper intends to fill this gap.

2.2 Smart technology for movement assessment

We use the term movement assessment as a way to evaluate or estimate the quality of movement to provide feedback to a user. According to this definition, we consider works on, for example, action recognition or movement modeling, if a user could use the result to assess their movement, such as telling whether it is correct or not. The use of Artificial Intelligence to analyze user movement and provide feedback to improve user performance, quality of life, and well-being has been a subject of research for more than a decade. Several reviews have been published to present the level of knowledge on the topic.

Caramiaux et al. [35] conducted a short review of machine learning approaches for motor learning. They focused on motor variability which requires the algorithms to differentiate between new movements and variations from known ones. They identified three types of adaptation: parameter adaptation in probabilistic models, transfer and meta-learning in deep neural networks, and planning adaptation by reinforcement learning. They also discussed challenges for applying the models, including variations of an already-trained skill, an adaptation that involves re-training procedures, and continuous evolution of motor variation patterns.

Rajšp and Fister [36] identified 109 articles related to intelligent data analysis for sports training. They focused on the competitive activity (e.g., no leisure training) in four training stages: planning, realization, control, and evaluation. They discussed the challenges of gathering data sets, working with coaches and players, and applying knowledge to practical situations. On the other hand, Gámez Díaz et al. [37] focused on digital twin coaching, which collects user data and provides personalized feedback. They categorized the works into sports, well-being, and rehabilitation domains. The discussed algorithms, devices, performance, and usability feedback of users. They discussed the challenges in evaluating the user's feedback and user interface.

As these reviews did not particularly focus on pose estimation, they mostly analyzed articles that used other techniques, such as sensors, and only some or a few articles that used pose estimation. Tsiouris et al. [38], for instance, reviewed 41 articles on virtual coaching for users with morbidity, but only one article monitored 3D pose. In this paper, we focus on works that use pose estimation only, as we deem it a promising technology for its capability to work on a normal laptop or mobile phone without extra equipment needed.

2.3 Augmented feedback and their effectiveness

Augmented feedback refers to feedback from an external source to a performer. While works on smart technology for movement assessment discuss feedback, a large body of knowledge about augmented feedback comes from various fields within sport science, in which the feedback is offered by human experts, such as instructors or trainers (simply referred to as instructor in this paper). Several ways to provide feedback, such as knowledge of results and knowledge of performance, have been discussed and studied their effectiveness in various settings. Lauber and Keller [39], for example, reviewed studies that have applied augmented feedback in exercise and prevention settings, focusing on the positive influence of augmented feedback on motor performance. They discussed the limitations of studies, which caused difficulties for practitioners in determining the best way to provide augmented feedback.

Recent reviews of works suggested more pieces of evidence for the effectiveness of different augmented feedback methods. Zhou et al. [40], for example, investigated the data supporting the value of feedback in physical education. Based on 23 studies, the effectiveness of feedback over no-feedback had strong evidence, the effectiveness of visual feedback over verbal feedback had limited evidence, and the effectiveness of information feedback compared with praise or corrective feedback was inconsistent. Meanwhile, Mödinger et al. [41] systematically reviewed the effectiveness of video-based visual feedback in physical education in schools. They found 11 articles in total and suggested that visual feedback was more effective than solely verbal feedback. Still, they found its practical usage required considerations of specific conditions.

The reviewed articles range from the usage of video recording with human instructors to the usage of smart technology to provide augmented feedback. In this paper, we investigate and discuss how pose estimation applications implement or extend existing knowledge of augmented feedback.

3 Methodology

We adopt PRISMA guidelines [42], an evidence-based minimum set of items for reporting in systematic reviews and meta-analyses. We searched for relevant articles twice. The original search was in 2022 and an additional search in 2024 was conducted to include additional articles from 2022 and 2024.

First, we searched Scopus and Web of Science databases on 27 July 2022, using the following query: (“human” OR “user” OR “teacher” OR “student” OR “children” OR “adult” OR “elder” OR “patient” OR “athlete”) AND (“pose estimation” OR “pose tracking”) AND (“exercise” OR “sport” OR “rehabilitation” OR “physical education” OR “motor learning” OR “movement” OR “fitness”) AND (“assistance” OR “correction” OR “guidance” OR “feedback” OR “coach”). The search was in title, abstract, and keywords (i.e., “topic” in Web of Science). We limited the period to between 2017 and 2022, and restricted the language to English only. For Scopus, we also exclude articles with the document type “Review”.

We used CADIMA [43] as a tool for data collection and selection. The inclusion criteria were: (1) human ‘body’ pose estimation is used; (2) user movement is assessed and communicated. Articles that simply classify movement or count correct repetitions without giving feedback to users were excluded. The feedback must be shown, explained, and/or discussed (e.g., articles that simply mentioned giving feedback without screenshots or other details are excluded). We also excluded review articles from the analysis.

The first author performed article selection and initial analysis. We first identified methods for pose estimation, methods for assessing movement, and methods for presenting augmented feedback. Once the methods were noted, we categorized the pose estimation and movement assessment using a bottom-up approach while we largely adopt classifications of augmented feedback from Lauber and Keller [39] and Magill and Anderson [44]. As we focus on the methods, not the result, we did not assess the risk of bias or the certainty (or confidence) of experiments that were conducted to evaluate proposed methods. Instead, we reported how experiments were conducted and what data was used.

We performed the search again using the same databases and query on 5 March 2024, but limited the period to between 2022 and 2024. Then, the second author applied the similar selection process and applied the categorizations derived from the first batch of papers to the second batch, allowing changes to be made if necessary. The second selection process allowed some works that partially fit the criteria, e.g., did not explain the feedback, if they are useful for our discussion. All authors discussed the results and wrote this paper.

4 Result

Fig. 1 presents the article selection process and the results. We found 104 articles from the first search (81 from Scopus and 23 from Web of Science) and 108 articles from the second search (106 from Scopus and 2 from Web of Science). Human pose estimation, movement assessment, and augmented feedback presentation of each article were analyzed and are presented in the next subsections. In the end, we found 45 articles in four main contexts: exercise (Table 2), sports (Table 3), performing arts (Table 4), and healthcare (Table 5). Note that the table is not comprehensive. For instance, the feedback of each system should be either concurrent or terminal or both, but it may be unclear to us so we did not note them in the table.Figure 1 Paper selection flow diagram.

Figure 1

Table 2 List of 26 articles with exercise context. The articles are listed with the publication year, context of use, as well as categories of pose estimation, movement assessment, and augmented feedback presentation. The movement assessment consists of pre-processing (yes), measurement (spatial and temporal), and assessment methods (mathematical formula, rule-based, and machine learning). The augmented feedback presentation consists of type of information (knowledge of result and knowledge of performance), format (visual verbal and visual non-verbal), and timing (concurrent and terminal). Note that the table is not comprehensive.

Table 2			Estimation	Movement assessment	Feedback Presentation	
Paper	Year	Context	Pre.	Measure	Method	Type	Format	Timing	
[16]	'19	General	OpenPose	Y		Math	KP	Nonverbal	Terminal	
[22]	'21	General	OpenPose		Spatial	Math	KR, KP	Verbal		
[25]	'21	General	OpenPose	Y		ML	KR, KP	Verbal, Nonverbal		
[45]	'22	General	MoveNet	Y	Spatial	Math	KR, KP	Verbal, Nonverbal	Concurrent, Terminal	
[46]	'22	General	Mediapipe	Y	Spatial	Math	KR	Verbal	Concurrent	
[47]	'22	General	Mediapipe	Y	Temporal	ML	KR	Verbal, Nonverbal	Terminal	
[48]	'22	General	Mediapipe	Y	Spatial	ML	KR			
[49]	'22	General	Mediapipe	Y	Spatial	ML	KR, KP	Verbal, Nonverbal	Concurrent	
[50]	'23	General	Mediapipe	Y	Spatial	ML				
[51]	'23	General	Mediapipe	Y	Spatial		KP	Nonverbal	Concurrent	
[52]	'23	General	OpenPose	Y	Spatial	ML	KR	Verbal, Nonverbal	Terminal	
[53]	'23	General	Mediapipe	Y	Spatial	Math	KR, KP	Nonverbal	Concurrent	
[54]	'23	General	Other CNN	Y	Spatial	Math	KR, KP	Verbal, Nonverbal	Terminal	
[23]	'21	Arm curls	PoseNet / MoveNet		Spatial, Temporal	Rule	KP	Nonverbal	Concurrent	
[26]	'22	General	mediapipe	Y	Spatial	ML	KP	Nonverbal	Terminal	
[17]	'19	Tai Chi	Other CNN	Y	Spatial, Temporal		KR, KP	Verbal, Nonverbal	Terminal	
[18]	'21	Tai Chi	OpenPose	Y	Spatial	Math	KR, KP	Verbal, Nonverbal	Terminal	
[55]	'22	Tai Chi	Others	Y	Spatial	ML, Math	KR		Terminal	
[24]	'21	Various	Other CNN	Y		Math	KR, KP	Verbal, Nonverbal	Concurrent	
[19]	'21	Yoga	PoseNet / MoveNet			ML	KR, KP	Verbal, Nonverbal	Concurrent	
[20]	'21	Yoga	Other CNN	Y		Math				
[21]	'21	Yoga	OpenPose		Spatial	Math	KP	Verbal, Nonverbal	Concurrent	
[56]	'23	Yoga	Mediapipe	Y	Spatial	ML	KR, KP	Verbal	Concurrent	
[57]	'23	Yoga	OpenPose, Mediapipe	Y	Spatial	ML	KR	Nonverbal, Verbal	Concurrent	
[58]	'23	Yoga	Mediapipe	Y	Spatial	ML	KR, KP	Verbal	Concurrent	
[59]	'23	Yoga	Mediapipe	Y	Spatial	Math	KR, KP	Nonverbal	Concurrent	

Table 3 List of 8 articles with sport context. The articles are listed with the publication year, context of use, as well as categories of pose estimation, movement assessment, and augmented feedback presentation. Note that the table is not comprehensive.

Table 3			Estimation	Movement assessment	Feedback Presentation	
Paper	Year	Context	Pre.	Measure	Method	Type	Format	Timing	
[60]	'22	Badminton	Mediapipe	Y	Spatial	ML	KR, KP	Nonverbal, Verbal	Terminal	
[13]	'21	Baseball	OpenPose		Spatial		KR, KP	Verbal, Nonverbal	Terminal	
[15]	'22	Baseball	OpenPose	Y	Temporal	Math	KR, KP	Verbal, Nonverbal	Terminal	
[61]	'23	Climbing	Others	Y	Temporal	Other	KR, KP	Nonverbal, Verbal	Terminal	
[14]	'21	Kickboxing	PoseNet / MoveNet			Rule, ML	KP	Verbal, Nonverbal		
[62]	'24	Martial arts	Other CNN	Y	Spatial	ML, Math	KR, KP	Nonverbal	Concurrent	
[12]	'19	Skiing	Other CNN		Spatial	ML	KR, KP	Nonverbal		
[11]	'18	Tennis	OpenPose			ML	KR	Nonverbal	Terminal	

Table 4 List of 4 articles with performing arts context. The articles are listed with the publication year, context of use, as well as categories of pose estimation, movement assessment, and augmented feedback presentation. Note that the table is not comprehensive.

Table 4			Estimation	Movement assessment	Feedback Presentation	
Paper	Year	Context	Pre.	Measure	Method	Type	Format	Timing	
[10]	'21	Ballet	OpenPose	Y	Spatial		KP	Terminal		
[9]	'19	Dance	OpenPose		Temporal	Math, Rule	KP	Nonverbal		
[63]	'23	Dance	Mediapipe	Y	Spatial	ML	KR, KP	Verbal, Nonverbal	Concurrent, Terminal	
[64]	'23	Dance	OpenPose	Y	Spatial	ML				

Table 5 List of 7 articles with healthcare context. The articles are listed with the publication year, context of use, as well as categories of pose estimation, movement assessment, and augmented feedback presentation. Note that the table is not comprehensive.

Table 5			Estimation	Movement assessment	Feedback Presentation	
Paper	Year	Context	Pre.	Measure	Method	Type	Format	Timing	
[65]	'22	Knee disorders	Mediapipe	Y	Spatial	ML	KR, KP	Nonverbal	Concurrent	
[27]	'19	Rehabilitation	Others	Y	Spatial		KR, KP	Verbal, Nonverbal	Concurrent	
[3]	'21	Rehabilitation					KP	Verbal, Nonverbal	Terminal, Concurrent	
[66]	'23	Rehabilitation		Y	Spatio-Temporal	ML	KP	Nonverbal, verbal	Terminal	
[67]	'23	Rehabilitation	Other CNN		Spatial	ML	KR, KP	Nonverbal	Terminal	
[68]	'24	Rehabilitation	Mediapipe	Y	Spatial	Math, ML	KR, KP	Nonverbal	Concurrent	
[2]	'20	Various	OpenPose			ML	KR, KP	Verbal, Nonverbal		

4.1 Human pose estimation

Human pose estimation module infers landmarks or keypoints of human figures, such as elbow locations, from an input. Human pose estimation could be broadly classified into 2D (X and Y coordinates) and 3D (X, Y, and Z coordinates). We discuss the input devices, techniques, and studies conducted to evaluate those techniques. It should be noted that our primary focus is on pose estimation for the end users in each article. Certain studies have utilized various methods for tasks such as data preparation. For example, Wu et al. [51] employed MVN Link [69], a professional-grade motion capture device, to generate 3D workout animation examples.

4.1.1 Input devices

We found the usage of a single general camera (including, but not limited to, a built-in laptop camera and an external webcam), mobile camera, binocular or multiple cameras, and depth-sensing camera (Kinect and RealSense). We summarize the input device with pose estimation techniques in Table 6. Note that some articles did not specify an input device but worked on typical videos or clips, listed as general camera/video.Table 6 Inputs and methods for pose estimation. Note that we only note the input(s) that are related to pose estimation. The asterisk (*) indicates that the method is further modified or extended.

Table 6Input	Technique	2D Examples	3D Examples	
General camera/video	OpenPose	[26], [25], [21], [16], [11], [22]*, [52], [64]*	[10], [18]*, [57]	
MediaPipe	[60], [46], [63], [68]	[49], [50], [51], [57]	
PoseNet / MoveNet	[14], [23], [19], [45]	NA	
Other CNN	[12], [20], [54]	NA	
Others	[27]	[27], [55]	
Mobile camera	OpenPose	[13]	NA	
Other CNN	[24]	NA	
Binocular/multiple cameras	OpenPose	NA	[9]*, [67]	
Depth-sensing camera	OpenPose	[15]	[2]	
Other CNN	NA	[17]	
Others	[27], [61]	[27], [61]	

It is clear that a single general camera/video was used the most, as it provided an advantage for deep learning-based pose estimation techniques - users can use the system without the need to purchase additional equipment. Additionally, alternative input devices span from commonplace tools like mobile cameras (e.g., [13], [24]) to specialized setups such as binocular or multiple cameras (e.g., [9]) and depth-sensing cameras (e.g., [15], [2]). Beltrán Beltrán et al. [61] recorded RGB-D video using an iPad Pro 4th Generation, which has a LiDAR sensor, which uses a laser to measure distances and movement in an environment.

Multiple inputs could be employed. Yamei et al. [62], for example, used inputs from cameras or sensor devices to track the positions and angles of users. Some systems used multiple inputs to offer functionalities beyond pose estimation. For example, Jan et al. [18] used shoes equipped with a digital compass to determine the orientation of users while practicing Tai chi.

4.1.2 Libraries or techniques used for pose estimation

We found the usage of OpenPose, MediaPipe, PoseNet/MoveNet, Convolutional Neural Network (CNN), and other techniques for human pose estimation. Table 6 summarizes the findings and links to the relevant articles.

OpenPose is a real-time multi-person pose estimation library based on Cao et al. [70]. The library2 used multi-stage CNN to detect 135 keypoints and reconstruct either 2D pose (e.g., [15], [25], [21], [13], [64], [52]) and 3D pose (e.g., [10], [2]). Though Nagarkoti et al. [16] and Kurose et al. [11] did not state using OpenPose directly, they cited Cao et al. [71], which OpenPose is based on.

Most articles used the library as it is while some extended the library for their purposes. Jan et al. [18] employed Lifting from the Deep [72] to translate a 2D pose into a 3D pose. Similarly, Zhang et al. [9] obtained 2D poses from multiple cameras, then implemented the binocular stereo matching principle to obtain a 3D pose. Yang et al. [22] used data augmentation, randomly cropped and rotated images, to address “abnormal human” situations and considered context information to address the uncertainty in the case of occlusion. Negi et al. [57] combined the MediaPipe and OpenPose through MediaPipe's mp.solutions.

MediaPipe is an open-source framework developed by Google [73] that provides tools, including pose estimation, for building real-time machine learning applications in Android, Python, and Web platforms. The library provides three models (i.e., lite, full, heavy) to detect 33 body landmark locations. It is able to reconstruct 2D pose (e.g., [60], [46], [63], [68]) and 3D pose (e.g., [49], [50], [51], [57]). Most articles used the library as it is.

PoseNet/MoveNet is an open-sourced library [74], [75] that uses a deep learning TensorFlow model to detect 17 2D-keypoints of human. The library supports multiple models. The lightest model could be run in real time on modern smartphones but with lower accuracy. MoveNet is regarded as the next-generation iteration of PoseNet. Articles that used PoseNet include [14], [23], [19], while [45] used MoveNet.

Other Convolutional Neural Networks are used in several works. Jeon et al. [24] used Mobilenetv2 [76] to optimize an HPE model [77], which added a few de-convolutional layers over the last convolution stage in the ResNet [78]. They implemented Online Pose Distillation to minimize performance drop. Wang et al. [12] proposed a structural-aware convolution module, which concatenated spatial and temporal relation modules to go through a convolution layer to reduce dimension. Shi and Jiang [20] proposed a framework containing two branches: one used a confidential map to estimate the positions of bone joint points; another used affinity domain to predict the positions and directions of the limbs. They iterated the above two branches to construct a human skeleton based on the confidence set. Kamel et al. [17] implemented CNN with four convolutional layers, which received input from an RGB-D camera and generated a 3D skeleton model. Mandic et al. [54] used PoseCamera [79], a real-time human pose estimation SDK that is based on work by Osokin [80] and based on MediaPipe for hand tracking. Zheng et al. [67] utilized a pre-trained human pose estimation model [81], which employed a deep convolutional neural network backbone and triangulation approaches to infer the 3D coordinates of joints from multiple views3.

Other techniques, such as Kinect and other deep learning models, were experimented with by Gu et al. [27]. Wei et al. [55] employed the You Only Look Once (YOLO) v4 network [82] to identify bounding boxes. Subsequently, they utilized the Time Series Deep Neural Network (TSDNN) to create heatmaps of key points from human body images. Finally, the Pose Regression Neural Network (PRNN) was employed to convert these heatmaps into keypoints. Beltrán Beltrán et al. [61] employed Apple's Vision framework [83] for extracting the skeleton.

4.1.3 Experiments

Experiments performed on the pose estimation module include accuracy and/or performance such as speed or frame rate. The accuracy is generally measured by comparing the positions of landmarks from the pose estimation module to some ground truth. Kamel et al. [17], for example, compared their pose estimation results with the results they gathered from Kinect. Similarly, Zhang et al. [9] evaluated dance movement reconstruction against previously measured actual 3D coordinates. On the other hand, some articles used publicly available datasets, including COCO [24], [22], MPII [22], Penn Action, and JHMDB [12].

A few works experimented with different models. Gu et al. [27] compared 2D pose and 3D pose generated from four deep learning models [84], [76], [72], [77] and Kinect. They selected Kinect for their system as it provided the most accurate result. Huang et al. [21] evaluated the accuracy and frame rate of different models, but there is no detail about the data used for the evaluation.

4.2 Movement assessment

This module replaces the need for human experts, such as instructors or trainers, to oversee user movement. In order to analyze the quality of movement and give users feedback, the system may process keypoints or the motion of keypoints and/or measure the amount or degree of user attributes, then assess how well a user performs. We discuss the techniques for each part and studies conducted to evaluate those techniques (see Table 7).Table 7 Movement assessment.

Table 7Technique	Examples	
Pre-processing		
Temporal alignment	Dynamic time warping [10], [18], [25], [27], [16], [61]	
Spatial alignment and normalization	Align using selected origin point and normalize [27], [17], [63], [64], [52], [49], [50], [55], [57], [51], [58], [56], [45], [47], [46], [60], [65], [54], [48][53]; Normalization only [20]	
Noise handling and smoothing	Filter [26], [10], [24]; Fill missing values [20]; Quarter-shift [24]	
Selection	Spatial [18], [26]; Temporal [15]	
Temporal segmentation	Peak detection [26]	


	
Measurement		
Spatial measurement	Angle of body parts [23], [10], [18], [21], [13], [12], [17], [22], [27], [68], [63], [64], [52], [49], [50], [56], [55], [57], [51], [58], [62], [45], [47], [46], [59], [60], [65], [54], [48], [53], [67]	
Temporal measurement	Direction/displacement [17], [15], [9], [61]; Range of motion/rotated angle [23], [9]; Number of repetition [26]	
Spatio-Temporal measurement	Angle of body parts mapping with time [66]	


	
Assessment methods		
Mathematical formula or model	Difference of a measurement [16], [9]; Angular similarity/distance [15], [24], [53], [21], [68]; Euclidean distance [45], [21], [20], [59], [54]; Others [18], [24], [22], [46], [47];	
Rule-based method	Grading [9]; Correct posture checking [23], [14], [23]	
Machine learning	Classification using Support Vector Machine(SVM) [11], [12], [2]; Classification using Random Forest Classifier (RFC) [49]; using k-nearest neighbor (KNN) [48]; using Neural Network based on LSTM [2]; using Artificial Neural Network (ANN) [14], [19]; using Deep Neural Network (DNN) [64], [57], [58]; using transformer-based model [62]; using multilayer perceptron(MLP) [52]; using time series classification [26]; using XGBoost [55]; Similarity score of embedding pairs from multi-stage CNN [85], [25], [66]; Deep learning-based [56], [60]; using multiple machine learning [50]; using Spatial-Temporal Graph Convolutional Network (ST-GCN) [67]; Others [63], [51], [65], [61];	

4.2.1 Pre-processing

Pre-processing involves translating keypoints or the motion of keypoints into a more desirable one, such as finding correct keypoints by removing wrong keypoints. Pre-processing techniques include temporal alignment, spatial alignment and normalization, noise handling and smoothing, selection, and temporal segmentation.

Temporal alignment aligns temporal sequences to cope with temporal issues. For instance, a delay typically occurs when a user is trying to imitate the instructor's movement. The speed of movement between a user and an instructor could differ. Temporal alignment attempts to minimize such temporal differences so the movement assessment can focus on comparing, for example, the posture. All papers with temporal alignment [10], [18], [25], [27], [16], [61] employ dynamic time warping (DTW). The technique has been used to find patterns in time series [86] in various domains. The technique optimizes a distance metric and nonlinearly maps a frame of the instructor to a frame of the user, as illustrated in Fig. 2. The distance metric can be customized. For instance, Nagarkoti et al. [16] used angles between the pair of limbs as the distance metric.Figure 2 Dynamic time warping by Programminglinguist [CC BY-SA 4.0], via Wikimedia Commons. Two sequences (the solid lines) are matched (the dash lines) with certain rules.

Figure 2

Spatial alignment and normalization processes and geometrically matches points within a frame to address spatial difficulties, such as differences in physique or camera distance. The alignment and/or normalization are particularly essential when user performance is inferred from the position of human keypoints. Gu et al. [27] and Shi and Jiang [20] used a pelvis joint or a middle point to be the origin and normalized coordinates before calculating the Euclidean distance between an instructor and a user. On the other hand, Kamel et al. [17] only normalized coordinates to neutralize the difference in body size. They used rotations and direction of motions to evaluate the user performance, thus the alignment was not necessary.

Noise handling and smoothing processes data to minimize the effect of imperfections in pose estimation. Singh et al. [26] used SavGol filter [87] to minimize fluctuations of coordinates before detecting repetitions. Li and Pulivarthy [10] applied a median filter to angle sequences to prevent poor results due to noisy data. Jeon et al. [24] applied heatmap-smoothing, quarter-shift, and the one-euro filter to minimize fluctuation. Shi and Jiang [20] predicted undetected joint coordinates based on the standard movement of the limbs.

Selection chooses a subset of relevant features before processing, which could be either temporal or spatial features. Jan et al. [18] used only elbow, knee, and foot information for evaluating Tai-Chi Chuan practice. In contrast, Singh et al. [26] removed ankle, knee, and other points that have low variability when evaluating CrossFit workouts. Akiyama and Umezu [15] excluded redundant frames before and after a baseball pitching motion by selecting only 30 frames after the system detected a specific body part at an angle.

Temporal segmentation splits a sequence into sub-sequences. Singh et al. [26], for example, used peak detection methods to segment workout exercises into repetitions.

4.2.2 Measurement

Measurement refers to the quantification of attributes of one person. For instance, a system could identify the value of the angle between body parts or quantify the speed of a movement. Measurement could be briefly categorized into spatial, temporal and spatio-temporal.

Spatial measurement. The measurement involves the attributes of one person within one frame. Yang et al. [22], for example, calculated the horizontal distance between the hip and heel and the score of thigh length of a user to infer the proper form of a squat. The angle of body parts seems to be the most common spatial measurement [23], [10], [18], [21], [13], [12], [17], [22], [27], [68], [63], [64], [52], [49], [50], [56], [55], [57], [51], [58], [62], [45], [47], [46], [59], [60], [65], [54], [48], [53], [67]. Ranasinghe et al. [23], for example, calculated hip-shoulder-elbow angle to check for proper elbow lock during an arm curl exercise.

Temporal measurement. The measurement involves the attributes of one person across frames. A straightforward measurement is a difference in keypoint position over time. The direction of motion was used by Kamel et al. [17] in practicing Tai Chi. Similarly, Akiyama and Umezu [15] calculated displacement direction and distance of the body joint to provide suggestions for improving a baseball pitching form. Ranasinghe et al. [23] used the range of motion, i.e. the angular distance of the movement around a joint, to evaluate an arm curl exercise. Zhang et al. [9] incorporated the distance of the joint point between frames and the corresponding rotated angle as a curvature of the joint point combination movement for analyzing dance actions. Akiyama and Umezu [15] visualized the amount of joint movement between two successive frames as the indication of speed. They also mentioned the acceleration of joints and timing of the motion, but the detail of how those attributes were measured is unclear. A number of repetitions counted from a number of segments such as [26] can also be seen as a temporal measurement. However, a number of correct and/or incorrect repetitions would typically involve some assessment methods.

Spatio-Temporal measurement. The spatio-temporal measurement captures and analyzes the spatial coordinates of body joints in 2D or 3D space with associated time changes across the sequence of frames. This ensures a thorough and correct understanding of the body postures and movements taking place over time. Garg et al. [66] have developed a new convolutional neural network-based model specifically for the task of patient performance evaluation. In other words, it has combined both the spatial and temporal dimensions, already in the first layers of the design, for the perfect improvement of the faculty of a set assessment, which is accurate and deep. Its first two layers take input with the angular (spatial) and temporal data and, thereby, capture high-order relationships representing complex patterns and changes in variations of patient performance through time.

4.2.3 Assessment methods

Assessment methods assess how well users perform. The system could implement multiple methods or conditionally select methods, for example, based on the type of movement as seen in [21]. Note that some methods could be explained in multiple ways. We categorized the methods as the authors expressed them.

Mathematical formula or model evaluates user performance mathematically. It could be as simple as the summation or average of the absolute difference of measurement (e.g., angle [16], [9] or displacement [9]). Variations of formulas have been adopted to find angular similarity/distance [15], [24], [21], [53], [68] as well as Euclidean distance [21], [20], [45], [59], [54]. Other mathematical formulas or models found include Gaussian function-based similarity metric [18], geometric analysis and comparison [46], dynamic time warping (DTW) [47] and the dot product between the joint dynamic vectors [24]. The mathematical formula or model typically results in a numerical value, which could be further classified into classes (e.g., good and bad) using a formula or other assessment methods. Yang et al. [22], for example, expressed the conditions of good angle and hip-heel distance in mathematical form.

Rule-based method evaluates or classifies user performance based on conditions. Zhang et al. [9], for example, graded user performance (excellent, good, pass, fail) based on the similarity score. The rule-based method could also lead to a numerical result. Ranasinghe et al. [23] detected a number of correct arm curl repetitions when shoulder-elbow-wrist angles are less than 90 degrees, then larger than 170 degrees repeatedly. Wessa et al. [14] and Ranasinghe et al. [23] tracked the time taken to perform a specific action by identifying the starting and ending point according to conditions of user posture.

Machine learning leverages data to evaluate user performance. It is typically used for classifying user performance into a predefined class. The technique used ranges from well-known supervised learning, such as Support Vector Machines (SVM) or K-nearest neighbors, to deep learning. Kurose et al. [11] first created feature vectors using joint position coordinates and classified the vectors with Gaussian Mixture Model (GMM). They then used SVM to predict a tennis shot result based on the posture class, movement amount, and play area. Wang et al. [12] used SVM with radial basis kernel function (RBF) to classify good and bad poses of skiing. Chalvatzaki et al. [2] used gait parameters, such as stride length or gait speed, as a feature vector for an SVM classifier with classes from Performance Oriented Mobility Assessment [88]. They also used Neural Network based on LSTM units for human activity recognition and gait stability assessment. Wessa et al. [14] and Tarek et al. [19] used the Artificial Neural Network (ANN) of 3 layers to classify keypoints into correct and incorrect poses of kickboxing and yoga. Singh et al. [26] implemented and compared four time series classification methods, including 1-nearest neighbors dynamic time warping (1NN-DTW), ROCKET [89], Fully Convolutional Network (FCN), and Residual Network (ResNet).

In addition to classifying a pose into a class, machine learning could be used for calculating a numerical value. Park et al. [85], Zhou et al. [25] proposed a body part embedding model for motion similarity. They decomposed joint points into 5 body parts, then encoded motion classes, skeletons, and camera views of each body part as an embedding using multi-stage CNN. The similarity score was then calculated from the average cosine similarity between the embedding pairs.

4.2.4 Experiments

Studies related to movement assessment include performance and/or its accuracy. For the performance, Singh et al. [26] reported training time. Shi and Jiang [20] mentioned that their system satisfied realtime requirements, but no actual time was given. On the other hand, we found various ways to test how the proposed automated assessment conforms to the correct value or standard.

One common way was to ask humans to put a label or value, such as a score or number of repetitions, on movement gathered from themselves or other participants, then compare human annotation with the system, as seen in [21], [20], [19], [24]. For user movement assessment that employed machine learning, labeled data could be also used for training. Zhou et al. [25] and Wang et al. [12] asked participants to perform the correct and incorrect movements, then split the data to train and test their classification. The ground truth could come from seen results, such as the hit distance of a baseball swing [13] or a tennis shot result [11], and the authors discussed the correlation between their result and the actual one [13], [11].

Some articles used less formal methods for the study. Li and Pulivarthy [10] and Kamel et al. [17] compared the scores of experienced and novice users, then inferred the effectiveness of their assessment module as the experienced users' scores were higher than the novice scores. Yang et al. [22] simply intentionally performed incorrect movement and presented the output. Lastly, we found only one article that used a publicly available dataset. Zhou et al. [25] used NTU RGB+D similarity annotations dataset to validate their results.

4.3 Augmented feedback presentation

Augmented feedback (or extrinsic feedback or external feedback) refers to information about performance from others. In our scope, it is the feedback provided by the system to a user. We discuss the type of information, format, and timing and frequency of the augmented feedback. Table 8 provides examples of works that implemented each category of feedback. Note that some categories are excluded from the table due to the lack of examples, and one work could provide multiple pieces of feedback.Table 8 Examples of augmented feedback, classified by type of information and format. We note the timing Concurrent and Terminal as superscripts. We also note how the instructor and/or user are visualized: Instructor only, User only, Juxtaposition, Superposition, and Relationship encoding. The asterisk (*) indicates different media types are put, for example, in juxtaposition. NA means we could not find any example from the reviewed papers.

Table 8Category	Knowledge of Results	Knowledge of Performance	
General	Error	
Audio-verbal	[56]C	[45]C	[58]C	


	
Visual-verbal	
- Number	[24]C, [27]C, [18]T, [13]T, [17]T, [25], [2]	NA	NA	
- Word(s)	[15]T, [13]T, [56]C	NA	NA	
- Phase	[60]T, [61]T	[19]C, [60]T, [61]T	[21]C, [10]T, [14], [22], [66]	


	
Visual-nonverbal	
- Video	NA	U: [24]C, [13]T, [9], [2]; U&R: [19]C, [17]T; J: [23]C*, [14]*, [45]C; J&R: [18]T*; S: [16]T*	I: [12]	
- Image	NA	S&R: [15]T	U: [26]T; J: [23]C*, [18]T*, [14]*	
- Animation	NA	U: [2], [9]; J: [27]C; S: [16]T*	J&R: [25], [51]C	
- Diagram	[45]T, [54]T, [61]T, [53]T	[47]T	[61]	
- Other	[19]C, [11]T	NA	NA	

4.3.1 Type of information

Type of information involves what content of the feedback is communicated to a user. Literature classifies augmented feedback into two main types: knowledge of results and knowledge of performance.

Knowledge of results (KR) gives information about the outcome of the user's movement. Common feedback of this type includes an indication whether a user does a movement correctly [13] or a score relating to movement quality [18], [24], [13], [2], [17]. The feedback could be a result of assessing overall performance or individual aspects. Li et al. [13], for example, color-coded a list of body parts according to the goodness of the posture of each part. A number of repetitions, as seen in [24], [27], or recognized pose name, as seen in [56], could be another feedback that informs the successful movement. One knowledge of results could lead to another knowledge of results. For instance, Cai et al. [45] associated the number of repetitions per minute with physical fitness, such as body strength and lower body strength.

Some knowledge of results is tied to the context of movement. Akiyama and Umezu [15], for example, compared the user's baseball swing timing with the instructor and displayed whether the user was faster, slower, or had similar timing. Kurose et al. [11] listed the characteristics of tennis shots for each prediction result (a score, a losing point, or a rally continuation).

Knowledge of performance (KP) gives information about the characteristics of the user's movement that lead to the result. Tarek et al. [19] informed whether and why the pose is correct or incorrect (e.g., a foot above kneecap for Yoga Tree Pose) by highlighting the related body parts over user video recordings. Video replay or live video of a user is a common approach to demonstrate user performance, and systems usually annotate them with information from pose estimation or movement assessment modules. For example, a system could show a user video, overlaid with its detected skeleton [24], [13], [2], [45]. A system could highlight parts with low scores [25] or color-coded body parts according to similarity score [18], then overlaid on a user video. Other types of knowledge of performance include joint angles over time, as seen in [47], and movement kinetics and kinematics, such as the amount of movement over time [15].

The information from an assessment module, as the annotation or other forms, could highlight correct aspects or error aspects of the performance. Correct aspects inform that a user is on track and encourages them to continue. Error aspects inform a user about their mistakes (descriptive) and/or directs how to correct the mistakes (prescriptive). Akiyama and Umezu [15] highlighted body parts that received minimum similarity score. They provided advice on form improvements by drawing arrows indicating the motion direction of the instructor on the user's image. Some systems, such as [14], [23], [10], [22], [12], [20], [66], suggested the correction or the instructor movement when the system detected the wrong movement.

Many systems provide both knowledge of performance and knowledge of results. For instance, [61], [60] explained characteristics of the user's movement, e.g., “your hips are very far away from the wall”, and the result of the movement, e.g., “so it pulls you out from the wall”. In addition to displaying information from an assessment module, the system could have a dedicated module for feedback. Ranasinghe et al. [23] used a reinforcement learning model to provide correct and incorrect instruction images. Li and Pulivarthy [10] surveyed correction feedback from experts, then used a decision tree classifier with Gradient Boosting to classify the feedback and present it to a user.

4.3.2 Format

The format involves the ways the feedback is communicated to a user. We broadly classify format according to signs and channels [90] into audio-verbal (word heard), visual-verbal (word read), and visual-nonverbal. There could be a system that provides audio-nonverbal, such as music or sound effects, but we have not found such a system in this review.

Audio-verbal is the use of words to communicate feedback through speech. This approach is the closest to traditional motor learning, where a coach or trainer verbally gives a learner feedback. To translate an output to speech, Chalvatzaki et al. [2] used a TTS (text-to-speech) system to give verbal feedback to the user. The systems usually provide audio feedback along with visual feedback. Lavanya et al. [56], for instance, displayed and announced a recognized yoga pose name when users performed the pose correctly. Similarly, Singh et al. [49], Anuradha et al. [58] provided audio feedback if improper forms were detected. Cai et al. [45] provided more general audio feedback, such as “You are doing well” or “Please follow the instruction and adjust your pose.”

Visual-verbal is the use of words to communicate the feedback visually (e.g., on-screen). Quantitative information is generally presented as Number, as found in [18], [25], [24], [13], [17]. For qualitative information, a system could display a category as simple words, e.g., “faster” or “slower” [15], or “front arm angle” or “back arm angle” [13]). Lastly, a system could output phases or sentences, as seen in [14], [10], [21], [22], [66], [60].

Visual-nonverbal does not use words to communicate the feedback. We found the use of videos (e.g., [21], [24], [13], [2]), images (e.g., [14], [23], [18], [26], [9]), 2D animations of skeleton (e.g., [27], [9]), diagram (e.g., ability radar charts [45], line graph [47], or bar chart [54]), and other graphic (e.g., correct/incorrect symbols [19] or movement paths and colormap [11]). Ajay et al. [53] proposed Gradient-Weighted Class Activation Mapping (Grad-CAM) for visualization and a quantitative metric called Overlap Ratio (OvR) to measure the quality of the visualization result. Some systems display multiple types of media. Jan et al. [18], for example, displayed a user video beside a 3D model of an instructor.

The system could display information about a user only, an instructor only, or both. Displaying user only could be used to provide KP feedback in general as well as highlight incorrect aspects of user movement. Systems, such as [24], [13], [2], visualized detected skeleton on a user video. Singh et al. [26] displayed frames from the user camera feed that were identified as incorrect. Displaying instructor only, on the other hand, is typically used to prescribe the correct movement. Wang et al. [12] offered advice to a user by sampling a related video clip with correct poses.

In case both instructor and user information are displayed, we could further consider visual comparison [91], as a user typically needs to identify differences or similarities of their movement compared to the movements of an instructor in order to see errors and/or improve their movement. Juxtaposition places a user and an instructor in separated, nearby spaces. Juxtaposition could be side-by-side, as seen in [21], [23], [18], [27], or picture-in-picture, as seen in [14]. Superposition places a user and an instructor in the same space. Akiyama and Umezu [15], for example, superimposed the user image on the instructor image. Nagarkoti et al. [16] displayed the user's video with their skeleton and overlaid the skeleton of an instructor on them when the error was above a threshold. Lastly, relationship encoding displays relationships between a user and an instructor, such as a score or classification results, directly. The relationship could appear as an annotation on the user's video, as seen in [19], [18], [25].

We also found the usage of the temporal component of a video to communicate the feedback that a user could perceive visually. Kamel et al. [17] provided motion replay, which would be suspended when the similarity score of a frame was lower than 50%, and the joints that have lower-than-average scores were highlighted with yellow. Additionally, a visual comparison could happen with a visual-verbal format. For example, a similarity score could be considered as relationship encoding since it represents a relationship between an instructor and a user.

4.3.3 Timing and frequency

Timing and frequency involve when and how often the feedback is communicated to a user. Users could benefit from appropriate timing and frequency of feedback, for which multiple approaches were used in the literature.

Timing could be generally classified into concurrent and terminal. Concurrent feedback is the feedback given while a person is performing a skill or movement, as seen in [23], [19], [21], [27], [56]. Feedback in audio format in our review was all provided concurrently in Lavanya et al. [56], Singh et al. [49], and Anuradha et al. [58]. Terminal feedback is the feedback given after a person has finished performing a skill or movement, as seen in [15], [26], [10], [18], [13], [16], [17], [11]. Blanchet et al. [63] offered two kinds of feedback to the user: one is concurrent feedback, which is given through a color-coded skeleton overlay, and the other is terminal feedback, which is given as a score and performance summary text after each attempt at the dance.

Frequency of feedback could be on every practice trial or less. In case of reduced frequency, a system could provide feedback only when user performance meets a certain value (performance-based bandwidths). Nagarkoti et al. [16], for example, projected an instructor movement on a user only when their error is above a threshold. Blanchet et al. [63] offered incremental part learning, i.e., learning each sub-task separately before learning the whole task. As the dance is being learned, the system gradually decreases the amount of assistance provided for the task. The frequency could be self-selected, or the feedback could be summarized or averaged after a number of practice trials. However, we have yet to find examples in this review.

4.3.4 Experiments

Most articles evaluated augmented feedback, specifically or as a whole system, through users. Wu et al. [51], for example, conducted user studies with 24 participants in a comprehension phase and 12 participants in a practice phase. They evaluated different types of representations (2D and 3D), visual cues (directional cues, measurement cues, skeleton joints, and visual metaphors), and availability of feedback (with and without feedback) across different types of exercises (isotonic exercises, including bicycle crunch, mountain climber, squat, and jumping jacks, and isometric exercises, including plank and AB hold).

Straightforward ways included asking participants to use a system, then conducting interviews or using questionnaires to investigate system usability, usability issues, user engagement, user preference, and/or user satisfaction [17], [18], [12], [27], [51]. Alternatively, user performance or improvement when using a system could be used to evaluate the overall system. A study could be short-term (within a day, e.g., [15], [14], [23], [17]) or long-term (e.g., [14], [27], [17]).

Few works employ quantitative measurements that do not require user testing. Ranasinghe et al. [23], which provided feedback based on the reinforcement learning model, reported the relevance and value of feedback on errors. Zheng et al. [67] developed a quantitative metric called Overlap Ratio (OvR) to measure how well the visualization method distinguishes between correct and incorrect movements.

5 Discussion

This section summarizes the findings of each module and the overall pipeline that we observed while reviewing the articles. We also suggest some future research directions for pose estimation-based physical movement applications.

5.1 Pose estimation

Open-source libraries that support pose estimation in real-time plays an important role for research on pose estimation-based physical movement applications, as supported by the number of works that used OpenPose and MediaPipe (see Table 6). It is also important to note that some pose estimation techniques may not have emerged yet or may appear to be underused in our review. For instance, we found no usage of transformer-based methods, which is deemed more potent compared to Convolutional Neural Networks (CNNs) due to its capacity of attention modules within transformers to capture long-range dependencies and provide global evidence for predicted keypoints [33]. One possible reason for this absence is the lack of easy-to-use open-source libraries of the methods. The libraries should enable researchers beyond the machine learning domain, such as HCI researchers, to readily adopt pose estimation techniques for evaluating and delivering feedback to users, eliminating the necessity for an in-depth understanding of intricate technical details.

Pose estimation libraries for a specific purpose may be needed, even though general-purpose pose estimation libraries are useful. Wang et al. [12], for example, proposed a pose estimation method that incorporates spatial and temporal relation of human keypoints to deal with the fast movement of ski technique and skiing skis. They trained the model, experimented with a dataset, and reported errors in pose estimation due to crossed snowboards that mixed with the background. As a physical movement application usually focused on one or a few specific types of movement, a pose estimation library that could be trained with a custom dataset could be useful for research in this area and possibly yield more accurate results.

Pose estimation may not have to be perfect to be useful in physical movement applications. We found very limited discussion about how pose estimation issues as well as the limit of sensitivity and accuracy of the pose estimation models could affect the assessment. In our prior work [92], we observed some inaccuracy of pose estimation results during the experiment with Tai Chi movements, but our participants did not notice it because they focused on other parts. Thus, in addition to continue improving pose estimation accuracy, one future research is to study the effect of imperfect pose estimation in physical movement applications on users.

5.2 Movement assessment

There is limited evaluation and discussion to tell which technique is better or which factors should be considered when applying a technique to a new context. For instance, differences in physique, camera angle, or camera distance could result in high Euclidean distance even if a user perfectly imitates a trainer. Though there could be articles outside our review that compare techniques, such as Srikaewsiew et al. [93], none of them were compared with proposed movement assessment methods. We thus encourage more experiments to evaluate and identify the limitations of movement assessment methods.

Sharing data about movement similarity or movement quality could be useful. We believe sharing more data about movement similarity or movement quality could be useful for evaluating and comparing the movement assessment methods and advancing this field. For instance, researchers can prompt users to perform movements, assess them using their system (0-100 score), and compare them with expert evaluations. In addition to the error (difference between the system and expert assessment), the researchers could consider publishing user recordings and/or pose estimation results along with the score given by the experts. As such, other researchers could leverage the pose estimation data and compare their assessment results with the experts without the need to gather new data, similar to how Zhou et al. [25] used the public dataset, NTU RGB+D 120 Similarity Annotations [85], to evaluate action similarity assessment for fitness assistance.

Transformer-based models could become integral to the evaluation capabilities of assessment methodologies. The development of transformer-based models in the paradigms of assessment is taking another massive leap for computational methodologies. The dynamic potentials of such models were sketched by Yamei et al. [62] in movement assessment where the dynamism falls upon the estimation of spatial measurements with special emphasis on the nuanced methods of edge contour and segments of molecular structures in regard to body posture analysis. Further, the addition of classification through artificial neural networks adds a realm of sophistication to this proposed transformer-based model and promises a subtle exploration of the correctness of the postures by the users. In this regard, the above features are reasons for it being revolutionary in design, but they also lay down the path for further elaboration regarding the effectiveness and reliability of contemporary models of assessment, which now promise an improved level of precision and robustness within computational frameworks of assessment.

5.3 Augmented feedback presentation

Relating results of automated feedback to traditional feedback is difficult due to the lack of feedback description in a number of reviewed papers. We mainly adopt feedback classification from literature in motor learning by Magill and Anderson [44]. The classification is well-established and includes research works where a human trainer gives feedback to a learner, which should allow us to compare research within our scope with others. However, we found difficulty in classifying feedback. For instance, Ranasinghe et al. [23] mentioned verbal encouragement phrases, but no detail about the channel was given. Chalvatzaki et al. [2] gave verbal feedback, but it is unclear what feedback information was given to users. We were unable to identify the timing of a number of works ([14], [25], [22], [2], [12], [9], [20]). This lack of description limits the replication, comparison, as well as contribution to general knowledge in motor learning. Thus, we encourage authors to always report content, format, and timing of feedback in their literature.

Researchers could experiment with different ways to provide augmented feedback in the design space. While it is not comprehensive, Table 8 can still give us an idea of trends and gaps that can be experimented with. We could observe more usage of a visual channel in communicating verbal feedback, compared to audio one. Different forms of visual objects, such as a number or a video, seem to be suitable for different types of content. Still, one could try to come up with items in NA cells or items excluded from the table. A system could provide correct aspects of knowledge of performance, kinetics, kinematics, and/or biofeedback.

Study about feedback prioritizing is still underexplored. In motor learning, selecting the skill component for knowledge of performance is recommended [44]. For instance, a trainer could tell a learner to correct arm movement first, and then leg movement later. We still found only a limited number of works (e.g., [63]) that adopt such an approach. For automated feedback, one area that we found interesting to explore is the usage of artificial intelligence to provide feedback according to user needs, similar to [23], [10]. Alternatively, designing a user interface that allows a user to customize provided feedback could take less time to implement and easier to test the effect of priority feedback in an automated system.

Erroneous feedback should be considered and discussed, even though there is no apparent discussion about the challenges of providing augmented feedback from the review. Erroneous feedback is wrong feedback, which could hinder motor learning as a user would use the wrong feedback and ignore their (correct) sensory feedback [94]. Though the accuracy of pose estimation has improved over the years, it is not perfect and could lead to erroneous feedback. A system could, for example, disable augmented feedback when the confidence value of pose estimation is below a threshold. Finding a way to mitigate the erroneous feedback could be one essential future research direction for pose-estimation-based physical movement applications to be widely adopted.

5.4 Overall pipeline

We could observe the dynamic landscape of this field as we performed the search twice. Notably, we found a shift in the popularity of pose estimation libraries. Initially, OpenPose emerged as the most widely used library, but in subsequent searches, MediaPipe took the lead. Furthermore, we noticed an increase in the number and variety of machine learning-based assessment methods. There has been an increase in the usage of audio with visual feedback from the growing interest in multimodal approaches for enhancing user experience. Additionally, we identified more examples to replace previously empty cells (marked as “NA”) in Table 8.

One notable advantage of pose-estimation-based methods lies in their adaptability to novel pose estimation techniques without requiring significant modifications to interconnected modules, as they operate on a set of body keypoints. For instance, assessing angles of body parts can be applied in both OpenPose and newer libraries like MediaPipe. Additionally, feedback visualization could still be a score. However, beyond improving accuracy, advancements in pose estimation methods could lead to a greater variety of movements that can be assessed. Novel assessment methods could then be developed, resulting in more comprehensive feedback presentations for the benefit of end users. While pose estimation, movement assessment, and augmented feedback presentation could each be individual research topics, we believe that having an overview of this pipeline is crucial for practical adoption by end users.

5.5 Limitations

Our work has two main limitations. First, there exist articles that addressed specific module(s), which may provide rich knowledge about the module, but were excluded from this review by our inclusion criteria or other reasons. Our prior works, for instance, either provided feedback with no movement assessment [92] or assessed movement without providing feedback [93]. Still, our review should be useful to see the entire pipeline of pose estimation-based systems in providing feedback for physical movement. Table 1 should be useful for those who are interested in a particular module of the systems.

Second, as we discussed in the previous section, some articles provided very limited descriptions of the user interface and feedback. We thus used the provided images as a source of information. It is possible that we could not study every aspect of the user interface. Also, it was sometimes unclear whether an image depicted concepts or actual implementation. Because the description was quite limited, we decided to consider those images as a way to communicate feedback to either users or readers and included them in our analysis.

6 Conclusion

The purpose of this study was to examine current methods for estimating human pose, assessing movement, and providing feedback using deep learning techniques. The review analyzed 45 articles on these topics. We found that systems used CNN (notably OpenPose and MediaPipe) for pose estimation and used either mathematical formula or model, rule-based method, or machine learning for assessing movement. The feedback, including knowledge of the result and knowledge of performance, was mostly presented visually in verbal forms (i.e., number, word, and phase) and nonverbal form (i.e., video, image, animation, diagram, and other graphics). Public datasets and human subjects were both used in the tests to assess each module.

We discuss the remaining gaps and areas for future research. We suggest a need for general-purpose pose estimation libraries as well as specialized ones. While there are still challenges to be addressed in pose estimation, it may not be necessary for it to be perfect to be useful in certain applications. For movement assessment, there is limited evaluation and discussion on which techniques are best and what factors should be taken into consideration when applying them to new contexts. Ready-to-use datasets should facilitate the evaluation. We suggest further study about feedback prioritization and erroneous feedback, which could possibly facilitate the adoption of pose-estimation-based physical movement applications.

CRediT authorship contribution statement

Atima Tharatipyakul: Writing – review & editing, Writing – original draft, Validation, Methodology, Investigation, Funding acquisition, Formal analysis, Data curation, Conceptualization. Thanawat Srikaewsiew: Writing – review & editing, Formal analysis, Data curation. Suporn Pongnumkul: Writing – review & editing, Resources, Project administration, Investigation, Funding acquisition, Conceptualization.

Declaration of Competing Interest

The authors declare the following financial interests/personal relationships which may be considered as potential competing interests: Suporn Pongnumkul reports financial support was provided by National Electronics and Computer Technology Center. Suporn Pongnumkul reports a relationship with National Electronics and Computer Technology Center that includes: employment.

Data availability statement

Data included in article/supp. material/referenced in article.

Acknowledgement

This research is funded by the 10.13039/501100011058 National Electronics and Computer Technology Center , Thailand.

2 https://github.com/CMU-Perceptual-Computing-Lab/openpose.
==== Refs
References

1 Erol A. Bebis G. Nicolescu M. Boyle R.D. Twombly X. Vision-based hand pose estimation: a review Comput. Vis. Image Underst. 108 1–2 2007 52 73
2 Chalvatzaki G. Koutras P. Tsiami A. Tzafestas C.S. Maragos P. i-Walk Intelligent Assessment System: Activity, Mobility, Intention, Communication 2020 School of E.C.E., National Technical University of Athens Athens, Greece 500 517
3 Blas H.S.S. Mendes A.S. Encinas F.G. Silva L.A. González G.V. A multi-agent system for data fusion techniques applied to the internet of things enabling physical rehabilitation monitoring Appl. Sci. 11 1 2021 1 19
4 Mousavi Hondori H. Khademi M. A review on technical and clinical impact of Microsoft kinect on physical therapy and rehabilitation J. Med. Eng. 2014 2014
5 Difini G.M. Martins M.G. Barbosa J.L.V. Human pose estimation for training assistance: a systematic literature review ACM International Conference Proceeding Series 2021 189 196
6 Stenum J. Cherry-Allen K.M. Pyles C.O. Reetzke R.D. Vignos M.F. Roemmich R.T. Applications of pose estimation in human health and performance across the lifespan Sensors 21 21 2021 7315 34770620
7 Badiola-Bengoa A. Mendez-Zorrilla A. A systematic review of the application of camera-based human pose estimation in the field of sport and physical exercise Sensors 21 18 2021 5996 34577204
8 Da Gama A. Fallavollita P. Teichrieb V. Navab N. Motor rehabilitation using kinect: a systematic review Games Health J. 4 2 2015 123 135 26181806
9 Zhang L. Liang X. Zhang W. Tang R. Fan Y. Nan Y. Song R. Behavior recognition on multiple view dimension International Conference on Wavelet Analysis and Pattern Recognition, vol. 2019-July 2019 State Grid Henan Electric Power Company Kaifeng Power Supply Company China
10 Li J.E. Pulivarthy H. BalletNetTrainer: an automatic correctional feedback instructor for ballet via feature angle extraction and machine learning techniques Proceedings of the International Conference on Industrial Engineering and Operations Management 2021 ARQuest SSERN International Kirkland, WA, United States 603 613
11 Kurose R. Hayashi M. Ishii T. Aoki Y. Player pose analysis in tennis video based on pose estimation 2018 International Workshop on Advanced Image Technology, IWAIT 2018, Matsudo Orthopedics Hospital Chiba, Japan 2018 1 4
12 Wang J. Qiu K. Peng H. Fu J. Zhu J. AI coach: deep human pose estimation and analysis for personalized athletic training assistance MM 2019 - Proceedings of the 27th ACM International Conference on Multimedia 2019 Microsoft Research Asia China 2228 2230
13 Li Y.C. Chang C.T. Cheng C.C. Huang Y.L. Baseball swing pose estimation using OpenPose 2021 IEEE International Conference on Robotics, Automation and Artificial Intelligence, RAAI 2021 2021 Physical Education Tunghai University Taichung, Taiwan 6 9
14 Wessa E. Ashraf A. Atia A. Can pose classification be used to teach Kickboxing? International Conference on Electrical, Computer, and Energy Technologies, ICECET 2021 2021 Helwan University, HCI-LAB, Faculty of Computers and Artificial Intelligence Egypt
15 Akiyama S. Umezu N. Similarity-based form visualization for supporting sports instructions LifeTech 2022 - 2022 IEEE 4th Global Conference on Life Sciences and Technologies 2022 Ibaraki University, Dept. Mechanical System Engineering Hitachi, Japan 480 484
16 Nagarkoti A. Teotia R. Mahale A.K. Das P.K. Realtime indoor workout analysis using machine learning computer vision Proceedings of the Annual International Conference of the IEEE Engineering in Medicine and Biology Society, EMBS 2019 Samsung Research Institute Bangalore, 560037, India
17 Kamel A. Liu B. Li P. Sheng B. An investigation of 3D human pose estimation for learning Tai Chi: a human factor perspective Int. J. Hum.-Comput. Interact. 35 2019 427 439
18 Jan Y.F. Tseng K.W. Kao P.Y. Hung Y.P. Augmented Tai-Chi Chuan practice tool with pose evaluation Proceedings - 4th International Conference on Multimedia Information Processing and Retrieval, MIPR 2021 2021 Graduate Institute of Networking and Multimedia, National Taiwan University Taipei, Taiwan 35 41
19 Tarek O. Magdy O. Atia A. Yoga trainer for beginners via machine learning Proceedings of the 2021 International Japan-Africa Conference on Electronics, Communications, and Computations, JAC-ECC 2021, HCI-LAB 2021 Faculty of Computers and Artificial Intelligence, Helwan University Egypt 75 78
20 Shi D. Jiang X. Sport Training Action Correction by Using Convolutional Neural Network 2021 Xinzhou Teachers Univ Xinzhou, Peoples R China AD
21 Huang X. Pan D. Huang Y. Deng J. Zhu P. Shi P. Xu R. Qi Z. He J. Intelligent yoga coaching system based on posture recognition Proceedings - 2021 International Conference on Culture-Oriented Science and Technology, ICCST 2021 2021 Communication University of China, School of Information and Communication Engineering Beijing, China 290 293
22 Yang L. Li Y. Zeng D. Wang D. Human exercise posture analysis based on pose estimation IEEE Advanced Information Technology, Electronic and Automation Control Conference (IAEAC) 2021 Sichuan Sports Industry Group Tanma Ai Technology Co. Ltd Chengdu, China 1715 1719
23 Ranasinghe I. Yuan C. Dantu R. Albert M.V. A collaborative and adaptive feedback system for physical exercises Proceedings - 2021 IEEE 7th International Conference on Collaboration and Internet Computing, CIC 2021 2021 University of North Texas, Computer Science and Engineering Denton, TX, United States 11 15
24 Jeon H. Kim D. Kim J. Human motion assessment on mobile devices International Conference on ICT Convergence, vol. 2021-October 2021 Electronics and Telecommunications Research Institute, Intelligent Robotics Research Division Deajeon, South Korea 1655 1658
25 Zhou J. Feng W. Lei Q. Liu X. Zhong Q. Wang Y. Jin J. Gui G. Wang W. Skeleton-based human keypoints detection and action similarity assessment for fitness assistance 2021 6th International Conference on Signal and Image Processing, ICSIP 2021 2021 Ningbo University of Technology Ningbo, China 304 310
26 Singh A. Le B.T. Nguyen T.L. Whelan D. O'Reilly M. Caulfield B. Ifrim G. Interpretable Classification of Human Exercise Videos Through Pose Estimation and Multivariate Time Series Analysis 2022 Output Sports Limited, NovaUCD Dublin, Ireland 181 199
27 Gu Y. Pandit S. Saraee E. Nordahl T. Ellis T. Betke M. Home-based physical therapy with an interactive computer vision system Proceedings - 2019 International Conference on Computer Vision Workshop, ICCVW 2019 Boston University, United States 2019 2619 2628
28 Niu Y. She J. Xu C. A survey on imu-and-vision-based human pose estimation for rehabilitation 2022 41st Chinese Control Conference (CCC) 2022 IEEE 6410 6415
29 Munea T.L. Jembre Y.Z. Weldegebriel H.T. Chen L. Huang C. Yang C. The progress of human pose estimation: a survey and taxonomy of models applied in 2d human pose estimation IEEE Access 8 2020 133330 133348
30 Desmarais Y. Mottet D. Slangen P. Montesinos P. A review of 3d human pose estimation algorithms for markerless motion capture Comput. Vis. Image Underst. 212 2021 103275
31 Wang J. Tan S. Zhen X. Xu S. Zheng F. He Z. Shao L. Deep 3d human pose estimation: a review Comput. Vis. Image Underst. 210 2021 103225
32 Gamra M.B. Akhloufi M.A. A review of deep learning techniques for 2d and 3d human pose estimation Image Vis. Comput. 114 2021 104282
33 Zheng C. Wu W. Chen C. Yang T. Zhu S. Shen J. Kehtarnavaz N. Shah M. Deep learning-based human pose estimation: a survey ACM Comput. Surv. 56 1 2023 1 37
34 Chen H. Feng R. Wu S. Xu H. Zhou F. Liu Z. 2d human pose estimation: a survey Multimed. Syst. 29 5 2023 3115 3138
35 Caramiaux B. Françoise J. Liu W. Sanchez T. Bevilacqua F. Machine learning approaches for motor learning: a short review Front. Comput. Sci. 2 2020 online, available: https://www.frontiersin.org/articles/10.3389/fcomp.2020.00016
36 Rajšp A. Fister I. A systematic literature review of intelligent data analysis methods for smart sport training Appl. Sci. 10 9 2020 3013
37 Gámez Díaz R. Yu Q. Ding Y. Laamarti F. El Saddik A. Digital twin coaching for physical activities: a survey Sensors 20 20 2020 1 21
38 Tsiouris K.M. Tsakanikas V.D. Gatsios D. Fotiadis D.I. A review of virtual coaching systems in healthcare: closing the loop with real-time feedback Front. Digit. Health 2 2020 567502 34713040
39 Lauber B. Keller M. Improving motor performance: selected aspects of augmented feedback in exercise and health Eur. J. Sport Sci. 14 1 2014 36 43 24533493
40 Zhou Y. De Shao W. Wang L. Effects of feedback on students' motor skill learning in physical education: a systematic review Int. J. Environ. Res. Public Health 18 12 2021 6281 34200657
41 Mödinger M. Woll A. Wagner I. Video-based visual feedback to enhance motor learning in physical education—a systematic review Ger. J. Exerc. Sport Res. 2021 1 14
42 Page M.J. McKenzie J.E. Bossuyt P.M. Boutron I. Hoffmann T.C. Mulrow C.D. Shamseer L. Tetzlaff J.M. Akl E.A. Brennan S.E. The prisma 2020 statement: an updated guideline for reporting systematic reviews Syst. Rev. 10 1 2021 1 11 33388080
43 Kohl C. McIntosh E.J. Unger S. Haddaway N.R. Kecke S. Schiemann J. Wilhelm R. Online tools supporting the conduct and reporting of systematic reviews and systematic maps: a case study on cadima and review of existing tools Environ. Evid. 7 1 2018 1 17
44 Magill R. Anderson D. Motor Learning and Control 2010 McGraw-Hill Publishing New York
45 Cai Z. Fernando O.N.N. Ong J.Y. Posebuddy: pose estimation workout mobile application 2022 International Conference on Cyberworlds (CW) 2022 IEEE 151 154
46 Taware G. Kharat R. Dhende P. Jondhalekar P. Agrawal R. Ai-based workout assistant and fitness guide 2022 6th International Conference on Computing, Communication, Control and Automation (ICCUBEA 2022 IEEE 1 4
47 Yue C.Z. Yong L. Shyan L. Exercise quality analysis using ai model and computer vision J. Eng. Sci. Technol. 2022 157 171
48 Bi R. Gao D. Zeng X. Zhu Q. Lazier: a virtual fitness coach based on ai technology 2022 IEEE 5th International Conference on Information Systems and Computer Aided Education (ICISCAE) 2022 IEEE 207 212
49 Singh V. Patade A. Pawar G. Hadsul D. trainer-an ai fitness coach solution 2022 IEEE 7th International Conference for Convergence in Technology (I2CT) 2022 IEEE 1 4
50 Dedhia U. Bhoir P. Ranka P. Kanani P. Pose estimation and virtual gym assistant using mediapipe and machine learning 2023 International Conference on Network, Multimedia and Information Technology (NMITCON) 2023 IEEE 1 7
51 Wu Y. Yu L. Xu J. Deng D. Wang J. Xie X. Zhang H. Wu Y. Ar-enhanced workouts: exploring visual cues for at-home workout videos in ar environment Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology 2023 1 15
52 Bernardo J.S. Divinagracia E.F. Go K.K. Determining exercise form correctness in real time using human pose estimation Proceedings of the 2023 5th International Conference on Image, Video and Signal Processing 2023 83 88
53 Ajay L. Biradar V.G. Chandu M. Bharath J. Ai tool as a fitness trainer using human pose estimation 2023 International Conference on Network, Multimedia and Information Technology (NMITCON) 2023 IEEE 1 12
54 Mandic S. Tracy R. Sra M. Arfit: pose-based exercise feedback with mobile ar Proceedings of the 2023 ACM Symposium on Spatial User Interaction 2023 1 3
55 Wei C. Wen J. Bi R. Yang H. Tao Y. Fan Y. He Y. Cheng D. Zhang Y. Online 8-form Tai Chi Chuan training and evaluation system based on pose estimation 2022 IEEE 24th Int Conf on High Performance Computing & Communications; 8th Int Conf on Data Science & Systems; 20th Int Conf on Smart City; 8th Int Conf on Dependability in Sensor, Cloud & Big Data Systems & Application (HPCC/DSS/SmartCity/DependSys) 2022 IEEE 366 371
56 Lavanya Y. Rajalakshmi N. Sumanth K. Gowrishankar S. Asha Rani K.P.S. A novel approach for developing inclusive real-time yoga pose detection for health and wellness using raspberry pi 2023 7th International Conference on Computation System and Information Technology for Sustainable Solutions (CSITSS) 2023 IEEE 1 7
57 Negi S. Garg M. Maindola H. Kansal V. Jain U. Bhatla S. Real-time human pose estimation: a mediapipe and python approach for 3d detection and classification 2023 3rd International Conference on Technological Advancements in Computational Sciences (ICTACS) 2023 IEEE 128 133
58 Anuradha M. Rao A. Shreyas S. Sanjaya K. Real time virtual yoga tutor 2023 IEEE 8th International Conference for Convergence in Technology (I2CT) 2023 IEEE 1 4
59 Elavarasi S.A. Kumar P.A. Jayanthi J. Development of ai-based posture monitoring system to assist yoga training 2023 International Conference on Research Methodologies in Knowledge Management, Artificial Intelligence and Telecommunication Engineering (RMKMATE) 2023 IEEE 1 4
60 Jian C.Z. Abdullah J. Lenando H. Dl-shuttle: badminton coaching training assistance system using deep learning approach 2022 International Conference on Digital Transformation and Intelligence (ICDI) 2022 IEEE 1 7
61 Beltrán Beltrán R. Richter J. Köstermeyer G. Heinkel U. Climbing technique evaluation by means of skeleton video stream analysis Sensors 23 19 2023 8216 37837046
62 Yamei L. Li G. Qiang G. Dynamic light collection system based on human posture estimation application in martial arts action teaching simulation Opt. Quantum Electron. 56 3 2024 376
63 Blanchet J. Hillis M.E. Lee Y. Shao Q. Zhou X. Kraemer D.J. Balkcom D. Learnthatdance: augmenting tiktok dance challenge videos with an interactive practice support system powered by automatically generated lesson plans Adjunct Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology 2023 1 4
64 Liu Y. Zhang T. Li Z. Deng L. Deep learning-based standardized evaluation and human pose estimation: a novel approach to motion perception Trait. Signal 40 5 2023
65 Wang G. Zhu B. Fan Y. Wu M. Wang X. Zhang H. Yao L. Sun Y. Su B. Ma Z. Design and evaluation of an exergame system to assist knee disorders patients' rehabilitation based on gesture interaction Health Inf. Sci. Syst. 10 1 2022 20 36032777
66 Garg B. Postlmayr A. Cosman P. Dey S. Short: deep learning approach to skeletal performance evaluation of physical therapy exercises Proceedings of the 8th ACM/IEEE International Conference on Connected Health: Applications, Systems and Engineering Technologies 2023 168 172
67 Zheng K. Wu J. Zhang J. Guo C. A skeleton-based rehabilitation exercise assessment system with rotation invariance IEEE Trans. Neural Syst. Rehabil. Eng. 2023
68 Pereira B. Cunha B. Viana P. Lopes M. Melo A.S. Sousa A.S. A machine learning app for monitoring physical therapy at home Sensors 24 1 2023 158 38203019
69 Movella MVN link online, available: https://www.movella.com/products/motion-capture/xsens-mvn-link 2023
70 Cao Z. Hidalgo Martinez G. Simon T. Wei S. Sheikh Y.A. Openpose: realtime multi-person 2d pose estimation using part affinity fields IEEE Trans. Pattern Anal. Mach. Intell. 2019
71 Cao Z. Simon T. Wei S.-E. Sheikh Y. Realtime multi-person 2d pose estimation using part affinity fields online, available: https://arxiv.org/abs/1611.08050 2016
72 Tome D. Russell C. Agapito L. Lifting from the deep: convolutional 3d pose estimation from a single image Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2017 2500 2509
73 Lugaresi C. Tang J. Nash H. McClanahan C. Uboweja E. Hays M. Zhang F. Chang C.-L. Yong M.G. Lee J. Mediapipe: a framework for building perception pipelines arXiv preprint arXiv:1906.08172 2019
74 TensorFlow PoseNet online, available: https://github.com/tensorflow/tfjs-models/tree/master/posenet 2019
75 TensorFlow MoveNet online, available: https://www.tensorflow.org/hub/tutorials/movenet 2024
76 Sandler M. Howard A. Zhu M. Zhmoginov A. Chen L.-C. Mobilenetv2: inverted residuals and linear bottlenecks Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2018 4510 4520
77 Xiao B. Wu H. Wei Y. Simple baselines for human pose estimation and tracking Proceedings of the European Conference on Computer Vision (ECCV) 2018 466 481
78 He K. Zhang X. Ren S. Sun J. Deep residual learning for image recognition Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2016 770 778
79 Posecamera, online, available: https://github.com/Wonder-Tree/PoseCamera 2021
80 Osokin D. Real-time 2d multi-person pose estimation on cpu: lightweight openpose arXiv preprint arXiv:1811.12004 2018
81 Iskakov K. Burkov E. Lempitsky V. Malkov Y. Learnable triangulation of human pose Proceedings of the IEEE/CVF International Conference on Computer Vision 2019 7718 7727
82 Bochkovskiy A. Wang C.-Y. Liao H.-Y.M. Yolov4: optimal speed and accuracy of object detection arXiv preprint arXiv:2004.10934 2020
83 Vision framework — apply computer vision algorithms to perform a variety of tasks on input images and video, online, available: https://developer.apple.com/documentation/vision 2023
84 Mehta D. Rhodin H. Casas D. Fua P. Sotnychenko O. Xu W. Theobalt C. Monocular 3d human pose estimation in the wild using improved cnn supervision 2017 International Conference on 3D Vision (3DV) 2017 IEEE 506 516
85 Park J. Cho S. Kim D. Bailo O. Park H. Hong S. Park J. A body part embedding model with datasets for measuring 2D human motion similarity IEEE Access 9 2021 36547 36558
86 Berndt D. Clifford J. Using dynamic time warping to find patterns in time series Workshop on Knowledge Discovery in Databases Seattle, WA, USA vol. 398 1994 359 370 no. 16, online, available: http://www.aaai.org/Papers/Workshops/1994/WS-94-03/WS94-03-031.pdf
87 Savitzky A. Golay M.J. Smoothing and differentiation of data by simplified least squares procedures Anal. Chem. 36 8 1964 1627 1639
88 Tinetti M.E. Williams T. Franklin Mayewski R. Fall risk index for elderly patients based on number of chronic disabilities Am. J. Med. 80 3 1986 429 434 3953620
89 Dempster A. Petitjean F. Webb G.I. ROCKET: exceptionally fast and accurate time series classification using random convolutional kernels Data Min. Knowl. Discov. 34 5 2020 1454 1495
90 Zabalbeascoa P. The nature of the audiovisual text and its parameters Didact. Audiovis. Transl. 7 2008 21 37
91 Gleicher M. Albers D. Walker R. Jusufi I. Hansen C.D. Roberts J.C. Visual comparison for information visualization Inf. Vis. 10 4 2011 289 309
92 Tharatipyakul A. Choo K.T. Perrault S.T. Pose estimation for facilitating movement learning from online videos Proceedings of the International Conference on Advanced Visual Interfaces 2020 1 5
93 Srikaewsiew T. Khianchainat K. Tharatipyakul A. Pongnumkul S. Kanjanawattana S. A comparison of the instructor-trainee dance dataset using cosine similarity, Euclidean distance, and angular difference 2022 26th International Computer Science and Engineering Conference ICSEC, Sakon Nakhon, Thailand 2022 235 240 10.1109/ICSEC56337.2022.10049368
94 Buekers M.J. Magill R.A. Hall K.G. The effect of erroneous knowledge of results on skill acquisition when augmented information is redundant Q. J. Exp. Psychol. Sect. A 44 1 1992 105 117
