
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)12425-5
10.1016/j.heliyon.2024.e36394
e36394
Research Article
Effects of grading rubrics on EFL learners’ writing in an EMI setting
Alghizzi Talal Musaed TMAlghizzi@imamu.edu.sa
⁎
Alshahrani Tahani Munahi
Department of English Language and Literature, College of Languages and Translation, Imam Mohammad Ibn Saud Islamic University, Saudi Arabia
⁎ Corresponding author. TMAlghizzi@imamu.edu.sa
15 8 2024
30 9 2024
15 8 2024
10 18 e3639427 6 2023
14 8 2024
14 8 2024
© 2024 The Authors
2024
https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
Despite considerable evidence that supports the use of grading rubrics (GRs) as tools for written corrective feedback, there is a paucity of research that investigates which of the different types of GRs best develops learners' International English Language Testing System (IELTS) writing scores in English as a medium of instruction (EMI)-contested settings. This study attempted to explore which rubric types (i.e., holistic, ESL composition profile, correction code, and IELTS) best assist English as a Foreign Language (EFL) learners in writing proficiency and which type leads to improving IELTS scores when such practice is embedded in EMI-disputed settings. Therefore, 351 male and female Saudi EFL learners were recruited to participate voluntarily. These participants were distributed equally among four groups corresponding to rubric type. For almost four months, the participants were exposed to a process-genre approach in which they were required to draft topics based on the comments received from their peer colleagues and teacher. The comments provided depended on the rubrics specified for their group type. The participants' pretest, midterm, and posttest scores were analyzed using one-way analysis of variance, t-tests, and paired samples t-tests. The results revealed that the ESL composition profile developed gradually, followed by the correction code group. However, the holistic groups did not improve. The tests were also assessed by specialists using the IELTS rubric. The findings revealed that the IELTS groups outperformed the other groups in all tests, followed by the female group in the ESL composition profile in the posttest. Meanwhile, other groups failed to improve. We discussed the results, considering the importance of GRs for improving EFL learners’ scores. Finally, we outlined the pedagogical implications for writing teachers in EMI settings. This study aimed to contribute to the growing research on EMI in relation to GRs, especially in the context of tertiary education in Saudi Arabia.

Keywords

Correction code
EMI
ESL composition profile
Holistic
IELTS
Written corrective feedback
==== Body
pmc1 Introduction

In recent years, researchers in the English as a Foreign Language (EFL) field have identified that time exposure to English, environmental contextual factors, composition teaching approaches, and feedback may play essential roles in developing EFL learners’ writing skills [1]. EFL teachers worldwide need to be equipped with recent feedback types, application strategies, and optimal practices to improve the process of teaching and learning English skills in general and in writing [2]. However, in the literature on English as a medium of instruction (EMI) settings, most of—if not all—the written production of EFL learners has been assessed using two methods: written corrective feedback (WCF) and grading rubrics (GRs). Studies investigating these types of writing assessments showed contradictory results of positive and negative effects [[3], [4], [5], [6], [7], [8], [9], [10], [11], [12]]. The issue with the WCF-focused research papers was that not only did they focus on the correction of metalinguistic aspects of the language (i.e., accuracy and grammar), but they also incorporated several WCF application strategies, such as the ones (e.g., direct, indirect, focused, unfocused, and reformulation) summarized by Ellis [13]. In addition, Bitchener and Ferris [3] highlighted other methodological issues such as “the absence of a control group, the failure to compare initial texts with new texts, the limited assessment of longitudinal effectiveness, and inconsistency in the selection of consistent instrumentation such as the type or genre of writing tasks” [5,7,8].

Al-Johani [14] concluded that Saudi EFL teachers neither provide students with authentic real-life examples nor welcome their ideas; instead, they offer immediate WCF and criticism. Furthermore, Knoch and Chapelle [15] argued that “when students make drafts or revise their works, they are usually left alone without any guidelines from the teacher” (p. 19). He stated that teachers most commonly do not follow up on their students' work to determine whether they have improved. This can hinder students’ English academic development. Li and Vuono [4] revealed that studies on WCF published in System over the past 25 years discussed issues related to comparing different groups that receive different types of WCF, the effectiveness of oral and written CF, and the evidence for supporting positive WCF. These studies provide insights into the pathway for the current study to investigate issues regarding the different types of WCF.

De Boer et al. [11] stated that EFL learners value the use of rubrics as a tool for WCF due to the “transparency in grading that is given by the teacher.” They motivate students and encourage participation and learner-centered learning since GRs provide students with the reason for the given grade. The GR studies had their own issues. Although GRs may comprise several fundamental aspects of writing, such as organization, content, style, spelling, word choice, and mechanics, they are dependent on the writing instructors’ subjectivities [11]. To the best of our knowledge, studies [[16], [17], [18], [19]] have yet to embark on applying GRs (i.e., holistic, ESL composition profile, correction code, and International English Language Testing System (IELTS) to EFL Saudi students using an application strategy (i.e., direct) specified for WCF and comparing their effectiveness in improving IELTS writing scores. For this reason, the purpose of the article is twofold: to compare the effectiveness of the different types of WCF by using some specified GRs to apply them to male and female Saudi students for future use to improve the writing skills of EFL students, and then to compare their effectiveness in developing IELTS scores.

2 Literature review

2.1 Theoretical foundations for WCF

The noticing theory is the theoretical argument that supports the use of WCF as an educational tool. It stated that it is only through conscious attention that the language they receive can be converted into learners' intake; noticing this process is a relevant step in language learning. Schmidt [20] related the importance of noticing to the process of WCF by emphasizing that when learners are aware of the gap between what they can produce and what they need to produce, as well as between what they produce and what target language speakers produce, such awareness would activate attention to the current interlanguage knowledge. Schmidt [20] asserted that learners’ conscious attention to the correct linguistic forms would enable them to identify the error and reconstruct the correct form of the target language. In this way, WCF functions as a noticing facilitator that helps learners bridge the gap between their interlanguage and the target language.

2.2 Rubrics as a tool of WCF

Improving EFL learners’ writing proficiency is a vital topic that attracts EFL researchers around the globe. Using rubrics as a tool for teaching and assessing student writing skills continues to be a most commonly discussed topic in the education field today [21]. De Boer et al. [11] defined a rubric as “an instrument for guidance and/or assessment that has the shape of a table with evaluation criteria to the left and a scale with different levels of achievement on top, plus descriptors in each cell. With these descriptors, each criterion and level (of performance) is made explicit.” A rubric is considered a supportive tool to assist students with self-assessment and peer evaluation. As for teachers, rubrics provide guidance on how to assess and grade, offering an efficient means of providing precoded feedback. Furthermore, Brooks [21] noted that rubrics reveal what is required in writing and define levels of student writing performance. He also pointed out that rubrics can be used both for teaching and evaluation. According to Mahmoudi and Buğra [22], rubrics serve as corrective feedback tools to assist students with their writing assignments. Instructional rubrics usually have two common features: (1) a list of criteria and (2) degrees of quality.

De Boer et al. [11] confirmed that students value the use of rubrics as a tool for WCF due to the “transparency in grading that is given by the teacher.” They are considered a motivating tool since they provide students with the reason for the given grade. Finson and Ormsbee [23] classified rubrics into two types: holistic or analytic. Holistic rubrics are used to assess the overall quality of the work, whereas analytic rubrics are used to assess different parts of the work based on established criteria. In this study, one holistic rubric, along with three types of analytic rubrics: ESL composition profile, correction codes, and IELTS, were used as WCF tools to improve the writing proficiency and IELTS writing scores of EFL learners. In this study, these different types of GR were used because the English language and translation college requires their members to employ WCF and to use one of the suggested GRs to effectively improve students’ writing. These types were selected specifically for their suitability to the requirements of the college, as they were proposed by the quality assurance unit to all faculty members. Below is a detailed description of these rubrics.

2.2.1 Holistic rubric

According to Brookhart [24], holistic rubrics generally provide students with an overall score for the written work. Teachers usually depend on a set of criteria to correct such work and judge its quality by giving a single score for the whole work. Hosseini and Mowlaie [25] encouraged the use of holistic scoring, as learners are bothered by the number of mistakes marked, and the rubric emphasizes only what was correctly done. In addition, the holistic rubric's greatest benefit is that teachers can correct many papers in a relatively short time. Other researchers also support its use [12,26,27]. Conversely, Brookhart [24] and Li [28] criticized this rubric in that it may not provide students with specific details of their mistakes and, thus, will not reflect the quality of the written work. Martin-Kniep [29] also demonstrated that teachers value elements of writing differently; hence, students' scores are subjectively assessed.

2.2.2 Analytic rubric

In this rubric, correcting and evaluating students’ work is based on scoring criteria. Each element of writing is included in the rubric with a specific score [28]. Bitchener and Ferris [3] noted that analytic WCF rubrics should include the basic elements of writing, such as content, organization, discourse, syntax, vocabulary, and mechanics. To be specific, content is related to the thesis statement and the arrangement and description of ideas. The organization consists of the introduction, sequence of ideas, and number of words. Furthermore, discourse elements cover topic sentences, paragraph unity, transition, discourse makers, reference, fluency, and variation. In addition, mechanics is related to spelling, punctuation, citation of references, and appearance. Three types of analytic rubrics are outlined below.

2.2.3 ESL composition profile

Jacobs et al.‘s [30] ESL composition profile is a scoring rubric that contains the basic elements of writing as developed by the authors. In this rubric, scores are obtained from a 100-point scale covering content (30 points), organization (20 points), vocabulary (20 points), language use (25 points), and mechanics (5 points). Each element is scored on one of four levels of proficiency to be judged subjectively by the raters. Lee, Gentile, and Kantor [31] asserted that this type of scoring tool is reliable, as it is validated among raters and researchers in the ESL field and is widely used by teachers [17,32] and ESL researchers [[33], [34], [35]]. Recently, researchers have used this rubric to assess the writing proficiency of EFL learners due to its popularity and validity [17,36]. According to Heidari et al. [37], Jacobs et al.‘s [30] ESL composition profile is the most commonly used scale and a well-designed rubric to assess L2 writing.

2.2.4 Correction codes

Correction code rubrics are tools for providing feedback in which learners are responsible for using their knowledge to correct the editing symbols highlighted and underlined by the teachers. Learners participate in self-correction and discover the correct alternatives [38]. The correction code rubric type is considered an effective feedback tool that guides students’ writing, provides reference for error types, and allows students to review and rewrite their work correctly [39]. Different studies have supported the use of correction code rubrics [[38], [39], [40], [41], [42]].

2.2.5 IELTS

IELTS is an international standardized test to measure the English language proficiency of nonnative English speakers. It consists of four parts that represent one of the four macroseius: speaking, listening, reading, and writing. The test assessment rubric is divided into nine bands from 1 to 9, representing the lowest and highest proficiency scores, respectively. Candidates should score at least 6 or 7 to be accepted for academic study at a higher education institution in an English-speaking country. The IELTS writing section is divided into two tasks, with higher scores rewarded for Writing Task 2. In Task 2, candidates are required to write a well-developed piece using the appropriate content, style, register, and organization [43]. The test is then assessed by trained and certified examiners using a specific rubric. Few studies have compared the different types of WCF rubrics to improve EFL learners' performance in IELTS. For instance, Sanavi and Nemati [18] explored the role of six different WCFs and error correction code rubrics (i.e., reformulation, direct form, indirect form, metalinguistic, peer correction, and error coding) to prepare EFL learners for their IELTS exams and demonstrated that the reformulation group achieved the highest scores in the IELTS’ writing section. Table 1 compares the four GRs used in the study.Table 1 Comparison of the four grading rubrics.

Table 1Type	Scoring	Features	Division of scores	Focus	
Holistic	Holistic scoring	Holistic rubrics are used to assess the overall quality of the work. Holistic rubrics are best to use when there is no single correct answer or response and the focus is on overall quality, proficiency, or understanding of a specific content or skills.	Holistic rubrics generally provide students with an overall score for the written work	Teachers usually depend on a set of criteria to correct such work and judge its quality by giving a single score for the whole work	
ESL composition profile	Analytic scoring	Analytic rubrics are used to assess different parts of the work based on established criteria. They are particularly useful for problem-solving or application assessments because a rubric can list a different category for each component of the assessment that needs to be included, thereby accounting for the complexity of the task.	100-point scale covering content (30 points), organization (20 points), vocabulary (20 points), language use (25 points), and mechanics (5 points)	It is a scoring rubric that contains the basic elements of writing as developed by the authors	
Correction codes	Analytic scoring	Grammar: 17 pts.
Spelling: 17 pts.
Word choice: 17 pts.
Organization: 16 pts.
Style: 16 pts.
Clarity: 17 pts.	Correction code rubrics are tools for providing feedback in which learners are responsible for using their knowledge to correct the editing symbols highlighted and underlined by the teachers. Learners participate in self-correction and discover the correct alternatives	
IELTS	Analytic scoring	See appendix D	IELTS is an international standardized test to measure the English language proficiency of nonnative English speakers	

2.3 EMI in the saudi context

EMI is viewed as a practical program to improve EFL learners' English proficiency and, therefore, result in a workforce that is more fluent in English. It refers to the use of English as the language of instruction in nonspeaking countries. The reason behind the introduction of EMI to higher education is most commonly related to the apparent advantages of this instructional methodology. The driving forces for EMI growth in EFL settings are not only its widely known benefits but also EMI's ability to improve graduates' job prospects and enhance institutional ranking and global competitiveness [44]. The term EMI used in this study refers to the teaching of a subject using the medium of the English language, but where there are no explicit language learning aims and where English is not the national language. Depending on this perspective, different studies support the use of EMI in the Saudi context. For example, Al-Kahtany, Faruk, and Al Zumor [45] conducted research to examine students' and teachers' attitudes toward EMI programs. The data collected from questionnaires and structured interviews showed positive attitudes toward the programs for their advantages in improving EFL students' four skills of English language learning. Similarly, Shamim et al. [46] investigated teachers' and learners' experiences and perceptions about the use of EMI in a preparatory year program. The interviews revealed that most teachers and students with higher proficiency levels favored the use of EMI.

2.4 EFL studies on using rubrics to improve writing proficiency

Mahmoudi and Buğra [22] noted that although many studies explored the effectiveness of rubrics in assessing EFL students writing [12,47,48], few studies [11,22,49,50] investigated the effectiveness of using instructional rubrics in teaching and learning writing skills. Qian [51] examined the use of process-oriented rubrics and explored the different variables that may affect the use of rubrics, such as English language proficiency, motivation to write, and previous experiences with writing programs. The analysis of the rubrics used by students and teachers and the students' written reflections revealed that the rubrics helped students practice the writing process. Mahmoudi and Buğra [22] conducted a recent study to explore how using rubrics in teaching writing skills can affect the writing performance of students. The qualitative analysis of the results revealed that using rubrics in teaching writing improved the students’ writing performance.

Some of the aforementioned studies [25,33,52] compared holistic and analytic GRs, and others investigated the effectiveness of one type of GR [38,39,41,49]. Le et al. [53] stated that gender-based studies may contribute to provide significant results in the WCF field, especially in an EFL setting. Few studies have investigated the differences between genders using GRs [[53], [54], [55]]. However, previous studies have not provided sufficient information regarding which types of GRs are more effective and their contribution to improving the IELTS band score in an EMI environment. What is ostensibly missing from the bulk of research in the Saudi context is an all-inclusive study to investigate the appropriate type of GRs for improving EFL learners’ writing proficiency overall in IELTS. Such findings can provide practical advice for EFL writing teachers in selecting the most effective GR and assisting EFL learners to achieve good scores in IELTS.

This study tackled the following research questions.1. What is the most effective GR for developing EFL learners' writing in general?

2. Which type of GR contributes to improving IELTS scores the most?

3 Methods and materials

3.1 Research design and setting

As this quasi-experimental quantitative study required specific types of participants, writing courses, flexibility in application, and time management, it was conducted in the second semester of the academic year 2021–2022 in an EMI setting. The experiment lasted for approximately four months. The implementation of English as the only medium of instruction in an EFL classroom setting in this study is to figure out the factors that lead to controversy over the rule of applying only EMI among EFL students by using different GRs as a mean of WCF. The researchers applied the rule of EMI to writing classes in the English Education Program. These rules were adopted from Prabjandee and Nilpirom [56].1. Embracing EMI with the “Right” attitude: a) English is the language used for instructional purposes, b) English is not itself the subject to be taught, c) English is emancipated from the native English speakers norm, d) English is owned by nonnative English speakers.

2. EMI lesson planning: Thus, with careful planning, an EMI lesson should be relevant to learners by using motivating materials that foster the real-life application of concepts studied. To achieve such a lesson, it is essential to explore learners' needs by using a variety of methods.

3. Executing EMI lessons: The authors suggested that EMI teachers help learners connect what they know to what they are learning, assist them in problem-solving, and promote retention of newly learned information.

4. Engaging in reflective practice refers to practitioners critically examining their philosophy, principles, theories, and techniques to take responsibility for their actions.

3.2 Participants

The participants were third-year students in an EMI program in the English Language and Literature Department and were recruited after official approval from the department chair. We recruited 762 students enrolled in Essay Writing II (Eng218). They were all Saudi students, and their ages ranged from 19 to 22. An email explaining the research topic, procedures, requirements, confidentiality, participation, and prospective benefits was sent to all students. Students interested in participating were required to sign an electronic consent form with their names and academic numbers, which allowed us access to their university transcripts and any other related information. Participants were divided evenly into four groups based on the following criteria: a) having passed all writing courses in their first attempt with a score between 71 and 85, which corresponds to the advanced level of the Common European Framework of Reference for Languages (CEFR); b) enrolled in Eng218 for the first time; and c) having not dropped or failed any other course before. Another email was sent to those who agreed to participate, suggesting the time and date for their first argumentative writing task or pretest, which was allotted 40 min for completion. After receiving 351 written samples, a preliminary holistic score of 100 was assigned to each by a rater (see the Analysis section for details). The scores were categorized based on the following group types: holistic group, ESL composition profile group, code correction group, and IELTS group. The scores were further analyzed with the analysis of variance (ANOVA) to identify any significant differences between the participants and across groups and genders. Table 2 displays all groups’ pretest preliminary holistic score comparison results, which showed no significant differences.Table 2 All groups’ pretest preliminary holistic score comparison results.

Table 2	All groups	Male	Female	
N	Mean	SD	N	Mean	SD	N	Mean	SD	
Holistic	88	50.04	25.08	24	42.83	24.75	64	52.75	24.85	
ESL composition profile	88	45.40	22.05	24	38.08	17.66	64	48.15	23.01	
Correction code	88	47.18	22.39	24	48.45	18.13	64	46.70	23.91	
IELTS	87	45.93	21.94	24	43.08	19.69	63	47.00	22.78	
F (ANOVA) and p-value	F (ANOVA) = .721, p-value = .540	F (ANOVA) = 1.052, p-value = .374	F (ANOVA) = .898, p-value = .443	

3.3 Teaching approach

During the experiment, all participants were enrolled in Eng218 and studied three chapters on argumentative, classification, and reaction essays from the mandatory textbook Effective Academic Writing by Liss and Davis [57]. These chapters were taught using one teaching approach in the EMI setting (i.e., process genre) to all participants by two fellow colleagues who agreed to participate voluntarily. These fellow colleagues were assistant professors of Applied Linguistics, specializing in teaching writing. Such a teaching approach is required not only to provide ad hoc samples of written texts to be analyzed in terms of sentences, words, audience, register, conventions, and others but also to write multiple drafts for every assigned topic [58]. On a regular weekly basis, each participant was asked to write an essay of 250 words on a topic that falls under the category of the previously mentioned chapters. Their first written drafts were corrected by one of their peers and designated external raters (see the section below) based on an individualized GR. The fundamental difference between the groups was that the GR sheet filled out by peers and raters in each group and discussed in class by the instructors with their students matched the group type.

3.4 Instruments

The instruments incorporated to collect the data were a pretest, a midterm test, and a posttest. In each of these tests, the students were asked to write an argumentative essay of no less than 250 words within 40 min. The time allotted for execution and the topic of the prompts were based on and/or adapted from the second question in the writing section of the IELTS, as follows.1. Some people think that the news media has a great effect on people's lives, which is a negative development. Do you agree or disagree with this statement? Give your opinions and relevant examples.

2. Many parents believe that COVID-19 vaccinations are required for students before they can attend public schools. Do you agree or disagree with this statement? Give your opinions and relevant examples.

3. Some Facebook users state that the platform should be allowed to collect data from its users to be able to improve. Do you agree or disagree with this statement? Give your opinions and relevant examples.

3.5 Data collection procedures and analysis

The written samples of the pretest, midterm test, and posttest reached 1404 essays, which were collected during the semester at weeks one, eight, and sixteen. Essays and writing assignments were corrected and commented on by experienced external raters based on the group-selected GR criteria (see Appendices A, B, C, and D). The only exception was that the rater designated for the IELTS group corrected all the tests of the other group types with the results used in the analysis. Significance was determined when p ≤ .05. Any significant increase in GR-type scores entailed its effectiveness and vice versa. Details of the raters are as follows.1. The holistic group rater was a nonnative assistant professor of Applied Linguistics specializing in teaching writing with 11 years of experience.

2. The ESL composition profile group rater was a nonnative associate professor of Applied Linguistics, specializing in teaching writing, with 13 years of experience.

3. The code correction group rater was a nonnative associate professor of Applied Linguistics, specializing in teaching writing, with 18 years of experience.

4. The IELTS group rater was a native assistant professor of English who was a trained and certified examiner of IELTS for 9 years.

4 Results

4.1 Descriptive statistics

The groups' descriptive statistics, as shown in Table 3, revealed that the highest mean belonged to the ESL composition profile group's posttest.Table 3 Descriptive statistics for all groups.

Table 3		N	Minimum	Maximum	Mean	SD	
Holistic	Pretest	88	10.00	95.00	50.04	25.08	
Midterm	88	10.00	95.00	53.21	24.93	
Posttest	88	14.00	95.00	59.70	24.83	
Pretest (IELTS)	88	.00	9.00	4.93	3.06	
Midterm (IELTS)	88	.00	9.00	5.05	2.78	
Posttest (IELTS)	88	.00	9.00	4.34	2.87	
ESL composition profile	Pretest	88	10.00	91.00	44.20	21.55	
Midterm	88	10.00	95.00	56.17	23.61	
Posttest	88	17.00	95.00	71.40	19.82	
Pretest (IELTS)	88	.00	9.00	4.51	2.79	
Midterm (IELTS)	88	.00	9.00	4.12	2.63	
Posttest (IELTS)	88	.00	9.00	5.82	2.70	
Correction code	Pretest	88	10.00	88.00	38.54	19.32	
Midterm	88	10.00	95.00	51.94	25.55	
Posttest	88	12.00	93.00	62.37	22.80	
Pretest (IELTS)	88	.00	9.00	4.14	2.91	
Midterm (IELTS)	88	.00	9.00	4.36	2.80	
Posttest (IELTS)	88	.00	9.00	4.42	2.90	
IELTS	Pretest	87	.00	9.00	3.95	2.43	
Midterm	87	.00	9.00	4.88	2.71	
Posttest	87	.00	9.00	6.19	2.54	

Table 4 shows the descriptive statistics for both men and women. The male group had the highest mean score on the ESL composition profile posttest. Similar results were found for the female groups, where the highest mean score was for the ESL composition profile posttest.Table 4 Descriptive statistics for the male and female groups.

Table 4	Male	Female	
		N	Mean	SD	N	Mean	SD	
Holistic	Pretest	24	42.83	24.75	64	52.75	24.85	
Midterm	24	52.08	25.17	64	53.64	25.02	
Posttest	24	55.45	25.82	64	61.29	24.47	
Pretest (IELTS)	24	5.29	2.86	64	4.79	3.14	
Midterm (IELTS)	24	4.29	2.92	64	5.34	2.70	
Posttest (IELTS)	24	4.37	2.94	64	4.32	2.86	
ESL composition profile	Pretest	24	38.50	18.55	64	46.34	22.34	
Midterm	24	61.45	20.55	64	54.18	24.52	
Posttest	24	81.87	9.63	64	67.48	21.24	
Pretest (IELTS)	24	3.79	2.87	64	4.78	2.74	
Midterm (IELTS)	24	3.87	2.57	64	4.21	2.67	
Posttest (IELTS)	24	5.20	2.85	64	6.06	2.63	
Correction code	Pretest	24	34.54	11.43	64	40.04	21.43	
Midterm	24	58.20	22.91	64	49.59	26.26	
Posttest	24	69.95	19.59	64	59.53	23.40	
Pretest (IELTS)	24	4.95	2.820	64	3.84	2.918	
Midterm (IELTS)	24	4.04	2.85	64	4.48	2.80	
Posttest (IELTS)	24	5.00	3.21	64	4.20	2.77	
IELTS	Pretest	24	4.41	2.20	63	3.77	2.510	
Midterm	24	4.87	2.57	63	4.88	2.78	
Posttest	24	6.95	1.85	63	5.90	2.72	

4.2 Pretest

Table 5 shows the results of the ANOVA test for the differences between groups (holistic, ESL composition profile, and correction code) in the pretests. The results revealed that there were no statistically significant differences between them.Table 5 Pretest ANOVA results.

Table 5	All groups	Male	Female	
N	Mean	SD	N	Mean	SD	N	Mean	SD	
Holistic	88	50.04	25.08	24	42.83	24.75	64	52.7500	24.85	
ESL composition profile	88	44.20	21.55	24	38.50	18.55	64	46.3438	22.34	
Correction code	88	38.54	19.32	24	34.54	11.43	64	40.0469	21.43	
F (ANOVA) and p-value	F (ANOVA) = 5.94, p-value = .053	F (ANOVA) = 1.13, p-value = .326	F (ANOVA) = 4.91, p-value = .058	

4.3 Holistic groups

Table 6 shows the results of the paired sample test for the differences between holistic groups (pretest-midterm) and (midterm-posttest). The results indicated that there was a statistically significant difference between them because of the significant difference found in the female sample in favor of the posttest. Otherwise, no significant differences were found between the tests.Table 6 Paired sample results for the holistic groups.

Table 6	Mean	N	SD	t	p-value	
All groups	Pretest	50.04	88	25.08	−.829	.409	
Midterm	53.21	88	24.93	
Midterm	53.21	88	24.93	−2.115	.037*	
Posttest	59.70	88	24.83	
Male	Pretest	42.83	24	24.75	−1.161	.258	
Midterm	52.08	24	25.17	
Midterm	52.08	24	25.17	−.540	.594	
Posttest	55.45	24	25.82	
Female	Pretest	52.75	64	24.85	−.205	.838	
Midterm	53.64	64	25.02	
Midterm	53.64	64	25.02	−2.170	.034*	
Posttest	61.29	64	24.47	
Note: *: significance at .05.

4.4 ESL composition profile groups

Table 7 shows the results of the paired sample test for the differences between the ESL composition profile groups (pretest-midterm) and (midterm-posttest). The results demonstrate that there was a statistically significant difference between pretest and midterm in favor of midterm in both male and female samples. Furthermore, there was a statistically significant difference between the midterm and posttest in favor of the posttest in both male and female samples.Table 7 Paired sample results for ESL composition profile groups.

Table 7	Mean	N	SD	t	p-value	
All groups	Pretest	44.20	88	21.55	−3.82	.000**	
Midterm	56.17	88	23.61	
Midterm	56.17	88	23.61	−5.09	.000**	
Posttest	71.40	88	19.82	
Male	Pretest	38.50	24	18.55	−4.80	.000**	
Midterm	61.45	24	20.55	
Midterm	61.45	24	20.55	−4.58	.000**	
Posttest	81.87	24	9.63	
Female	Pretest	46.34	64	22.34	−2.06	.043*	
Midterm	54.18	64	24.52	
Midterm	54.18	64	24.52	−3.54	.001**	
Posttest	67.48	64	21.24		
Note: **: significance at .01, *: significance at .05.

4.5 Correction code groups

Table 8 shows the results of the paired sample test for the differences between correction code groups (pretest-midterm) and (midterm-posttest). The results indicated a statistically significant difference between pretest and midterm in favor of midterm in both male and female samples. Furthermore, there was a statistically significant difference between the midterm and posttest in favor of the posttest in both male and female samples.Table 8 Paired sample results for code correction groups.

Table 8	Mean	N	SD	t	p-value	
All groups	Pretest	38.54	88	19.32	−3.79	.000**	
Midterm	51.94	88	25.55	
Midterm	51.94	88	25.55	−3.19	.002*	
Posttest	62.37	88	22.80	
Male	Pretest	34.54	24	11.43	−4.21	.000**	
Midterm	58.20	24	22.91	
Midterm	58.20	24	22.91	−2.19	.038*	
Posttest	69.95	24	19.59	
Female	Pretest	40.04	64	21.43	−2.22	.030*	
Midterm	49.59	64	26.26	
Midterm	49.59	64	26.26	−2.46	.017*	
	Posttest	59.53	64	23.40			
Note: **: significance at .01, *: significance at .05.

4.6 IELTS results for the holistic groups

Table 9 shows the results of the paired sample test for the differences between the results of the IELTS test of the holistic groups (pretest-midterm) and (midterm-posttest). The results revealed no statistically significant difference between pretest-midterm or midterm-posttest in all groups, which indicated that they did not improve.Table 9 Paired sample results for the IELTS test for the holistic groups.

Table 9	Mean	N	SD	t	p-value	
All groups	Pretest	4.93	88	3.06	−.282	.779	
Midterm	5.05	88	2.78	
Midterm	5.05	88	2.78	1.711	.091	
Posttest	4.34	88	2.87	
Male	Pretest	5.29	24	2.86	1.326	.198	
Midterm	4.29	24	2.92	
Midterm	4.29	24	2.92	−.117	.908	
Posttest	4.37	24	2.94	
Female	Pretest	4.79	64	3.14	−1.025	.309	
Midterm	5.34	64	2.70	
Midterm	5.34	64	2.70	2.003	.051	
Posttest	4.32	64	2.86	

4.7 IELTS results for the ESL composition profile groups

Table 10 shows the results of the paired sample test for the differences between ESL composition profile groups in the IELTS results (pretest-midterm) and (midterm-posttest). The results identified a statistically significant difference between midterm-posttest groups due to the female groups’ results, which revealed that the female group improved in the posttest. However, there were no statistically significant differences between the groups.Table 10 Paired sample IELTS results for the ESL composition profile groups.

Table 10	Mean	N	SD	t	p-value	
All groups	Pretest	4.51	88	2.79	.906	.367	
Midterm	4.12	88	2.63	
Midterm	4.12	88	2.63	−4.468	.000**	
Posttest	5.82	88	2.70	
Male	Pretest	3.79	24	2.87	−.098	.922	
Midterm	3.87	24	2.57	
Midterm	3.87	24	2.57	−1.465	.156	
Posttest	5.20	24	2.85	
Female	Pretest	4.78	64	2.74	1.136	.260	
Midterm	4.21	64	2.67	
Midterm	4.21	64	2.67	−4.583	.000**	
Posttest	6.06	64	2.63	
Note: **: significance at .01.

4.8 IELTS results for the correction code groups

Table 11 shows the results of the paired sample test for the differences between the IELTS results for the correction code groups (pretest-midterm) and (midterm-posttest). The results demonstrated that there were no statistically significant differences between pretest-midterm or midterm-posttest in all groups.Table 11 Paired sample IELTS results for the correction code groups.

Table 11	Mean	N	SD	t	p-value	
All groups	Pretest	4.14	88	2.91	−.490	.625	
Midterm	4.36	88	2.80	
Midterm	4.36	88	2.80	−.125	.901	
Posttest	4.42	88	2.90	
Male	Pretest	4.95	24	2.82	1.158	.259	
Midterm	4.04	24	2.85	
Midterm	4.04	24	2.85	−1.051	.304	
Posttest	5.00	24	3.21	
Female	Pretest	3.84	64	2.91	−1.228	.224	
Midterm	4.48	64	2.80	
Midterm	4.48	64	2.80	.541	.590	
Posttest	4.20	64	2.77	

4.9 IELTS groups

Table 12 shows the results of the paired sample test for the differences between the IELTS groups (pretest-midterm) and (midterm-posttest). The results indicated a statistically significant difference between the groups. There was a statistically significant difference for the male group between the midterm and posttest, in favor of the posttest, whereas no statistically significant differences were found between the pretest and midterm. In addition, the results revealed a statistically significant difference between the pretest and midterm in favor of the midterm test for the female group. Furthermore, there was a statistically significant difference between the midterm and posttest, in favor of the posttest.Table 12 Paired sample results for IELTS group.

Table 12	Mean	N	SD	t	p-value	
All groups	Pretest	3.95	87	2.43	−2.92	.004**	
Midterm	4.88	87	2.71	
Midterm	4.87	87	2.73	−3.56	.001**	
Posttest	6.18	87	2.55	
Male	Pretest	4.41	24	2.20	−1.18	.246	
Midterm	4.87	24	2.57	
Midterm	4.82	24	2.62	−4.73	.000**	
Posttest	6.95	24	1.89	
Female	Pretest	3.77	63	2.51	−2.68	.009**	
Midterm	4.88	63	2.78	
Midterm	4.88	63	2.78	−2.15	.035*	
Posttest	5.90	63	2.72	
Note: **: significance at .01, *: significance at .05.

4.10 ANOVA results for all groups

To investigate the scores difference between all GR groups, One way ANOVA test was performed. Results showed that there was a statistically significant difference between all groups with (p-value <.05). Thus, to know which specific groups differ, we used the multiple comparisons table with the results of the LSD post hoc test. Results showed that the ESL composition profile group outperformed other groups followed by the correction codes group. See Table 13.Table 13 ANOVA results for the comparison of all GR groups.

Table 13	Group	Mean	Std. Deviation	F	P-value	
Grading rubric results for all groups	Holistic (N = 88)	59.70	24.83	6.49	.002	
ESL Composition Profile (N = 88)	71.40	19.82	
Correction codes (N = 88)	62.37	22.80	
Total (N = 264)	64.49	23.04	

Table 14 explores the scores difference for the IELTS results between all groups. The results for the one way ANOVA test showed that there was a statistically significant difference between all groups with (p-value <.05). Thus, to know which specific groups differ, we used the multiple comparisons table with the results of the LSD post hoc test. Results showed that the IELTS group outperformed other groups followed by the ESL composition group. The holistic group had the lowest mean score.Table 14 ANOVA results for the comparison of all groups for the IELTS scores.

Table 14	Group	Mean	Std. Deviation	F	p-value	
IELTS results for all groups	Holistic (N = 88)	4.34	2.87	10.456	.001	
ESL Composition Profile (N = 88)	5.82	2.70	
Correction codes (N = 88)	4.42	2.90	
IELTS (N = 87)	6.19	2.54	
	Total	5.19	2.87	

5 Discussion

This study investigated the effects of four different types of GRs in an EMI setting on improving Saudi EFL learners' writing performance in general and on IELTS writing scores. The results indicated that among holistic groups, men exhibited no improvement, whereas women improved in the posttest but not significantly. In addition, there was no significant improvement in IELTS scores among the holistic groups. Although previous research on GR argued the benefits of using a holistic approach to enhancing learners’ writing proficiency [37,53,[59], [60], [61]], the disadvantages of holistic scoring have also been noted [62,63]. Holistic scoring was reported to be subjective and dependent, to some extent, on the mood of the teacher [62,63]. Moreover, scores do not reflect the types or number of mistakes; therefore, learners may use them as a reference for WCF. Hosseini and Mowlaie [25] commented, “it is also difficult to interpret the meaning of a composite score to the raters and to the users” (p. 33).

The results of the ESL composition profile groups reflected significant positive changes for both male and female learners. Setyowati et al. [17] reported that the ESL composition profile is an efficient and practical rubric to improve EFL learners’ writing. Turgut and Kayaoğlu [19] compared the written productions of 16 EFL learners in an experimental group exposed to the ESL composition profile of 22 EFL learners in a control group and concluded that the former outperformed the latter. The remarkable improvement results in this study can be attributed to the fact that this rubric is most commonly used among EFL university educators, which makes it popular and well-known among students. It is easy to follow for students, as they have already been trained to conform to the rubric standards.

Although EFL learners improved significantly when using this rubric, the improvement in IELTS scores was not significant, with men exhibiting no improvement in all tests and women showing improvement in the posttest only. This may be because these groups were not instructed according to IELTS rubrics, which is in line with Alghizzi's [64,65] findings that the type of mistake identification and correction was restricted to the rubric elements EFL learners were introduced to. Other factors, such as learners' engagement with the feedback, nature of errors, and type of WCF, which were examined in this study, may have affected the lack of improvement in IELTS scores [37,66].

The results for correction codes demonstrated significant improvement among men and women in all the tests. Comparing all groups in all tests, correction code groups outperformed holistic groups, which revealed that using correction codes gradually improved learners' writing proficiency. This was in line with Hosseiny [67], Buckingham and Aktuğ-Ekinci [39], and Lee et al. [68], who concluded that using correction codes significantly improved EFL learners' writings. However, the use of correction codes did not yield positive results for IELTS. This could be attributed to what Yu et al. [16] referred to as the negative effects of WCF, which suggest that EFL learners might respond negatively to positive feedback and that relevant factors, such as students' individual differences and writing needs, were neglected when providing feedback. Thus, when teachers assessed learners’ essays, they seemed to focus on fixing the errors mentioned in the correction code rubric without paying attention to errors that were not included there. Accordingly, learners worked on reformulating these errors specifically, which led to high scores when comparing the different types of GRs. As for comparing the GR groups, the results showed that the ESL composition profile group had the highest scores followed by the correction codes group. This is because these two rubrics are commonly used by EFL university teachers as mentioned earlier. As for the results of the IELTS scores, it was evident that the IELTS group had the highest mean score followed by the ESL composition profile group. The holistic group achieved the lowest mean score which can be attributed to the nature of this GR since it does not provide specific feedback for improvement.

It is difficult to determine whether EMI can improve students' English. According to Galloway, Kriukow and Numajiri [69] and Le et al. [53], average IELTS scores can be considered an indicator of students’ increase or decrease in English language proficiency. Therefore, indications of improvement in the IELTS scores of EFL students in this study can be attributed to the effect of the types of GR in the EMI setting. This study did not include the IELTS group in the comparison of the effectiveness of rubrics due to the difference and incompatibility of IELTS scores with other tests (i.e., IELTS = 9, others = 100). After correcting the tests of all four groups based on IELTS GR, the results revealed that the IELTS group achieved the highest score, followed by the ESL composition profile group. The gender differences revealed that the male group developed only in this posttest, whereas the female group developed in all tests. The other groups did not improve in any of the tests. EFL researchers have confirmed that EMI plays a fruitful role in how much English students learn and how much input they receive [37]. The results were consistent with Sanavi and Nemati [18], who examined the role of different WCF types in preparing EFL learners for the IELTS test and found that students improved significantly when using different types of WCF.

6 Theoretical and pedagogical implications

While previous GR research has provided copious findings that advance our understanding of this field, we believe that there is a need to expand the scope of investigation by combining the application of GRs and WCF strategies with EFL learners. Certain theoretical and pedagogical implications are drawn from this study. Its main theoretical implication is that it revolves around language and reformulates writing errors under the umbrella of the interaction approach and sociocultural theory. Interaction theory stresses the importance of learning a language through input, output, and feedback. The sociocultural theory assumes that language learning is the product of communicative pressure that connects communication with acquisition [3]. GRs are a form of communication that tells learners what is unacceptable when learning a language and how to monitor and modify their output. Sanavi and Nemati [18] stated that during the process of WCF, learners could identify the divergence between the current state of knowledge and the target language.

From a pedagogical perspective, the feedback field is directly relevant to frontline writing teachers. Teachers can benefit from the results of the research findings and apply them in classroom settings. The results for female and male undergraduates did not increase in holistic feedback but increased gradually in the ESL composition profile, correction code feedback, and IELTS rubrics [[17], [18], [19]]. This study revealed that gender differences are not a major factor in reflecting the differences between men and women. Some researchers argue, however, that even when average performance is equal, gender discrepancies may still exist at the highest levels of writing ability [11]. Therefore, it is suggested to conduct another study for professional students in English. For example, it is advised to compare the variability in female’ and male’ writing scores. This implies that students' writing can be developed if they are guided through written feedback rubrics, regardless of gender. Additionally, the findings of this study revealed that the ESL composition profile and the code correction groups improved in group tests, but not when their essays were assessed using the IELTS rubric. This implies that the type of rubrics imposed on students determines the type of development that they will exhibit, as emphasized by Alghizzi [64,65]. In both studies, he maintained that the types of written mistakes EFL learners could identify and correct accordingly were those introduced and taught to them by their teachers. Therefore, if EFL writing instructors and faculties in English-major departments intend to develop their students’ writing abilities in IELTS, they can either adopt IELTS rubrics designated for writing or modify the ESL composition profile and/or the correction code rubrics according to the Common European Framework [70].

7 Conclusion

To conclude, GRs are considered one of the most vibrant streams in EMI research. Although teachers play a pivotal role in appropriately implementing GRs in a way that benefits EFL learners. It demonstrated that different types of GRs provide different results. It showed that the IELTS groups outperformed the other groups in all tests, followed by the female group in the ESL composition profile in the posttest. Meanwhile, other groups failed to improve. We discussed the results, considering the importance of GRs for improving EFL learners’ scores. This study aimed to contribute to the growing research on EMI in relation to GRs, especially in the context of tertiary education in Saudi Arabia.

Most importantly, improving students' writing proficiency is a complex issue that cannot be changed overnight. It is the teacher's responsibility to employ inclusive pedagogies, diversify the use of GRs and WCF application strategies, and create supportive learning environments that maximize students' engagement in writing and revision.

8 Recommendations

The findings suggest that teachers should diversify the use of both GRs and WCF to ensure comprehensiveness. Lee [71] confirmed that teachers should do this according to learners’ needs, error types, proficiency levels, and application strategies. Additionally, teachers should focus on certain error types, depending on the level of students. This will allow them to meet the focal requirement for an authentic classroom environment. In addition, the results reveal that teachers should, regardless of the WCF type, provide prewriting and postwriting instruction in support of feedback, allow learners to ask questions about the WCF received, and provide ample time to discuss errors.

By employing the aforementioned GRs, current hot spots and research trends were objectively and systematically presented. According to the results of this review, researchers may be encouraged to do more qualitative studies and use a more scientific and convincing measurement approach. Language teachers can have a better understanding of WCF and adjust their feedback strategies in an EMI setting. Therefore, future research should continue to explore WCF to provide academic support for EFL learners, educators, and the development of education.

9 Limitations

This study has some limitations that can be addressed by future research to enrich the field of GRs and WCF. First, it only examined gender differences belonging to one academic level at one university. To widen its scope, future studies need to apply this approach to different universities and academic levels. Second, this study compared only four GR types with participants. Thus, future studies can compare other types, vary their use and applications, and compare the results with high-stakes tests, such as the Test of English as a Foreign Language in an EMI environment. Finally, students received GRs for argumentative essays. Hence, more studies could investigate other types of essays, such as classification, narrative, and reaction essays.

Funding

This study was not funded by any organizations.

10 Employment

No organization will gain or lose from the publication of this study.

Financial interests

There is no financial interest of any kind.

10.1 Non-financial interests

There is no non-financial interest of any kind.

11 Ethics statement

This study was reviewed and approved by the university ethical board, with the approval number: 638,225,421,184,066,978. All participants provided informed consent to participate in the study.

12 Data availability statement

The research data are available on Figshare Repository at https://figshare.com/s/c4647c530dad352a64e3.

CRediT authorship contribution statement

Talal Musaed Alghizzi: Writing – review & editing, Writing – original draft. Tahani Munahi Alshahrani: Writing – review & editing, Writing – original draft.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Appendix A Holistic rubric

The overall writing score		

Appendix B ESL Composition Profile Criteria and Scores

	Level	Criteria	Score and comments	
Content	30–27
26–22
21–17
16–13	EXCELLENT TO VERY GOOD: knowledgeable • substantive • thorough development of thesis • relevant to assigned topic
GOOD TO AVARAGE: some knowledge of subject • adequate range • limited development of thesis • mostly relevant to topic, but lacks detail
FAIR TO POOR: limited knowledge of subject • little substance • inadequate development of topic
VERY POOR: does not show knowledge of subject • non-substantive • not pertinent • OR not enough to be evaluated		
Organization	20–18
17–14
13–10
9–7	EXCELLENT TO VERY GOOD: fluent expression • ideas clearly stated/supported • succinct • well-organized • logical sequencing • cohesive
GOOD TO AVARAGE: somewhat choppy • loosely organized but main ideas stand out • limited support • logical but incomplete sequencing
FAIR TO POOR: non-fluent • ideas confused or disconnected • lacks logical sequencing and development
VERY POOR: does not communicate • no organization • OR not enough to be evaluated		
Vocabulary	20–18
17–14
13–10
9–7	EXCELLENT TO VERY GOOD: sophisticated range • effective word/idiom choice and usage • word form mastery • appropriate register
GOOD TO AVARAGE: adequate range • occasional errors of word/idiom form, choice, usage but meaning not obscured
FAIR TO POOR: limited range • frequent errors of word/idiom form, choice, usage • meaning confused or obscured
VERY POOR: essentially translation • little knowledge of English vocabulary, idioms, word form • OR not enough to be evaluated		
Language use	25–22
21–18
17–11
10–5	EXCELLENT TO VERY GOOD: effective complex constructions • few errors of agreement, tense, number, word order/function, articles, pronouns, prepositions
GOOD TO AVARAGE: effective but simple constructions • minor problems in complex constructions • several errors of agreement, tense, number, word order/function, articles, pronouns, prepositions but meaning seldom obscured
FAIR TO POOR: major problems in simple/complex constructions • frequent errors of negation, agreement, tense, number, word order/function, articles, pronouns, prepositions and/or fragments, run-ons, deletions • meaning confused or obscured
VERY POOR: virtually no mastery of sentence construction rules • dominated by errors • does not communicate • OR not enough to be evaluated		
Mechanics	5
4
3
2	EXCELLENT TO VERY GOOD: demonstrates mastery of conventions • few errors of spelling, punctuation, capitalization, paragraphing
GOOD TO AVARAGE: occasional errors of spelling, punctuation, capitalization, paragraphing but meaning not obscured
FAIR TO POOR: frequent errors of spelling, punctuation, capitalization, paragraphing • poor handwriting • meaning confused or obscured
VERY POOR: no mastery of conventions • dominated by errors of spelling, punctuation, capitalization, paragraphing • handwriting illegible • OR not enough to evaluate		

Appendix C Correction codes rubric

Category	Score	Code	Content	
Grammar	17	G	All kinds of grammatical mistakes	
Spelling	17	Sp	Spelling mistakes	
Word Choice	17	WCh	Using the appropriate word	
Organization	16	O	Thesis statement - topic sentence - paragraph	
Style	16	S	Nativelike-effective by L1	
Clarity	17	C	Correct phrases and clauses, proper punctuation
Applying correct language conventions and usage	

Appendix D IELTS Rubric. Writing Task 2: Band Descriptors

IELTS is jointly owned by the British Council, IDP: IELTS Australia and the University of Cambridge ESOL Examinations (Cambridge ESOL).

https://www.ielts.org/-/media/pdfs/writing-band-descriptors-task-2.ashx.	Task achievement	Coherence and cohesion	Lexical resource	Grammatical range and accuracy	Comments	
9	-fully addresses all the parts of the task
- presents a fully developed position in answer to the question with relevant, fully extended and well supported ideas	- uses cohesion in such a way that it attracts no attention
- skillfully manages paragraphing	- uses a wide range of vocabulary with very natural and sophisticated control of lexical features; rare minor errors occur only as “slips”	-uses a wide range of structures with full flexibility and accuracy; rare minor errors occur only as “slips”		
8	-sufficiently addresses all parts of the task
- presents a well-developed response to the question with relevant, extended and supported ideas	- sequences information and ideas logically
- manages all aspects of cohesion well
- uses paragraphing sufficiently and appropriately	- uses a wide range of vocabulary fluently and flexibly to convey precise meanings
- skillfully uses uncommon lexical items but there may be occasional inaccuracies in word choice and collocation
- produces rare errors in spelling and/or word formation	- uses a wide range of structures
- the majority of sentences are error-free
- makes only very occasional errors or inappropriacies		
7	-addresses all parts of the task
-presents a clear position throughout the response
-presents, extends and supports main ideas, but there may be a tendency to over-generalize and/or supporting ideas may lack focus	-logically organizes information and ideas; there is clear progression throughout -uses a range of cohesive devices appropriately although there may be some under-/over-use
-presents a clear central topic within each paragraph	-uses a sufficient range of vocabulary to allow some flexibility and precision
-uses less common lexical items with some awareness of style and collocation
-may produce occasional errors in word choice, spelling and/or word formation	-uses a variety of complex structures
-produces frequent error-free sentences
-has good control of grammar and punctuation but may make a few errors		
6	-addresses all parts of the task although some parts may be more fully covered than others
-presents a relevant position although the conclusions may become unclear or repetitive
-presents relevant main ideas but some may be inadequately developed/unclear	-arranges information and ideas coherently and there is a clear overall progression
-uses cohesive devices effectively, but cohesion within and/or between sentences may be faulty or mechanical • may not always use referencing clearly or appropriately
-uses paragraphing, but not always logically	-uses an adequate range of vocabulary for the task
-attempts to use less common vocabulary but with some inaccuracy
-makes some errors in spelling and/or word formation, but they do not impede communication	-uses a mix of simple and complex sentence forms
-makes some errors in grammar and punctuation but they rarely reduce communication		
5	-addresses the task only partially; the format may be inappropriate in places
-expresses a position but the development is not always clear and there may be no conclusions drawn
-presents some main ideas but these are limited and not sufficiently developed; there may be irrelevant detail	-presents information with some organization but there may be a lack of overall progression
-makes inadequate, inaccurate or over-use of cohesive devices
-may be repetitive because of lack of referencing and substitution
-may not write in paragraphs, or paragraphing may be inadequate	-uses a limited range of vocabulary, but this is minimally adequate for the task
-may make noticeable errors in spelling and/or word formation that may cause some difficulty for the reader	-uses only a limited range of structures
-attempts complex sentences but these tend to be less accurate than simple sentences
-may make frequent grammatical errors and punctuation may be faulty; errors can cause some difficulty for the reader		
4	-responds to the task only in a minimal way or the answer is tangential; the format may be inappropriate
-presents a position but this is unclear
-presents some main ideas but these are difficult to identify and may be repetitive, irrelevant or not well supported	-presents information and ideas but these are not arranged coherently and there is no clear progression in the response
-uses some basic cohesive devices but these may be inaccurate or repetitive
-may not write in paragraphs or their use may be confusing	-uses only basic vocabulary which may be used repetitively or which may be inappropriate for the task
-has limited control of word formation and/or spelling; errors may cause strain for the reader	-uses only a very limited range of structures with only rare use of subordinate clauses
-some structures are accurate but errors predominate, and punctuation is often faulty		
3	-does not adequately address any part of the task
-does not express a clear position
-presents few ideas, which are largely undeveloped or irrelevant	-does not organize ideas logically
- may use a very limited range of cohesive devices, and those used may not indicate a logical relationship between ideas	-uses only a very limited range of words and expressions with very limited control of word formation and/or spelling
-errors may severely distort the message	-attempts sentence forms but errors in grammar and punctuation predominate and distort the meaning		
2	-barely responds to the task
-does not express a position
-may attempt to present one or two ideas but there is no development	-has very little control of organizational features	-uses an extremely limited range of vocabulary; essentially no control of word formation and/or spelling	-cannot use sentence forms except in memorized phrases		
1	-answer is completely unrelated to the task	-fails to communicate any message	-can only use a few isolated words	-cannot use sentence forms et al.		
0	-does not attend
-does not attempt the task in any way
-writes a totally memorized response					

Acknowledgements

We would like to thank Elsevier (www.elsevier.com) for English language editing and SAGE for formatting editing service (https://languageservices.sagepub.com/en/).
==== Refs
References

1 Alghizzi T.M. Complexity, Accuracy, and Fluency (CAF) Development in L2 Writing: the Effects of Proficiency Level, Learning Environment, Text Type, and Time Among Saudi EFL Learners [Doctoral Dissertation] 2017 University College Cork
2 Joseph S. Rickett C. Northcote M. Christian B.J. ‘Who are you to judge my writing?’: student collaboration in the co-construction of assessment rubrics N. Writ Int. J. Pract. Theor. Creativ. 17 2020 31 49 10.1080/14790726.2019.1566368
3 Bitchener J. Ferris D.R. Written Corrective Feedback in Second Language Acquisition and Writing 2011 Routledge Oxfordshire 10.4324/9780203832400
4 Li S. Vuono A. Twenty-five years of research on oral and written corrective feedback in System System 84 2019 93 109 10.1016/j.system.2019.05.006
5 Moser A. Written Corrective Feedback: the Role of Learner Engagement: a Practical Approach 2020 Springer New York
6 Mujtaba S.M. Reynolds B.L. Parkash R. Singh M.K.M. Individual and collaborative processing of written corrective feedback affects second language writing accuracy and revision Assess. Writ. 50 2021 100566 10.1016/j.asw.2021.100566
7 Rezaei S. Corrective Feedback in Task-Based Grammar Instruction 2011 Lap Lambert Academic Publishing Sunnyvale, CA
8 Sheen Y. Corrective Feedback, Individual Differences and Second Language Learning 2011 Springer New York
9 Williams J. The potential role(s) of writing in second language development J. Sec Lang. Writ. 21 2012 321 331 10.1016/j.jslw.2012.09.007
10 Biber D. Nekrasova T. Horn B. The effectiveness of feedback for L1‐English and L2‐writing development: a meta-analysis ETS Res. Rep. Ser. 2011 2011 i 99 10.1002/j.2333-8504.2011.tb02241.x
11 de Boer I. de Vegt F. Pluk H. Latijnhouwers M.A.H.E. Rubrics – a Tool for Feedback and Assessment Viewed from Different Perspectives: Enhancing Learning and Assessment Quality 2021 Springer New York 10.1007/978-3-030-86848-2
12 Rezaei A.R. Lovorn M. Reliability and validity of rubrics for assessment through writing Assess. Writ. 15 2010 18 39 10.1016/j.asw.2010.01.003
13 Ellis R. A typology of written corrective feedback types ELT J. 63 2009 97 107 10.1093/elt/ccn023
14 Al-Johani H.M. Finding a Way Forward: the Impact of Teachers' Strategies, Beliefs and Knowledge on Teaching English as a Foreign Language in Saudi Arabia [Doctoral Dissertation] 2009 The University of Strathclyde
15 Knoch U. Chapelle C.A. Validation of rating processes within an argument-based framework Lang. Test. 35 2018 477 499 10.1177/0265532217710049
16 Yu S. Geng F. Liu C. Zheng Y. What works may hurt: the negative side of feedback in second language writing J. Sec Lang. Writ. 54 2021 100850 10.1016/j.jslw.2021.100850
17 Setyowati L. Sukmawan S. El-Sulukkiyah A.A. Exploring the use of ESL composition profile for college writing in the Indonesian context Int. J. Lang. Educ. 4 2020 171 182 10.26858/ijole.v4i2.13662
18 Sanavi R.V. Nemati M. The effect of six different corrective feedback strategies on Iranian English language learners' IELTS writing task 2 Sage Open 4 2014 10.1177/2158244014538271
19 Turgut F. Kayaoğlu M.N. Using rubrics as an instructional tool in EFL writing courses J. Lang. Linguist. Stud. 11 2015 47 58
20 Schmidt R. Attention Robinson P. Cognition and Second Language Instruction 2001 Cambridge University Press 3 32
21 Brooks C. Carroll A. Gillies R.M. Hattie J. A matrix of feedback Aust. J Teach. Educ. 44 2019 14 32 10.14221/ajte.2018v44n4.2
22 Mahmoudi F. Buğra C. The effects of using rubrics and face to face feedback in teaching writing skill in higher education Int. Online J. Educ. Teaching. 7 2020 150 158
23 Finson K.D. Ormsbee C.K. Rubrics and their use in inclusive science Interv. Sch. Clin. 34 1998 79 88 10.1177/105345129803400203
24 Brookhart S.M. How to Create and Use Rubrics for Formative Assessment and Grading 2013 ASCD Alexandria, VA
25 Hosseini M. Mowlaie B. Effect of holistic vs. analytic assessment on improving Iranian intermediate EFL learners' writing skill J. Lang Transl. 6 2016 31 41
26 Bitchener J. Young S. Cameron D. The effect of different types of corrective feedback on ESL student writing J. Sec Lang. Writ. 14 2005 191 205 10.1016/j.jslw.2005.08.001
27 Chandler J. The efficacy of various kinds of error feedback for improvement in the accuracy and fluency of L2 student writing J. Sec Lang. Writ. 12 2003 267 296 10.1016/S1060-3743(03)00038-9
28 Li W. Scoring rubric reliability and internal validity in rater-mediated EFL writing assessment: insights from many-facet Rasch measurement Read. Writ. 35 2022 2409 2431 10.1007/s11145-022-10279-1
29 Martin-Kniep G.O. Becoming a Better Teacher: Eight Innovations that Work 2000 ASCD Alexandria, VA
30 Jacobs H.L. Zinkgraf S.A. Wormuth D.R. Hartfiel V.F. Hughey J.B. Testing ESL Composition: a Practical Approach 1981 Newbury House Publishers New York
31 Lee Y.W. Gentile C. Kantor R. Analytic scoring of TOEFL® CBT essays: scores from humans and e‐rater ETS Res. Rep. Ser. 2008 2008 i 71 10.1002/j.2333-8504.2008.tb02087.x
32 Wang W. Using rubrics in student self-assessment: student perceptions in the English as a foreign language writing context Assess Eval. High Educ. 42 2017 1280 1292 10.1080/02602938.2016.1261993
33 Bacha N. Writing evaluation: what can analytic versus holistic essay scoring tell us? System 29 2001 371 383 10.1016/S0346-251X(01)00025-2
34 Ghanbari B. Barati H. Moinzadeh A. Rating scales revisited: EFL writing assessment context of Iran under scrutiny Lang. Test. Asia 2 2012 83 100 10.1186/2229-0443-2-1-83
35 Marzban A. Jalali F.E. The interrelationship among L1 writing skills, L2 writing skills, and L2 proficiency of Iranian EFL learners at different proficiency levels Theor. Pract. Lang. Stud. 6 2016 1364 1371 10.17507/tpls.0607.05
36 Al-Mudhi M.A. Evaluating Saudi university students' English writing skills using an analytic rating scale J. Appl. Linguist. Lang. Res. 6 2019 95 109
37 Heidari N. Ghanbari N. Abbasi A. Raters' perceptions of rating scales criteria and its effect on the process and outcome of their rating Lang. Test. Asia 12 2022 20 10.1186/s40468-022-00168-3
38 Sampson A. Coded and uncoded error feedback: effects on error frequencies in adult Colombian EFL learners' writing System 40 2012 494 504 10.1016/j.system.2012.10.001
39 Buckingham L. Aktuğ-Ekinci D. Interpreting coded feedback on writing: Turkish EFL students' approaches to revision J. Engl. Acad. Purp. 26 2017 1 16 10.1016/j.jeap.2017.01.001
40 Ferris D. Roberts B. Error feedback in L2 writing classes? J. Sec Lang. Writ. 10 2001 161 184 10.1016/S1060-3743(01)00039-X
41 Han Y. Hyland F. Exploring learner engagement with written corrective feedback in a Chinese tertiary EFL classroom J. Sec Lang. Writ. 30 2015 31 44 10.1016/j.jslw.2015.08.002
42 Tang C. Liu Y.T. Effects of indirect coded corrective feedback with and without short affective teacher comments on L2 writing performance, learner uptake and motivation Assess. Writ. 35 2018 26 40 10.1016/j.asw.2017.12.002
43 Green A. Washback to the learner: learner and teacher perspectives on IELTS preparation course expectations and outcomes Assess. Writ. 11 2006 113 134 10.1016/j.asw.2006.07.002
44 Dimova S. Hultgren A.K. Jensen C. English-Medium Instruction in European Higher Education vol. 4 2015 Walter de Gruyter GmbH Berlin
45 Al-Kahtany A.H. Faruk S.M.G. Al Zumor A.W.Q. English as the medium of instruction in Saudi higher education: necessity or hegemony? J. Lang. Teach. Res. 7 2016 49 58 10.17507/jltr.0701.06
46 Shamim F. Abdelhalim A. Hamid N. English medium instruction in the transition year: case from KSA Arab World Engl. J. 7 2016 32 47 10.24093/awej/vol7no1.3
47 Ghaffar M.A. Khairallah M. Salloum S. Co-constructed rubrics and assessment for learning: the impact on middle school students' attitudes and writing skills Assess. Writ. 45 2020 100468 10.1016/j.asw.2020.100468
48 Reynders G. Lantz J. Ruder S.M. Stanford C.L. Cole R.S. Rubrics to assess critical thinking and information processing in undergraduate STEM courses Int. J. STEM Educ. 7 2020 9 10.1186/s40594-020-00208-5
49 Bradford K.L. Newland A.C. Rule A.C. Montgomery S.E. Rubrics as a tool in writing instruction: effects on the opinion essays of first and second graders Early Child. Educ. J. 44 2016 463 472 10.1007/s10643-015-0727-0
50 Bui M.C. Vuong T.M.K. The effect of using instructional rubrics on EFL students' writing performance: a high school case in the Mekong Delta of Vietnam Eur. J. Engl. Lang. Teach. 7 2022 11 30 10.46827/ejel.v7i1.4112
51 Qian Y. Using rubrics in a university EFL process writing program: an exploratory case study ASIAN TEFL 3 2018 81 94
52 Klimova B.F. Evaluating writing in English as a second language Procedia Soc. Behav. Sci. 28 2011 390 394 10.1016/j.sbspro.2011.11.074
53 Le X.M. Phuong H.Y. Phan Q.T. Le T.T. Impact of using analytic rubrics for peer assessment on EFL students' writing performance: an experimental study Multicult. Educ. 9 2023 41 53 10.5281/zenodo.7750831
54 Alqarni T.M. The Effect of Peer Assessment on the EFL Students' Writing Skills in the Saudi Context: a Case Study [Doctoral Dissertation] 2021 King Abdulaziz University
55 Kiasi G.A. Rezaie S. The effect of peer assessment and collaborative assessment on Iranian intermediate EFL learners' writing ability J. Engl. Lang. Teach. Appl. Linguist. 3 2021 8 16 10.32996/jeltal.2021.3.13.2
56 Prabjandee D. Nilpirom P. Pedagogy in English-Medium Instruction (EMI): some recommendations for EMI teachers Reflections 29 2022 421 434
57 Liss R. Davis J. Effective Academic Writing: the Researched Essay second ed. 2012 Oxford University Press Oxford
58 Badger R. White G. A process genre approach to teaching writing ELT J. 54 2000 153 160 10.1093/elt/54.2.153
59 Barkaoui K. Rating scale impact on EFL essay marking: a mixed-method study Assess. Writ. 12 2007 86 107 10.1016/j.asw.2007.07.001
60 Charney D. The validity of using holistic scoring to evaluate writing: a critical overview Res. Teach. Engl. 18 1984 65 81
61 Knoch U. Rating scales for diagnostic assessment of writing: what should they look like and where should the criteria come from? Assess. Writ. 16 2011 81 96 10.1016/j.asw.2011.02.003
62 Halleck G.B. Assessing oral proficiency: a comparison of holistic and objective measures Mod. Lang. J. 79 1995 223 234 10.1111/j.1540-4781.1995.tb05434.x
63 Jafarpur A. Can naive EFL learners estimate their own proficiency? Eval. Res. Educ. 5 1991 145 157 10.1080/09500799109533306
64 Alghizzi T.M. The Role of English Writing Instruction Methodologies on the types of Written Mistakes/errors EFL Graduate Diploma Students Can Identify in Their Writings 2011 Dublin International Foundation College [Unpublished Graduate Diploma thesis]
65 Alghizzi T.M. The Role of English Writing Instruction Methodologies on the Types of Written Mistakes/errors Saudi EFL Pre-university Students Can Identify in Their Writings 2012 University College Cork [Unpublished Master’s thesis]
66 Al-Ahdal A.A.M.H. Alfallaj F.S. Al-Awaied S.A. Al-Hattami A.A. A comparative study of proficiency in speaking and writing among EFL learners in Saudi Arabia Am. Int. J. Contemp. Res. 4 2014 141 149
67 Hosseiny M. The role of direct and indirect written corrective feedback in improving Iranian EFL students' writing skill Procedia Soc. Behav. Sci. 98 2014 668 674 10.1016/j.sbspro.2014.03.466
68 Lee I. Luo N. Mak P. Teachers' attempts at focused written corrective feedback in situ J. Sec Lang. Writ. 54 2021 100809 10.1016/j.jslw.2021.100809
69 Galloway N. Kriukow J. Numajiri T. Internationalisation Higher Education and the Growing Demand for English: an Investigation into the English Medium of Instruction (EMI) Movement in China and Japan 2017 The British Council UK
70 Council of Europe, Council for Cultural Co-operation Common European Framework of Reference for Languages: Learning, Teaching, Assessment 2001 Cambridge University Press Cambridge, UK
71 Lee I. Utility of focused/comprehensive written corrective feedback research for authentic L2 writing classrooms J. Sec Lang. Writ. 49 2020 100734 10.1016/j.jslw.2020.100734
