
==== Front
Heliyon
Heliyon
Heliyon
2405-8440
Elsevier

S2405-8440(24)12811-3
10.1016/j.heliyon.2024.e36780
e36780
Research Article
Conceptions of Assessment Among Pakistani Teachers of English: Implications for Policy and Professional Development
Hidri Sahbi shidri@hct.ac.ae
a⁎
Aziz Musharraf musharrafazizkaifi@gmail.com
b
Qutub Manal mqutub@uj.edu.sa
c
a Department of Education, Higher Colleges of Technology, Abu Dhabi Campus, Abu Dhabi, United Arab Emirates
b University Utara Malaysia, Malaysia
c English Language Institute, University of Jeddah, Saudi Arabia
⁎ Corresponding author. shidri@hct.ac.ae
24 8 2024
15 9 2024
24 8 2024
10 17 e3678021 10 2022
20 8 2024
22 8 2024
© 2024 Published by Elsevier Ltd.
2024

https://creativecommons.org/licenses/by-nc-nd/4.0/ This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/).
In this study, we investigated conceptions of assessment among 735 English language teachers from different educational institutions and levels across Pakistan, with most holding a master’s degree. Using a validated inventory (Brown, 2006) administered through a random sampling technique, we explored the alignment between teachers’ assessment conceptions, rooted in socio-cognitive contexts, and the prevailing language policies enforced by educational authorities. Derived from the Structural Equating Model, our findings indicated that, despite the significant correlations between perceptions of student accountability and improvement (r = .92) and between school accountability and both student accountability (r = .59) and improvement (r = .55), the statistical and practical discrepancies in the loadings of factors and indicators limited the applicability and construct validity of the original inventory within the Pakistani context. These differences highlighted the mismatch between how assessments are perceived and how they are applied on the ground. In order to resolve inequities and provide teachers with greater professional support, we recommended policymakers recalibrate current assessments to address disparities and better support teachers professionally. The study contributed to the broader discourse on language classroom practices, suggesting pathways for further investigations so that educational outcomes could be maximized within socio-cognitive environments.

Keywords

Assessment
Student & teacher accountability
Improvement
Irrelevance
Structural Equating Model
==== Body
pmc1 Introduction

Assessment is fundamental to every teaching-learning scenario, serving as a crucial measure of student performance, teaching effectiveness, and program accountability. The exploration of teachers’ conceptions of assessment (TCoA) within their socio-cognitive (political, educational, and social) frameworks has garnered a lot of attention in the last three decades. TCoA shape relevant pedagogical practices ([1] Baidoo-Anu et al., 2023 [2]; Brown, 2006 [3]; Brown & Remesal, 2017 [4]; Brown et al., 2015 [5]; DeLuca et al., 2019 [6]; Deneen & Brown, 2016 [7]; Lutovac & Flores, 2022 [8]; Opre, 2015), affect language policy and planning ([9] Shohamy, 2007 [10]; Hidri et al., 2023), and impact stakeholders in a powerful and coercive way ([10] Hidri et al., 2023). These factors imply that the way teachers understand, perceive, and use assessments in educational contexts is influenced by a combination of social and cognitive factors. While the social factors (interpersonal, cultural, and institutional of where teachers work) are shaped by the norms of the school community, educational standards, and the ways the wider communities perceive assessments; the cognitive factors refer to the different processes that influence how teachers perceive assessments.

Research (e.g. Ref. [2], Brown, 2006 [11]; Gerbil & Brown, 2014 [12]; Hidri, 2016) suggests that the way diverse conceptions of assessment are perceived can critically determine the quality of teaching, student learning, performance, and program accountability. Other researchers (e.g. Ref. [13], Barnes et al., 2017) have argued that TCoA form the cornerstone of guiding classroom management, instructional methodologies, and subsequent assessments, where classroom assessment tools serve to a) determine student accountability, b) enhance student learning, and c) critically inform and shape instructional strategies and planning. Beyond the classroom, assessment tools serve to a) contribute to school accountability by providing information on what and how students process learning, b) play a gatekeeping role for students ([11] Gebril & Brown, 2014 [14]; Shohamy, 2001 [9], 2007), and c) provide information on how educational outcomes are achieved ([15] Takele & Melese, 2022). The nature of these assessment tools and their purposes is anchored in the practical realities of TCoA, which encapsulates the complexities of educational scenarios and often a de facto language policy ([10] Hidri et al., 2023 [14]; Shohamy, 2001). However, despite the presence of such assessment tools and purposes, teachers sometimes find it challenging to remain impartial when carrying out such assessments since their conceptions of assessment can contrast sharply with the status quo of language policies and the underlying intentions of policymakers ([10] Hidri et al., 2023). This interconnection underscores the mapping of TCoA with their assessment practices and how policymakers perceive assessment.

Since its independence, Pakistan has faced many challenges in establishing a clear second language policy in terms of curricula since initially, the social and educational scenarios were “complicated by languages and language groups competing to be recognized as national language” ([16] Mahboob, 2002, p. 20). The general conceptions about student assessments are associated with Assessment of Learning (AoL), which is used for gatekeeping purposes, and although the concept of Assessment for Learning (AfL) is gaining ground, most educational institutions, mainly in the public sector, are still employing traditional patterns of assessment. In general, language teachers in Pakistan attach much importance to assessment, even though they still view it as inflexible, norm-referenced, and impressionistic ([17] Warsi, 2004). In the current educational landscape, high academic achievement is predominantly associated with the curricula of elite private institutions that prioritize mathematical and English language skills ([18] Alderman et al., 2001). Conversely, state schools often use imprecise norm-referenced assessments of the oral and auditory language skills, i.e., speaking and listening. Historically, most educational institutions have emphasized the teaching of reading and writing in English, with the provision of English language conversation classes confined to private missionary schools, where offerings were shaped by prescribed curricula and rigorous standards.

Language teachers in Pakistan manifest two TCoA paradigms: Traditional and contemporary. These paradigms are intertwined in a complex context shaped a de facto language policy. According to Ref. [19] Coleman and Capstick (2012, p. 14), educational institutions and syllabi fall into several types: a) elite private schools with English as the core medium of instruction; b) armed forces’ schools with English as the medium of instruction; c) state public schools, where English is taught traditionally; d) non-elite private English-medium schools; and d) seminaries, with religious education of less focus on English. A recent addition to the educational scene is the branches of international schools that use English. This categorization has significantly impacted language assessment across the country [20]. Coleman (2010) contends that the Pakistani educational institutions primarily tend to train learners to pass second language examinations, where language teachers assess learners using a limited number of assessment tools such as reading English texts aloud, translating English passages into Urdu, writing meanings of difficult English vocabulary, constructing sentences from English words, and reproducing timed memorized essays and stories. Unfortunately, these types of assessments remain at odds with how learners should be assessed in, for instance, using more comprehensive constructs such as the assessment of communicative abilities or 21st-century soft skills. Other teachers adopt modern teaching methodologies, leveraging technology to foster students’ communicative abilities, by crafting outcome-driven English curricula that offer learners ample opportunities to become more participative through presentations, communicative activities, role-playing, debates, and student-oriented in-class discussions. However, despite these sporadic attempts, the perception of assessment has remained negative among teachers and students to the extent that common sense has labelled it as “a necessary evil”.

This study investigates conceptions of assessment and how they have been held by English language teachers in Pakistan, a critical yet underexplored area with significant implications for educational effectiveness and teacher development. The study aims to address the misalignment between current assessment conceptions, often rooted in outdated language policies, and the evolving needs of modern education in Pakistan. This misalignment raises concerns about the efficacy of English language teaching and the professional growth of teachers in the country. The diversity in TCoA among English language teaching practitioners underscores the need to empirically address these conceptions to evaluate the quality and relevance of assessment practices ([21] Horgan & Gardiner-Hyland, 2019), where such diversity can result in an unalignment between the objectives and learning outcomes ([22] Fletcher et al., 2012). We planned to investigate assessment conceptions among Pakistani teachers of English, by applying the already validated inventory of four-factor TCoA, student accountability, student accountability, improvement, and irrelevance ([2] Brown, 2006). Despite the abundant volume of research on TCoA in different contexts using this inventory (e.g. Ref. [2], Brown, 2006 [23], 2011 [5]; DeLuca et al., 2019 [6]; Deneen & Brown, 2016 [22]; Fletcher et al., 2012 [11]; Gebril & Brown, 2014 [12]; Hidri, 2016 [21]; Horgan & Gardiner-Hyland, 2019 [7]; Lutovac & Flores, 2022 [14]; Shohamy, 2001 [9], 2007), there remains a significant research gap in probing the status of language assessment among Pakistani teachers of English. We also sought to explore how such conceptions of assessment align with the Pakistani language policy requirements and how this policy should contribute to teacher professional development. While the existing literature predominantly focuses on TCoA at different educational levels in many contexts, such conceptions have not been given due momentum in an ESL and exam-driven context such as Pakistan ([24] Ahmad & Saeed, 2021). This study is unique in its focus on the specific context of Pakistan, with its high stakes testing culture and complex linguistic landscape, by providing a more nuanced understanding of TCoA. Additionally, the study goes beyond theoretical understanding to examine the practical implications of these conceptions for language policy and teacher professional development, aiming to bridge the gap between research and practice and offer actionable recommendations for educational reforms in Pakistan. The study then set out to answer the following questions:a. To what extent can teachers’ conceptions of assessment be admissible to the Pakistani context?

b. What implications does language policy have on teachers’ conceptions of assessment and their professional development in Pakistan?

2 Review of the literature

Language assessment is embedded in its social milieu, reflecting unremitting educational shifts towards the assessments of 21st-century competency-based soft skills. The interpretation of TCoA is inherently nuanced and complex, denoting multifaceted and controversial belief systems in different contexts, where such beliefs about what to assess, how to assess, why to assess, and when to assess ([25] Ferretti et al. (2021)) are shaped by the teaching and learning context. This nuanced context lays the ground for adopting frameworks, models, or inventories to probe TCoA. Brown’s inventory of TCoA (2006) is a case in point [26]. Brown (2004 [2]; 2006) designed a four-factor inventory to investigate TCoA, a) assessment for improvement purposes of both teaching and learning, b) assessment for accountability of the learner, c) assessment for accountability of the teacher and institution, and d) irrelevance of assessment for the teachers’ teaching, and learners’ learning [27]. Remesal (2011) investigated [2] Brown’s (2006) inventory among primary and secondary school teachers, taking into account the pedagogical and societal interconnection of such conceptions [28]. Azis (2015) adapted Brown’s model including examination as an important construct. While some studies (e.g. Ref. [29], Leong et al., 2018 (Hong Kong)) highlighted the important nature of AfL, other studies on this original inventory (e.g., ([11] Gebril & Brown, 2014 (Egypt) [10]; Hidri, 2016 (Tunisia) indicated that these contexts used summative assessment as the most effective tool of learning to the extent that teachers thought of assessment as irrelevant. This summative nature of assessment stands in contrast with TCoA in the American context that [30] McMillan and Nash (2000) investigated, using a survey and found that the teachers supported innovative assessment, rather than traditional assessment. A recent mixed-method study carried out in Europe ([31] Vogt et al., 2020) conducted a needs analysis survey for “Teachers’ Assessment Literacy Enhancement” in Cyprus, Germany, Greece, and Hungry about English teaching and learning and found that TCoA changed with the change of a context yet traditional approaches to assessment were used more than the innovative ones [32]. Chan and Luk (2021) claimed that TCoA are aligned with learners’ academic achievement; however, they attempted to conceptualize teachers’ focus on assessment in terms of holistic competency-based assessment, by highlighting the student-teacher relationship and assessment subjectivity and its crucial role in holistic competency-based assessment. During COVID-19 Pandemic [25], Ferretti et al. (2021) investigated Italian EFL teachers’ conceptions of assessment in a distance teaching-learning scenario and found that teachers faced uncertainties about understanding purposes and methods of assessment. Such investigations are important since they show that teachers’ varied assessment beliefs and the relevant literacy levels exert a strong impact on their teaching, as well as assessment practices. These contextual divergences, therefore, highlight that, in socio-cognitive milieus, teachers may assign different beliefs and values to assessment. We hypothesized that in our context, teachers had different views and beliefs about assessment.

Contextual language policies impose a profound influence TCoA, shaping teachers’ beliefs about assessment within two assessment environments; one marked by low-stakes assessment environment and the other shaped by high-stakes examinations ([33] Brown, Gebril & Michaelides, 2019). The instructional and assessment policy guidelines in these environments manifest the values teachers assign to assessment and impact their interactions with students, while rigorously adhering to established language policies. In low-stakes settings, teachers may rely on formative assessment and innovative practices to measure students’ learning. However, in other contexts such as Pakistan, the prevailing assessment policy is on AoL, where policymakers dictate the use of summative assessment and immediate scoring ([34] Ishaq et al., 2020). In such examination-dominated environments, language teachers conceive of assessment as producing overall grades through mid-course tests and final examinations ([35] Hoodbhoy, 2021). The educational scene in Pakistan has been marked by many controversies, including considerable failures in the English classes, high dropout rates after elementary schooling, large classes, and a centralized exam system that considers summative assessment only ([34] Ishaq et al., 2020). This backlog has resulted in shaky ground, where teachers face the dilemma of whether to carry out assessments as per their beliefs or to adhere to the language policy of the country.

The volatile nature of policy contexts can exert considerable influence on TCoA, assisting teachers in transforming their assessment beliefs ([36] Ford et al., 2018). One pertinent example of this shift is observed in Hong Kong, where despite the fact that policymakers and teachers conceive of summative language assessment as inescapable; there have taken significant efforts to integrate diagnostic and formative assessment in the examination system. This integration has influenced methods and purposes of assessing students. To empower teachers to possess quality assessment conceptions necessitates equipping them with effective skills to conduct and use valid assessments aimed at improving learning, rather than an AoL. While assessment is primarily implemented to probe the learner performance, understanding the main features of such performances necessitates exploring TCoA [37]. Vandeyar and Killen (2007) addressed the impact of TCoA on their practices and found that different TCoA can lead to different approaches in assessment. For instance, teachers who conceptualize assessment as a tool to diagnose learner performance will practically use assessment for student improvement; however, a lack of assessment literacy can lead to an over-reliance on traditional assessment methods ([38] Cagasan et al., 2020). Understanding TCoA in Pakistan can provide valuable insights to address the areas for improvement, especially teachers’ practices of assessment, and how they are formed by their underlying assessment beliefs ([39] Rasooli et al., 2023 [25]; Ferretti et al. (2021). We inferred that replacing outdated teaching conceptions with more dynamic assessment-literate conceptions can lead to significant changes in teaching practices. However, we also highlighted the fact that for this process to be smooth, the stakeholders, such as the ministry people, need to embrace this change and trust teachers to lead some educational reforms.

3 Methodology

3.1 Research design and participants

We explored the applicability of [2] Brown’s inventory of TCoA (2006) to the Pakistani context, by using the random sampling technique to administer the survey inventory via the Survey Monkey platform to teachers of English (n = 980), from which we received 735 responses with a return rate of 75 %. On the cover page, we informed the respondents that by agreeing to reply to the survey, they were providing their consent, and we had previously informed them via email about the purpose and significance of the study, and confidentiality of their responses. We obtained the prior ethical approval from the Committee of Validators1 in March 2020.

3.2 Data collection instrument

We employed Brown’s inventory of TCoA [2] (2006) in the Pakistani context. This inventory includes four factors, student accountability (students are responsible for their assessment performance), school accountability, (schools are responsible for the assessment outcomes), improvement (assessment is meant to enhance learning and teaching), and irrelevance (assessment results are perceived as inaccurate, negative, and they should be ignored). Both school accountability and student accountability have three indicators each. Improvement has four levels: a) assessment “describes students’ abilities”, b) “improves learning”, c) “improves teaching”, and d) “assessment is valid” each of which includes three indicators. Factor four, irrelevance, deals with three levels each of which includes three items, where assessment is generally held to be bad, ignored, and inaccurate (see Table 3 for the inventory factors and indicators).

3.3 Sample characteristics

Table 1 describes the respondents’ profiles, indicating that the sample had 43 % and 57 % of females and males respectively. More than 65 % of the population had an age range of 20 and more than 30 % had a range of 31–40. Most teachers (85 %) were either primary (43 %) or secondary (42 %) with a small population of vocational (3.8 %) and university teachers (11.2 %). Most of the respondents (72.2 %) had a master’s degree, while only 1 % were PhD holders. While 59.5 % of the respondents claimed that they are assessment literate, 58.4 % and 64.8 % of them claimed that they neither had any courses nor any assessment training respectively.Table 1 Respondents’ profile data (n = 735).

Table 1Data	N	%	
Gender	
Female	313	42.6	
Male	422	57.4	
Age	
20–25	231	31.4	
26–30	248	33.7	
31–35	121	16.5	
36–40	105	14.3	
41–45	26	3.5	
46–50	3	.4	
Above 50	1	.1	
Teaching levels	
Primary	316	43.0	
Secondary	309	42.0	
Vocational	28	3.8	
University	82	11.2	
Educational background	
BA	74	10.1	
MA	531	72.2	
MPHIL	123	16,7	
PhD	7	1.0	
Courses in assessment	
Yes	306	41.6	
No	429	58.4	
Training in assessment	
Yes	259	35.2	
No	476	64.8	
Assessment literacy	
Yes	437	59.5	
No	298	40.5	

3.4 Data analysis

We based our analyses of the data on the Structural Equating Model (SEM) to test our research hypotheses and the theoretical framework of the relationship between the different variables, school accountability, student accountability, improvement, and irrelevance. The SEM model hypotheses that there is a) a direct relationship between the variables, and therefore, this relationship directly influences the path from one variable to another, b) an indirect mediation hypothesis that there is an indirect relationship between variables, c) a moderation hypothesis that the relationship between two variables is affected by a third variable, d) a latent variable hypothesis that the variables cannot be directly observed but only inferred from the observed variables, and e) a goodness-of-fit hypothesis that the overall model fits the observed data which we tested through different fit indices of the SEM model ([40] Hu & Bentler, 1998 [41]; Yuan et al., 2016). Therefore, we planned to test whether [2] Brown’s inventory (2006) fits the observed data on TCoA in Pakistan (see the Findings section on these hypotheses).

We considered different levels of analysis of the SEM model, Exploratory Factor Analysis (EFA), Parallel Analysis (PA), R v2-Menu Analysis, and Confirmatory Factor Analysis (CFA), using SPSS 25.0, R menu v2, and AMOS 22.0. We implemented EFA, PA, and R-menu v2 as a precursor to explore and extract the number of factors and we used CFA to confirm the loadings of factors and indicators. For a model to have CFA good fit indices [42], Hu and Bentler (1999) suggested that Chi-Square/df should range from good (<3) to sometimes permissible (<5); the p-value should be > .05; the comparative fit index (CFI) > .95 (great); >.90 traditional; >.80 (sometimes permissible). The goodness-of-fit index (GFI) > .95; adjusted goodness-of-fit index >.80; standardized root means square residual (RMSR) < .09; root mean square error of approximation (RMSEA) < .05; (good); .05-.10 (moderate); >.10 (bad); and p of close fit (PCLOSE) > .05.

3.5 Findings

We carried out the internal consistency analysis to check the adequacy of the sample, by using the reliability analyses of the four factors. The internal consistencies of Cronbach’s Alpha values were satisfactory for school accountability, improvement, and irrelevance with values of α.611, α.727, and α.842 respectively but low for student accountability (α.422). To check the adequacy of the population sample, we used two statistical tools of EFA, descriptives (coefficients, significance level, determinant and KMO and Bartlett’s test of sphericity) and principal components of extraction, unrotated factor solution, and scree plot with an eigenvalue greater than 1. Table 2 shows that there is no issue with the sample size (n = 735) since the sampling adequacy value was estimated at .823. The test of sphericity was significant with a value of p = .001, X2 and, df of 325. Researchers (e.g. Ref. [42], Hu & Bentler, 1999) claimed that for a sample that should have no adequacy issue, its population should not be less than 400, with a sample size value greater than .5 (>.5).Table 2 KMO and Bartlett’s test (PCA) (n = 735).

Table 2Kaiser-Meyer-Olkin Measures of Sampling Adequacy	.823	
Bartlett’s Test of Sphericity	χ2: 4688 df: 325Sig.: .001	

Table 3 Communalities of factor loadings.

Table 3	Inventory items	Extraction	α.	M	SD	d	
1	“Assessment provides information on schools.	.597	.610	4.03	1.00	4.03	
2	Assessment is a quality accurate indicator.	.632	.610	3.81	1.15	3.30	
3	Assessment is a good way to evaluate a school.	.517	.610	3.93	1.05	3.74	
4	Assessment places students into categories	.507	.647	3.95	.99	3.98	
5	Assessment is assigning a grade to student work.	.448	.647	4.06	.93	4.32	
6	Assessment determines qualification standards.	.395	.647	4.14	.93	4.42	
7	Assessment determines how learned and teaching.	.529	.647	4.06	.91	4.46	
8	Assessment establishes what students have learned.	.331	.647	4.10	.93	4.39	
9	Assessment measures students thinking skillsa.	.644	.675	3.46	1.29	2.68	
10	Assessment provides feedback about performance.	.306	.595	4.11	.82	5.02	
11	Assessment feeds back to students needs.	.491	.595	4.07	.92	4.41	
12	Assessment helps students improve their learning.	.457	.647	4.24	.85	4.98	
13	Assessment is integrated with teaching practice.	.294		4.00	2.46	1.56	
14	Assessment information modifies ongoing teaching.	.427	.595	3.89	.96	4.05	
15	Assessment allows students to get instructionsa.	.602	.595	3.85	1.08	3.55	
16	Assessment results are trustworthya.	.552	.675	3.54	1.25	2.83	
17	Assessment results are consistenta.	.696	.675	3.19	1.34	2.38	
18	Assessment results can be depended ona.	.376	.595	3.80	1.05	3.61	
19	Assessment forces teachers to teach against beliefs.	.495	.835	2.96	1.25	2.35	
20	Assessment is unfair to studentsa.	.618	.835	2.60	1.31	1.99	
21	Assessment interferes with teachinga.	.534	.835	2.97	1.28	2.30	
22	Teachers make little use of the results.	.473	.835	3.27	1.23	2.65	
23	Assessment results are filed and ignoreda.	.586	.835	3.18	1.31	2.43	
24	Assessment has little impact on teachinga.	.514	.835	3.18	1.29	2.46	
25	Assessment results have measurement errors.	.603	.835	3.49	1.22	2.86	
26	Teachers should consider assessment imprecision.	.621	.835	3.56	1.17	3.03	
27	Assessment is an imprecise process”a.	.598	.675	2.93	1.39	2.10	
	Total	.512	.763	3.64	.45	3.32	
Extraction Method: Principal Component Analysis.

All items were significant at p=< .001.

a Items with low loadings were not retained in the model (Fig. 5).

To check factor loadings and estimate their number, the PCA of data communalities, Table 3, indicates factor loadings of the 27 [2] Brown’s inventory (2006) items, where loadings ranged from .696 (item 17) to .294 (item 13). Due to reliability concerns, we discarded the loadings (column 1) with values of ≤ .30 from EFA analysis. Cronbach’s Alpha of the inventory items (n = 27) was α.763 (mean = 3.64, SD = .45, d = 3.32), which was said to be high since it was estimated at a level of more than .60. For the extraction values, items of more than .30 were 6, 8, 10, and 18 with values of .395, .331, .306. and .376 respectively. Six items had values of more than .4: 5(.448), 11(.491), 12(.457), 14(.427), 19(.495) and 22(.473). Nine items had values of above .5: 1 (.597), 3 (.517), 4 (.507), 7 (.529), 16 (.552), 21 (534), 23 (.586), 24 (.514) and 27 (.598). Items 9 (.644), 15 (.602), 17 (.696), 20 (.618), 25 (.603) and 26 (.621) had values above .6. The Cronbach’s Alpha values ranged from .610 to .835. Factor 1 (α. 835), 8 indicators (19–26); factor 2 (α. 675) 4 indicators (9, 16, 17, and 27); factor 3 (α. 647), 6 indicators (4, 5, 6, 7, 8, and 12); factor 4 (α.610), 3 indicators (1, 2, and 3); factor 5 (α.595), 5 indicators (10, 11, 14, 15, and 18); factor 6 (had one item only, item 13), and, therefore, it was not possible to run the reliability analysis. The mean and SD values ranged from 2.60 (item 20) to 4.24 (item 12) and from .82 (item 10) to 1.29 (item 27) respectively. Cohen’s d, a measure of effect size, ranged from 1.56 (item 13) to 5.02 (item 10).

In analyzing the total variance of the eigenvalues, Table 4, the PCA data indicated that there are six factors whose eigenvalues were ≥1 and that these values ranged from 4.132 (factor 1) to 1.044 (factor 6). The first six factors had an eigenvalue greater than 1, with values of 4.132, 3.950, 2.143, 1.488, 1.088 and 1.044 for items 1, 2, 3, 4, 5, and 6 respectively. The scree plot in the graphical representation (Fig. 1) of the PCA and eigenvalues showed six factors to retain with factor six almost being on the borderline. To explore the exact number of factors and whether the data had statistically significant eigenvalues, we conducted both parallel analysis (PA) and Monte Carlo (MC), using [43] O’connor’s syntax (2000). Contrary to the EFA, the analysis of the eigenvalues indicated four instead of six factors. Table 5 presents the PA of the eigenvalues of the raw data, mean, and percentile for the 735 respondents (see Fig. 2).Table 4 Principal component analysis.

Table 4Component	Initial Eigenvalues	
Total	% of Variance	Cumulative %	
1	4.132	15.304	15.304	
2	3.950	14.629	29.933	
3	2.143	7.936	37.869	
4	1.488	5.513	43.381	
5	1.088	4.030	47.411	
6	1.044	3.868	51.279	
7	.967	3.582	54.861	
8	.935	3.462	58.323	
9	.856	3.171	61.495	

Fig. 1 PCA of the eigenvalues (6 factors).

Fig. 1

Fig. 2 Graphical representation of the scree plot eigenvalues (4 factors).

Fig. 2

Table 5 Principal components and random Normal data generation (PA).

Table 5Root	Raw data	Means	Prcntyle	
1.00	4.13	1.37	1.42	
2.00	3.95	1.32	1.36	
3.00	2.14	1.28	1.31	
4.00	1.49	1.24	1.27	
5.00	1.09	1.21	1.24	
6.00	1.04	1.18	1.21	
7.00	.97	1.15	1.18	
8.00	.93	1.13	1.15	
9.00	.86	1.10	1.12	

Table 5 indicates a selection data sample of nine out of the 27 inventory items, a calculation of 1000 data sets, and a percentile of 95, where column one, root, is the number of the inventory items of the TCoA inventory. Column two, raw data, analyzed the eigenvalues of PCA that were aligned with the actual data of the study. Columns 3 and 4 indicated the values of the means and percentile respectively. To get the exact number of factors that loaded on the matrix, the value of the raw data in column 2 should be greater than the mean values in column 3. Therefore, the first four factors in column 2 had raw data values of 4.13, 3.95, 2.14, and 1.49 and means of 1.37, 1.32, 1.28, and 1.24 respectively. The graphical representation of the scree plot of the Monte Carlo PCA in Fig. 1 supports the PA of Table 5 since the data showed four factors. This result is based on the curve, where it shows four factors above the elbow of the curve. The blue, green, and light brown lines indicate the raw data of the 27 questionnaire items, the mean values of all the items, and the eigenvalues of all the items respectively.

We then used the R menu V-2 analysis to confirm the PA number of factors and to investigate the exact number of factors further, as well as their loadings. Table 6 deals with Velicar’s MAP values, Velicar’s minimum average partial in Table 7, fit to comparison data, Table 8, and comparison data in Table 9, as well as Fig. 3 on the eigenvalues, PA, OC, and AC and fit to comparison data in Fig. 4. These analyses particularly addressed the number of factors in the inventory. In Table 7, the MAP output showed a distinct 4th step minimum squared average partial correlation of .022. This is already highlighted in Table 6 with the squared map of .022 (factor 4) and its fourth power map of .012. In addition, the Velicar’s minimum components to retain were 4 factors (Fig. 3). In Fig. 3, the largest number of factors in the data was fixed at 6 factors; however, the parallel analysis (PA), optimal coordinates (OC), and acceleration factor (AF) analyses displayed 4 factors. Table 8, the fit to comparison data, shows that moving from factor one to two, from two to three, and from three to four was statistically significant at .000 (p = .000). However, moving from four to five provides a value of 1.000 (p = 1.000) (column 3, p-value). Table 9 presents the comparison data of factor number estimation, and it indicates 4 factors, which is also confirmed in Fig. 4, fit to comparison data.Table 6 Velicar’s map values.

Table 6	Squared average partial correlations	4th average partial correlations	
0	.079	.032	
1	.080	.029	
2	.077	.031	
3	.103	.063	
4	.022	.012	
5	.023	.012	
6	.024	.012	
7	.025	.012	
8	.027	.012	
9	.028	.012	

Table 7 Velicer's minimum average partial test.

Table 7Velicer's Minimum Average Partial Test	
	Velicer's Minimum	
Minimum	Components to retain	
Squared MAP	.022	4	
4th power MAP	.012	4	

Table 8 Fit to comparison data.

Table 8Fit to Comparison Data	
Nb of factors	RMSR Eigenvalue	p-value	
1 factor	1.257	NA	
2 factor	.818	.000	
3 factor	.518	.000	
4 factor	.133	.000	
5 factor	.139	1.000	

Table 9 Comparison data.

Table 9Correlations	Number of factors to retain	
Spearman	4	

Fig. 3 Eigenvalue, PA, OC and AF values.

Fig. 3

Fig. 4 Fit to comparison data.

Fig. 4

Fig. 5 Fit indices: χ2: 218.438, df: 112, χ2/df: 1.950, SRMR: .052, GFI: .966, AGFI: .953, RMSEA: .036, CFI: .936, TLI: .923, PCLOSE: 1.000.

Fig. 5

Given the variability in factor numbers, we opted for CFA mainly due to its accuracy and precision in delineating loadings and factor structure. As mentioned in the data analysis section, the fit indices of CFA, Chi-square/df, CFI, GFI, AGFI, RMSR, RMSEA, and PCLOSE, were practical in evaluating the (in)admissibility of original inventory of TCoA ([2] Brown, 2006) to the Pakistani context.

CFA data analyses showed that the original inventory was inadmissible to the Pakistani context. Despite the model in Fig. 5 indicated four factors, the factor loadings significantly deviated from the original inventory values. Therefore, items that had low factor loadings (1, 2, 3, 21, 22, 23, 24, 25, 26, and 27) were subsequently removed from the analysis., The original 27 inventory items (Table 3) were reduced to 17. School accountability loaded to the original three items, “assessment provides information” (r=.57), “assessment is an accurate indicator” (r=.68), and “assessment evaluates school” (r=.52). School accountability loaded to the three original items, “assessment categorizes students” (r=.41), “assessment is about assigning a grade” (r=.49), and “assessment is about qualification standards” (r=.44). Unlike the original inventory, factor three, improvement, had 2 s-order factors, the first second-order factor, “assessment improves learning”, loaded directly onto five items, “learning and teaching” (r=.51), “students’ learning” (r=.48), “students’ performance” (r=.48) and “feedback learning” (r=.48) and “improves learning” (r=.58). The other second-order factor, “assessment improves teaching”, loaded to three items, “teaching practice” (r=.24), “ongoing teaching” (r=.54) and “assessment is trustworthy” (r=.29). Factor four, irrelevance, was directly linked to three indicators, “against beliefs” (r=.67), “measurement error” (r=.73) and “assessment makes little use of results” (r=.60). The correlations between the four factors were positive in school accountability and improvement (r = .92), school accountability and student accountability (r = .59), school accountability and improvement (r = .55); however, it was negative between school accountability and irrelevance with a value of r = −.16.

All factor loadings of the Pakistani model were significant at p = .000. The fit indices (Fig. 5), X2/df, were determined at 1.950. The CFI value is traditional since it was fixed at .93. If the TLI is estimated at almost .90 then the model is said to be good ([42] Hu & Bentler, 1999). In this study, it is estimated at .923. The GFI value is .96, and the AGFI, originally estimated at more than .80, is estimated at .95, which should originally be more than .80. For the model to fit well, an RMSEA good level should be less than .05. In this study it is fixed at .03. The SRMR should be less than .09. In this study, it is at .05. The p of CLOSE fit is at 1.000, and, as mentioned in the data analysis section, it should be more than .05. The loadings here show that the original inventory is inadmissible to the Pakistani context.

4 Discussion

In this study, we examined the extent to which TCoA were applicable in the Pakistani context and the implications of such conceptions on language policy and teacher professional development. To answer the first research question, “To what extent can teachers’ conceptions of assessment be admissible to the Pakistani context?”, we found that, for accountability, the indicators’ covariance loaded onto the factors between school accountability and student accountability, thus showing that schools and teachers could be held accountable for the nature and purpose of assessment. The acceptable fit of the model and the positive factor loadings indicated that teachers of English perceived student accountability and school accountability as reliable constructs that could profoundly shape assessment practices [2]. Brown’s (2006) inventory accountability intersects with the Pakistani teachers’ perceptions of the accountable nature of assessment. Also, teachers agreed with student accountability since they thought that assessment could help students to be placed into categories, assign them a grade, and determine if they could meet the different qualification standards. The strongly endorsed high correlation between student accountability and improvement (we retained eight out of the 12 indicators) reflected teachers’ conceptualization of formative assessment in that they agreed with most of the items that promoted formative assessments. The high positive correlation meant that teachers tended to believe that assessment makes both school and student accountable for its quality and impact and that school is accountable for the improvement nature of assessment. This formative assessment perspective is further reinforced in the improvement factor, where the analysis indicated a retention of 8 out of 12 items, thereby endorsing formative assessment further. This endorsement contradicts the status of assessment in most schools by being traditional. Teachers had high agreement percentages (agree and strongly agree) with most of the improvement indicators since they thought of assessment as a way to define the learning amount from teaching, determine what students have learned, provide feedback on student performance, and use assessment to improve learning. Teachers did not attribute much importance to the integrated nature of assessment with teaching, and this stance is also reflected in the lack of trust in the assessment results, where most of teachers claimed that the “assessment results are not trustworthy”. These endorsements meant that teachers had an awareness of the formative nature of assessment and their roles as assessors ([44] Xu & Brown, 2016) of students’ works. This self-awareness has been highlighted in research (e.g. Ref. [45], Scarino, 2013). Perhaps teachers thought of formative assessment as a way to improve learning. The improvement role of assessment in learning has also been addressed in research ([7] Lutovac & Flores, 2022).

The high agreement percentages and positive correlation values observed in the study meant stressed the importance teachers assigned to the formative nature of assessment in supporting learners to improve their assessment strategies, determine what they learned, provide feedback on their assessment performance, improve their learning, and change the teaching practices to meet the dynamic needs of students. This is a partial admissibility of the original inventory to the Pakistani context. These findings are in agreement with what [3] Brown and Remesal (2017) found in New Zealand and Spain, especially in the relationship between accountability and improvement (e.g., New Zealand [23], Brown, 2011; Hong Kong [46], Brown & Harris, 2009); Spain [3], Brown & Remesal, 2012), USA [47], Bonner & Chen, 2009)). Consistent with other studies in different contexts was the finding that TCoA generally impacted their practices of assessment and teaching (see Ref. [48] Griffiths et al., 2006). However, unlike other studies (e.g. Ref. [11], Gebril & Brown, 2014 [12]; Hidri, 2016), this survey study indicated that teachers still perceived accountability positively. Similar findings (e.g. Ref. [49], Darmody et al., 2020) were found in the literature on school accountability.

To answer the second research question, “What implications does language policy have on teachers’ conceptions of assessment and their professional development in Pakistan?”, we found that, consistent with the results from other contexts, our study indicated the impact of language policy in Pakistan on TCoA and their professional development. Despite the variation in the correlation between factors and indicators, findings of the study concurred with other studies [10] (Hidri et al., 2023), highlighting a persistent tension between teachers and decision-makers, specifically, the Ministry of Education (MoE) people. In this study, Pakistani teachers of English exhibited conflicting views about improvement and irrelevance, which was reflected in low correlations. Teachers’ skepticism emanates from the trustworthiness of results and the ways assessment was integrated with teaching. Consequently, teachers thought that the examinations did not measure what they were supposed to measure, indicating a lack of trust in the assessment results. In this educational context, the examination boards affiliated with the MoE are exclusively responsible for exam design, predominantly focusing on summative assessments, which stands in sharp contrast with the formative nature of classroom assessment. We deduced that teachers had skeptical views towards the diverse policy decisions as reflected in a moderately high correlation in irrelevance factor, with three direct indicators, “assessment forces teachers to teach in a way against their beliefs”, “assessment results should be treated cautiously given measurement error”, and “teachers conduct assessments but make little use of the results”. Two of these indicators are related to exam validity, indicating that the irrelevant part of assessment may stem from the challenges teachers face in item design and the inflexible nature of the MoE people’s stances toward implementing innovative assessment strategies such as formative assessment. Also, it appears that teachers lack expertise in item design, interpret and use test scores, make inferences about their impact on test-takers, and make informed decisions accordingly. This narrow scope of assessment on the part of teachers underscores a significant deficiency in their professional development, highlighting the urgent need for training in assessment literacy. This finding indicated that teachers could not change the assessment techniques probably, likely due to the prescriptive nature of educational practices and language policy as dictated by decision-makers. Excessive monitoring by policymakers can result in negative washback effects on teaching and assessment practices. Nevertheless, teachers can mitigate these evolving classroom-based assessment challenges by engaging with national assessment policies more actively. For this initiative to be effective, teachers need to be assessment literate, empowering them to report results and draw valid inferences about students’ abilities in examinations. However, as stated in the inventory, teachers “make little use of results” in examination-driven systems since they do not have any control over the assessment process.

The findings of the study indicated that the good fit indices of the model (see Fig. 5) were not congruent with the original inventory (see Ref. [2] Brown, 2006). This incongruity was evident since ten (10) indicators on improvement and irrelevance did not align with the model. Due of their factors’ low loadings, we removed four items from improvement associated with measuring critical thinking, providing differentiated instruction alternatives, and maintaining reliability and assessment consistency. The deleted items reflected teachers’ views of the role of diverse assessment paradigms allowing for differentiated instruction and favouring memorization over competency-based assessment of higher-order critical thinking skills. These construct-irrelevance conceptions of assessment show teachers’ limited scope of professional development in assessment literacy. It would then appear that assessments primarily measure discrete-point test items reliant on memorization. TCoA at this level are impacted and shaped by the frequency and quality of teacher professional development, especially when it is administered by the MoE people who lack expertise language assessment. Similar to the findings in other studies (e.g., ([49] Darmody et al., 2020 [12]; Hidri, 2016), this study identified low factor and indicator loadings. However, unlike other studies that showed strong endorsements between factors and their indicators (e.g. Ref. [33], Brown et al., 2019 [50]; Brown et al., 2011) with high endorsement values, this study found high correlations only between student accountability and improvement. Because of the inadmissibility of some indicators in factors 3 and 4 to the Pakistani context, teachers reflected their views on the nature and purpose of assessment results, suggesting that in that such results could not be consistent, nor depended on, as long as the MoE people oversee these assessment tasks.

Teachers hail from diverse social backgrounds, requiring the implementation of varied and appropriate methodologies to implement assessment. Despite these efforts, language assessment at the school level remains in dire need of comprehensive reforms. Unfortunately, the majority of students entering tertiary education institutions do not possess the required competence in the English language competencies particularly among students who belong to state and low-income schools, leading to a significant learning deficit. This lack of equity underscores the challenges of implementing a democratic stance to language assessment ([9] Shohamy, 2007). While [2] Brown’s inventory (2006) has proved to be admissible to different educational contexts as discussed above, our study indicated that the construct validity of the inventory requires further exploration. Although the TCoA inventory is based on solid theoretical underpinnings of four factors, school accountability, student accountability, improvement, and irrelevance, and shows significant and strong correlations between factors and their indicators, its inadmissibility to the Pakistani context underscores construct-irrelevance and inconsistencies across other contexts such as Egypt ([11] Gebril & Brown, 2014), Tunisia ([12] Hidri, 2016). This construct-irrelevance of the inventory may reflect conflicting views of assessment in particular and language learning in general.

5 Conclusion

The objectives of this study were twofold: a) to examine the assessment intricate views among the Pakistani language teachers of English and b) to explore how these conceptions relate to and are shaped by the educational policies in the country and how such policies can impact the quality of professional development. The study showed that teachers generally held positive views of assessment in terms of its role in student progress and school accountability. However, we, at the same time, raised concerns about the practicality of this inventory in Pakistan highlighting clear deficiencies in its accuracy and validity, which could be due to unique cultural and educational differences. Going back to the hypotheses of the SEM model, we concluded that there is a direct relationship in the covariance between school accountability and student accountability, thus confirming the interlink of these variables. The SEM model also hypothesizes the indirect relationship between student accountability and improvement, thus supporting the mediation hypothesis we stated earlier, as well as the latent variable hypothesis in that this relationship could not be directly observed. We inferred this latent relationship through some direct variables of how teachers perceived formative assessment and student accountability, which meant that teachers’ conceptions of accountability impacted their beliefs about the improvement aspect of assessment. Considering the goodness-fit hypothesis of the overall model, we confirmed that [2] Brown’s (2006) inventory is not admissible to the Pakistani context. Also, the moderation hypothesis was supported by the observed low correlation between improvement and irrelevance and the conflicting views on the trustworthy aspect of the assessment results. Specifically, teachers’ perceptions of certain assessment practices may finally moderate their beliefs of the role assessment plays in the improvement aspect of language assessment in Pakistan. Finally, we concluded that the inadmissibility of specific factors and indicators to the Pakistani context calls into question the construct validity of the inventory, necessitating a further investigation of the inventory in Pakistani educational institutions and beyond. While the analyses showed strong correlations of specific factors and loadings in TCoA, we think that there are potential limitations in the universal implementation of [2] Brown’s (2006) inventory. These limitations call for implementing qualitative approaches to probe TCoA.

Incorporating interviews with senior teachers responsible for professional development, along with the current survey inventory, would have elicited reliable qualitative datasets, resulting in more insights into TCoA. Engaging with decision-makers through interviews could have also enlightened us to have a holistic understanding of the assessment landscape in Pakistan, and therefore, explore why language policies do not support teachers in being assessment literate. Given the current limitations, it would be imprudent to claim the generalisability of our findings across the broader Pakistani context, suggesting the need for future research to revisit TCoA in Pakistan.

The study elucidated three major implications that significantly enhance the status of assessment in Pakistan. Firstly, the political implications revealed a profound disconnect between TCoA and the language policies of the country, necessitating a policy reform to empower teachers to achieve professional growth and be assessment literate. The analysis of TCoA is beneficial as it supports stakeholders to critically evaluate the used assessment, bolster the quality of teaching and learning, and enhance the pace of professional development ([2] Brown, 2006) ([26] Brown, 2004 [7]; Lutovac & Flores, 2022 [51]; Stiggins, 2010 [52]; Xu & He, 2019) in refining the academic programs. Secondly, from a pedagogical perspective, since the assessment strategies are shaped by the relevant teaching methodologies and practices ([13] Barnes et al., 2017), teachers should move away from exam-centric curricula and traditional pedagogical approaches. Instead, embracing new teaching methodologies and competency-based assessment approaches can equip students with practical life skills, transcending rote memorization of isolated bits of the English language. Lastly, methodologically, the study used a multi-faceted approach to investigate the status of TCoA in Pakistan, by integrating EFA, R-Menu Analysis, and CFA. By using similar meticulous data analyses, we planned to offer future researchers to opportunity to check the replicability of the study to this context and other related contexts.

Educational institutions in Pakistan must prioritize intensive professional development training in assessment, particularly in English. As the role of the English language in Pakistan grows ([53] Shamim, 2011), the nature and conceptions and practices of assessment are changing rapidly, thus calling for substantial reforms in how second language assessment, teaching, and learning are perceived and used in Pakistan. Teachers should be assessment-literate to be able to assess learners objectively ([17] Warsi, 2004). Similarly, policymakers should embrace flexibility and openness to change, empower teachers to spreadhead assessment reforms, trust their insights on test reports and the uses of test scores. The current situation requires profound educational reforms in the teaching and assessment of English, by upskilling teachers’ beliefs and practices of assessment. TCoA are inextricably linked to their socio-cognitive contexts, carrying out assessment will always remain the linchpin in any constructive alignment of the course and program learning outcomes. Thus, teachers should not be marginalized, and they should take the lead in any educational reform despite the ongoing tension engendered by the exiting educational policies in Pakistan.

Data availability

Data privately shared on a safe archived drive and will be available from the corresponding author on reasonable request.

on request.

Funding

This study received no funding.

CRediT authorship contribution statement

Sahbi Hidri: Writing – review & editing, Writing – original draft, Visualization, Validation, Supervision, Software, Resources, Project administration, Methodology, Investigation, Formal analysis, Data curation, Conceptualization. Musharraf Aziz: Writing – review & editing, Writing – original draft, Visualization, Investigation, Data curation. Manal Qutub: Writing – review & editing, Validation, Supervision, Methodology, Formal analysis, Data curation, Conceptualization.

Declaration of competing interest

The authors declare that they have no competing interests.

Acknowledgements

The authors are grateful to the participants.

1 For the names of the validators, please contact the authors.
==== Refs
References

1 Baidoo-Anu D. Rasooli A. DeLuca C. Cheng L. Conceptions of classroom assessment and approaches to grading: teachers' and students' perspectives Education Inquiry 2023 10.1080/20004508.2023.2244136
2 Brown G.T.L. Teachers' conceptions of assessment: validation of an abridged version Psychol. Rep. 99 1 2006 166 170 17037462
3 Brown G.T. Remesal A. Teachers' conceptions of assessment: comparing two inventories with Ecuadorian teachers Stud. Educ. Eval. 55 2017 68 74
4 Brown G.T.L. Chaudhry H. Dhamija R. The impact of an assessment policy upon teachers' self-reported assessment beliefs and practices: a quasi-experimental study of Indian teachers in private schools Int. J. Educ. Res. 71 2015 50 64
5 DeLuca C. Coombs A. LaPointe-McEwan D. Assessment mindset: exploring the relationship between teacher mindset and approaches to classroom assessment Stud. Educ. Eval. 61 2019 159 169 10.1016/j.stueduc.2019.03.012
6 Deneen C.C. Brown G.T.L. The impact of conceptions of assessment on assessment literacy in a teacher education program Cogent Education 3 1 2016 1225380
7 Lutovac S. Flores M.A. Conceptions of assessment in pre-service teachers' narratives of students' failure Camb. J. Educ. 52 1 2022 55 71 10.1080/0305764x.2021.1935736
8 Opre D. Teachers' conceptions of assessment Procedia-Social and Behavioral Sciences 209 2015 229 233 10.1016/j.sbspro.2015.11.222
9 Shohamy E. The power of language tests, the power of the English language and the role of ELT Cummins J. Davison C. International Handbook of English Language Teaching 2007 Springer New York, NY 521 531
10 Hidri S. Coombe C. Aljahromi D. Rethinking language assessment literacy in the MENA region: critical perspectives The Journal of Asia TEFL 20 4 2023 935 941 10.18823/asiatefl.2023.20.4.13.93
11 Gebril A. Brown G.T. The effect of high-stakes examination systems on teacher beliefs: Egyptian teachers' conceptions of assessment Assess Educ. Princ. Pol. Pract. 21 1 2014 16 33
12 Hidri S. Conceptions of assessment: investigating what assessment means to secondary and university teachers Arab Journal of Applied Linguistics 1 1 2016 19 43
13 Barnes N. Fives H. Dacey C.M. U.S. teachers' conceptions of the purposes of assessment Teach. Teach. Educ. 65 2017 107 116
14 Shohamy E. The Power of Tests: A Critical Perspective on the Uses of Language Tests 2001 Pearson Education Harlow
15 Takele M. Melese W. Primary school teachers' conceptions and practices of assessment and their relationships Cogent Education 9 1 2022 2090185 10.1080/2331186X.2022.2090185
16 Mahboob A. No English, no future!. Language policy in Pakistan Obeng S. Hartford B. Political Independence with Linguistic Servitude: the Politics about Languages in the Developing World 2002 15 39
17 Warsi J. Conditions under which English is taught in Pakistan: an applied linguistic perspective Sarid Journal 1 1 2004 1 9
18 Alderman H. Orazem P.F. Paterno E.M. School quality, school cost, and the public/private school choices of low-income households in Pakistan J. Hum. Resour. 2001 304 326
19 Coleman H. Capstick A. Language in Education in Pakistan: Recommendations for Policy and Practice 2012 British Council Islamabad
20 Coleman H. Teaching and Learning in Pakistan: the Role of Language in Education 2010 The British Council Islamabad 1 56
21 Horgan K. Gardiner-Hyland F. Irish student teachers' beliefs about self, learning and teaching: a longitudinal study Eur. J. Teach. Educ. 42 2 2019 151 174
22 Fletcher R.B. Meyer L.H. Anderson H. Johnston P. Rees M. Faculty and students' conceptions of assessment in higher education High Educ. 64 1 2012 119 133 10.1007/s10734-011-9484-1
23 Brown G.T.L. Teachers' conceptions of assessment: comparing primary and secondary teachers in New Zealand Assessment Matters 3 2011 45 70 10.18296/am.0097(DONETILLHERE
24 Ahmad A. Saeed M. Development and validation of instrumentation to assess university academics' research and teaching performance in Punjab, Pakistan Journal of Institutional Research Southeast Asia 19 2 2021
25 Ferretti F. Santi G.R.P. Del Zozzo A. Garzetti M. Bolondi G. Assessment practices and beliefs: teachers' perspectives on assessment during long distance learning Educ. Sci. 11 6 2021 264
26 Brown G.T. Teachers' conceptions of assessment: implications for policy and professional development Assess Educ. Princ. Pol. Pract. 11 3 2004 301 318
27 Remesal A. Primary and secondary teachers' conceptions of assessment: a qualitative study Teach. Teach. Educ. 27 2 2011 472 482
28 Azis A. Conceptions and practices of assessment: a case of teachers representing improvement conception Teflin Journal 26 2 2015 129 154 10.15639/teflinjournal.v26i2/129-154
29 Leong W.S. Ismail H. Costa J.S. Tan H.B. Assessment for learning research in East Asian countries Stud. Educ. Eval. 59 2018 270 277 10.1016/j.stueduc.2018.09.005
30 McMillan J.H. Nash S. Teachers' classroom assessment and grading decision making Paper Presented at the Annual Meeting of the National Council of Measurement in Education 2000 New Orleans 10.1111/j.1745-3992.2001.tb00055.x
31 Vogt K. Tsagari D. Csépes I. Green A. Sifakis N. Linking learners' perspectives on language assessment practices to teachers' assessment literacy enhancement (TALE): insights from four European countries Lang. Assess. Q. 17 4 2020 410 433 10.1080/15434303.2020.1776714
32 Chan C.K. Luk L.Y. A four-dimensional framework for teacher assessment literacy in holistic competencies Assess Eval. High Educ. 2021 1 15 10.1080/02602938.2021.1962806
33 Brown G.T. Gebril A. Michaelides M.P. Teachers' conceptions of assessment: a global phenomenon or a global localism Frontiers in Education 4 2019 16 10.3389/feduc.2019.00016 Frontiers Media SA
34 Ishaq K. Rana A.M.K. Zin N.A.M. Exploring summative assessment and effects: primary to higher education Bull. Educ. Res. 42 3 2020 23 50
35 Hoodbhoy P. Pakistan's higher education system 38 Handbook of Education Systems in South Asia vol. 977 2021 10.1007/978-981-15-0032-9_64
36 Ford T.G. Urick A. Wilson A.S. Exploring the effect of supportive teacher evaluation experiences on US teachers' job satisfaction Educ. Pol. Anal. Arch. 26 2018 10.14507/epaa.26.3559 59-59
37 Vandeyar S. Killen R. Educators' conceptions and practice of classroom assessments in post-apartheid South Africa S. Afr. J. Educ. 27 1 2007 101 115
38 Cagasan L. Care E. Robertson P. Luo R. Developing a formative assessment protocol to examine formative assessment practices in the Philippines Educ. Assess. 25 4 2020 259 275
39 Rasooli A. DeLuca C. Cheng L. Beginning teacher candidates' approaches to grading and assessment conceptions: implications for teacher education in assessment Educ. Res. Pol. Pract. 22 1 2023 63 90 https://eric.ed.gov/?id=EJ1363076
40 Hu L.T. Bentler P.M. Fit indices in covariance structure modeling: sensitivity to underparameterized model misspecification Psychol. Methods 3 4 1998 424 453 10.1037/1082-989X.3.4.424
41 Yuan K.-H. Bentler P.M. Chan W. Structural equation modeling with many variables: a systematic review of issues and developments Front. Psychol. 7 2016 1761 10.3389/fpsyg.2016.01761 27895604
42 Hu L. Bentler P.M. Cut-off criteria for fit indices in covariance structure criteria versus new alternatives Structural Equation Modelling: A Multidiscip. J. 6 1999 1 55
43 O’connor B.P. SPSS and SAS programs for determining the number of components using parallel analysis and Velicer's MAP test Behav. Res. Methods Instrum. Comput. 32 3 2000 396 402 11029811
44 Xu Y. Brown G.T.L. Teacher assessment literacy in practice: a reconceptualization Teach. Teach. Educ. 58 2016 149 162 10.1016/j.tate.2016.05.010
45 Scarino A. Language assessment literacy as self-awareness: understanding the role of interpretation in assessment and in teacher learning Lang. Test. 30 3 2013 309 327 10.1177/0265532213480128
46 Brown G.T.L. Harris L. Unintended consequences of using tests to improve learning: how improvement-oriented resources Journal of Multi-Disciplinary Evaluation 6 12 2009 68 91
47 Bonner S. Chen P. Teacher candidates' perceptions about grading and constructivist teaching Educ. Assess. 14 2 2009 57 77 10.1080/10627190903039411
48 Griffiths T. Gore J. Ladwig J. Teachers' fundamental beliefs, commitment to reform, and the quality of pedagogy Paper Presented at the Australian Association for Research in Education (AARE) 2006
49 Darmody M. Lysaght Z. O’Leary M. Irish post-primary teachers' conceptions of assessment at a time of curriculum and assessment reform Assess Educ. Princ. Pol. Pract. 27 5 2020 501 521 10.1080/0969594X.2020.1761290
50 Brown G.T.L. Michaelides M.P. Ecological rationality in teachers' conceptions of assessment across samples from Cyprus and New Zealand Eur. J. Psychol. Educ. 26 3 2011 319 337
51 Stiggins R.J. Essential formative assessment competencies for teachers and school leaders Andrade H.L. Cizek G.J. Handbook of Formative Assessment 2010 Taylor & Francis New York, NY 233 250
52 Xu Y. He L. How pre-service teachers' conceptions of assessment change over practicum: implications for teacher assessment literacy Frontiers in Education 4 2019 145
53 Shamim F. English as the language for development in Pakistan: issues, challenges and possible solutions Dreams and realities: Developing countries and the English language 2011 291 310
