DOI: 10.1002/rev3.70211 ISSN: 2049-6613

Assessing the trustworthiness of experimental research in education—A reliability study with an adapted version of the EEF padlock tool

Hamish Chalmers, Jo‐Anne Baird, Michelle Meadows, Jonathan Kay, Nuo (Richard) Chen, Sarah Miller

Abstract

Systematic reviews have become a widely used methodology in education. Understanding the extent to which rigorous, trustworthy research is included in the reviews is crucial to the dependability of conclusions drawn from this method. Here, the reliability of judgements of research trustworthiness was investigated using an adapted version of the Education Endowment Foundation's padlock security rating tool as a rubric; this being a tool designed specifically for appraising trustworthiness in educational research. Fifty research manuscripts using experimental or quasi‐experimental designs were selected for review. Ten raters attended a training course on the application of the tool and independently rated the manuscripts. Moderate inter‐rater reliability was found (ICC = 0.54) for the final ratings, meaning that a single rater would not produce dependable ratings. Generalisability Theory analyses showed that approximately half of the variance in ratings was due to the manuscripts and over a third was associated with the manuscript by rater interaction. A low proportion of the variance was due to systematic rater effects (0.08). Estimated dependability coefficients indicated that including four raters in a study would likely reach generalisability coefficient of 0.82. Score resolution methods would improve the reliability but were not part of this research since the aim was to investigate independent ratings. Many manuscripts received low trustworthiness ratings due to incomplete reporting. This study establishes the need for further methodological work to improve judgements of trustworthiness of research to underpin systematic review methodology.

Context and implications

Rationale for this study Systematic reviews rely on trustworthiness appraisal tools, yet it is unclear whether these tools produce consistent judgements across raters in education research.

Why the new findings matter: Reliability was only moderate, indicating that single‐rater judgements may not be dependable and may undermine conclusions in systematic reviews.

Implications for systematic reviewers, authors of primary research, and journal editors: Systematic reviews should use multiple independent raters to improve reliability of trustworthiness appraisal. The authors of primary studies should adopt clearer, more complete reporting to reduce judgemental uncertainty. Journal editors should require duplicate trustworthiness appraisal in any systematic review they consider for publication and promote the use of reporting standards in reports of primary research, as both directly affect the credibility of evidence syntheses.

More from our Archive