Risk-of-bias assessment in prognostic cohort studies: a methodological comparison study of structured appraisal tools in acute pulmonary embolism research
Ludovica Anna Cimini, Frederikus A Klok, Marc Carrier, Kerstin de Wit, Scott C Woller, Andrea Galeazzo Rigutini, Roupen Odabashian, Rosa Talerico, Maria Giovanna Ranalli, Cecilia BecattiniObjectives
To compare the reliability of four commonly used structured tools for risk-of-bias assessment in prognostic cohort studies and evaluate their agreement with expert appraisal.
Design
Methodological comparison study.
Setting
Secondary analysis of published prognostic cohort studies in patients with acute pulmonary embolism from a previously conducted systematic review.
Participants
Sixty-three cohort studies assessing the prognostic role of echocardiography in patients with acute pulmonary embolism.
Primary and secondary outcome measures
Four independent reviewers assessed study quality/risk of bias using the Newcastle-Ottawa Scale (NOS), Risk Of Bias in Non-Randomised Studies of Interventions (ROBINS-I), Quality In Prognostic Studies (QUIPS) and Quality Assessment of Diagnostic Accuracy Studies (QUADAS-2). Concomitantly and independently, expert reviewers’ implicit global judgements were used as an external comparator. Inter-rater reliability, agreement with expert appraisal, internal consistency and floor/ceiling effects were assessed.
Results
Inter-rater agreement was fair for NOS (Gwet’s agreement coefficient 1, 0.25, 95% CI 0.15 to 0.34), ROBINS-I (0.38, 95% CI 0.28 to 0.47) and QUADAS-2 (0.35, 95% CI 0.27 to 0.43) and almost perfect for QUIPS (0.85, 95% CI 0.73 to 0.93). Agreement with expert appraisal was limited for all tools except ROBINS-I, which showed fair concordance. QUIPS frequently classified studies as low risk of bias, suggesting potential overestimation of study quality. Internal consistency was generally low across tools, while ceiling effects were observed for NOS and QUIPS. ROBINS-I showed the most balanced distribution of ratings.
Conclusions
Structured tools for risk-of-bias assessment in prognostic cohort studies have variable reliability and limited agreement with expert appraisal. ROBINS-I showed the strongest concordance with expert judgement, whereas QUIPS may provide optimistic ratings. These findings support further refinement and standardisation of risk-of-bias assessment methods for observational prognostic research.