Fixed or Random Effects: Analysis of Clustered Data
Kevin He, Xiangeng Fang, Yubo Shao, Nicholas Hartman, John D. KalbfleischABSTRACT
In analyzing clustered data, random effects (RE) and fixed effects (FE) models are two primary approaches. The RE model assumes that the cluster‐specific effects are random and uncorrelated with individual characteristics. If this assumption is violated, the RE approach can lead to biased estimates of the coefficients of individual covariates and cluster effects. In contrast, the FE approach, which models cluster effects as fixed effects, yields unbiased estimates even if the cluster effects are correlated with individual covariates. Numerous studies have compared the FE and RE approaches. However, the practical importance of the biases observed in RE‐based methods remains unclear. Some studies have demonstrated significant biases in RE models when the assumption of no correlation is violated, whereas others have found these biases to be negligible. These differences may be due to variations in study design, with some evaluations focusing on settings that may not reflect realistic data scenarios. In this paper, we compare FE and RE models through theoretical evaluation, simulation studies, and real‐data applications. Specifically, we derive an approximate formula to quantify the asymptotic bias in the RE model and verify its accuracy through simulations. These results help to explain the different conclusions reached in the existing literature. Our goal is not to propose a new estimator or a fundamentally new inferential framework, but to provide an analytical clarification and practical quantification of the bias incurred by the standard RE estimator under a specific, interpretable dependence model, and to clarify how this bias relates to fixed‐effects and correlated random‐effects/between‐within specifications.