Methods for Addressing Preferential Sampling in Semiempirical Liquefaction Modeling
Jonathan Schmidt, Shideh Dashti, Cristina Torres-MachiAbstract
Semiempirical probabilistic liquefaction models (ELMs) are widely used in geotechnical earthquake engineering for evaluating the potential for seismic soil liquefaction manifestation and its consequences. Estimating ELM parameters by statistical approaches relies heavily on curated case history data sets. Because of limited time and resources and unavoidable constraints, these case histories are often preferentially sampled—that is, the locations at which field data are collected are directly or indirectly linked to the observed liquefaction outcome. A key question for liquefaction modelers is what effect, if any, the different potential mechanisms of preferential data collection have on parameter estimates (and by extension, predicted probabilities and risks). Although past work has considered this question, many modelers adopted default approaches without critical examination of their assumptions with respect to the data. This paper provides a general Monte Carlo simulation procedure to illustrate how preferential data collection can affect parameter inference in ELM development, along with theoretical background. We show the consistency of parameter estimates under three adjustment mechanisms: no adjustment, the standard weighted likelihood approach, and a variation of weighted likelihood in which the weights are chosen to balance the data set. We also provide an illustration of how potentially incorrect parameter estimates can be translated into practical impacts. Based on the simulation studies and theoretical considerations, we show that the appropriate adjustment for a given data set always depends on the actual mechanism of data collection and that assumptions based on past recommendations may not always work (or even produce a conservative result). Therefore, we do not generally recommend algorithmic application of past assumed case weights or balancing the data set. Instead, we recommend specific approaches based on whether the data collection is known to be ignorable or nonignorable. For cases in which the data collection is unknown, we provide a proposed modification to the existing weighted likelihood approach in which the appropriate case weights (inverse inclusion probabilities) are estimated based on an auxiliary data set.