Statistical Methods for Population Pharmacokinetic Model Evaluation, Individualized Prediction, and Automated Model Development
Population pharmacokinetic (popPK) models describe drug concentrations over time while accounting for between- and within-individual variability through a hierarchical modeling framework. There are many components of a popPK model that need to be specified, including the structural and residual error models, covariate relationships, and random effects structure. Appropriately specifying these components is critical for accurate prediction and applications such as therapeutic drug monitoring (TDM). High-dose methotrexate (HDMTX) is a chemotherapy drug for which TDM is routinely used to identify delayed drug clearance and guide clinical management due to the risk of adverse events associated with high concentrations.
The first aim of this dissertation was to develop a popPK model for HDMTX in an adult patient population at ¹ú²ú´«Ã½ Medical Center and compare its performance with externally developed models. Model performance varied across the therapy monitoring period, with our model providing the most accurate median percent prediction error and median absolute percent prediction error at the first measured concentration, when concentrations are typically high. In contrast, a model by Hui et al. demonstrated the most accurate median individual percent prediction error from 36 hours onward, supporting the use of complementary models across different phases of HDMTX TDM.
The second aim was to develop a freely available TDM tool using an R Shiny application to support clinical monitoring after HDMTX administration. Based on the external validation results, the tool incorporated our newly developed model alongside models by Taylor et al. and Hui et al.
The final aim was to evaluate the ability of the automated popPK model selection tool pyDarwin to recover known data-generating structures under varied search settings. Different fitness functions and machine-learning-based algorithms were compared, with the Akaike Information Criterion and Bayesian Information Criterion demonstrating the greatest robustness to changes in model complexity. In the presence of a modest number of covariates, where exhaustive search becomes challenging, the random forest algorithm provided a more effective alternative for exploring the covariate space.
Together, this work addresses multiple stages of the popPK modeling process, spanning model development, clinical implementation, and automated model selection.