Survival is poor for patients with metastatic cancer, and it is vital to examine new biomarkers that can improve patient prognostication and identify those who would benefit from more aggressive therapy. In metastatic prostate cancer, 2 new assays have become available: one that quantifies the number of cancer cells circulating in the peripheral blood, and the other a marker of the aggressiveness of the disease. It is critical to determine the magnitude of the effect of these biomarkers on the discrimination of a model-based risk score. To do so, most analysts frequently consider the discrimination of 2 separate survival models: one that includes both the new and standard factors and a second that includes the standard factors alone. However, this analysis is ultimately incorrect for many of the scale-transformation models ubiquitous in survival, as the reduced model is misspecified if the full model is specified correctly. To circumvent this issue, we developed a projection-based approach to estimate the impact of the 2 prostate cancer biomarkers. The results indicate that the new biomarkers can influence model discrimination and justify their inclusion in the risk model; however, the hunt remains for an applicable model to risk-stratify patients with metastatic prostate cancer.
Cancer remains the second most prevalent cause of death in the United States, claiming 605,213 lives in 2021, surpassing COVID-19 deaths. The cancer mortality rate continued to decline between 2019 and 2020, dropping by 1.5%, marking a significant 33% decrease since 1991. This ongoing improvement primarily mirrors advances in treatment, allowing patients to achieve clinical remission and recovery. Now, a cancer patient is simultaneously exposed to the risk of primary cancer as well as other risks, such as other cancer(s) or other diseases, leading to a competing risks scenario. Analysis of survival data under competing risks and the presence of cured patients have been extensively studied individually, but there is limited work in the current literature that models the possibility of cure from one risk in the presence of competing risks. Moreover, such a model should allow for the possibility of cure from the cause-specific risk of the primary cancer; however, the overall survival probability should eventually approach zero, thereby incorporating the prevalent belief of eventual failure with certainty. We propose a novel unified competing risks cure model, based on the cause-specific hazard approach, that satisfies the aforementioned desired properties. The conditions required to establish model identifiability are studied in detail. To find the maximum likelihood estimates of the model parameters, a computationally efficient expectation maximization algorithm is developed. An extensive simulation study is carried out to demonstrate the performance of the proposed model and estimation method under different parameter settings and in the presence of multiple competing risks. Finally, an application is illustrated using breast cancer data from the SEER cancer database.
The use of digital devices to collect data in mobile health studies introduces a novel application of time series methods, with the constraint of potential data missing at random or missing not at random (MNAR). In time-series analysis, testing for stationarity is an important preliminary step to inform appropriate subsequent analyses. The Dickey-Fuller test evaluates the null hypothesis of unit root non-stationarity, under no missing data. Beyond recommendations under data missing completely at random for complete case analysis or last observation carry forward imputation, researchers have not extended unit root non-stationarity testing to more complex missing data mechanisms. Multiple imputation with chained equations, Kalman smoothing imputation, and linear interpolation have also been used for time-series data, however such methods impose constraints on the autocorrelation structure and impact unit root testing. We propose maximum likelihood estimation and multiple imputation using state space model approaches to adapt the augmented Dickey-Fuller test to a context with missing data. We further develop sensitivity analyses to examine the impact of MNAR data. We evaluate the performance of existing and proposed methods across missing mechanisms in extensive simulations and in their application to a multi-year smartphone study of bipolar patients.
In many contexts, particularly when study subjects are adolescents, peer effects can invalidate typical statistical requirements in the data. For instance, it is plausible that a student's academic performance is influenced both by their own mother's educational level as well as that of their peers. Since the underlying social network is measured, the Add Health study provides a unique opportunity to examine the impact of maternal college education on adolescent school performance, both direct and indirect. However, causal inference on populations embedded in social networks poses technical challenges, since the typical no interference assumption no longer holds. While inverse probability-of-treatment weighted (IPW) estimators have been developed for this setting, they are often highly unstable. Motivated by the question of maternal education, we propose doubly robust (DR) estimators combining models for treatment and outcome that are consistent and asymptotically normal if either model is correctly specified. We present empirical results that illustrate the DR property and the efficiency gain of DR over IPW estimators even when the treatment model is misspecified. Contrary to previous studies, our robust analysis does not provide evidence of an indirect effect of maternal education on academic performance within adolescents' social circles in Add Health.
Evaluating the reproducibility or agreement of microbiome measurements is often a crucial step to ensure rigorous downstream analyses in microbiome studies. In this paper, we address this need by developing adaptations of Lin's concordance correlation coefficient (CCC) tailored to microbiome studies. We introduce a general formulation of the new CCC measures upon the use of a distance function appropriately characterizing the discrepancy between microbiome compositional measurements. We thoroughly study the special cases that adopt Euclidean distance and Aitchison distance. Our proposals appropriately account for the unique features of microbiome compositional data, including high-dimensionality, dependency among individual relative abundances, and the presence of many zeros. We further investigate a practical compound approach to help better understand the sources of data inconsistency. Extensive simulation studies are conducted to evaluate the utility of the proposed methods in realistic scenarios. We also apply the proposed methods to a microbiome validation dataset from the Feeding Infants Right.. from the STart (FIRST) study. Our analyses offer useful insight about the extent of data variations resulted from two different experiment procedures as well as their heterogeneous patterns across genera.
In precision medicine, there is much interest in estimating the expected-to-benefit (EB) subset, i.e. the subset of patients who are expected to benefit from a new treatment based on a collection of baseline characteristics. There are many statistical methods for estimating the EB subset, most of which produce a 'point estimate' without a confidence statement to address uncertainty. Confidence intervals for the EB subset have been defined only recently, and their construction is a new area for methodological research. This article proposes a pseudo-response approach to EB subset estimation and confidence interval construction. Compared to existing methods, the pseudo-response approach allows us to focus on modelling a conditional treatment effect function (as opposed to the conditional mean outcome given treatment and baseline covariates) and is able to incorporate information from baseline covariates that are not involved in defining the EB subset. Simulation results show that incorporating such covariates can improve estimation efficiency and reduce the size of the confidence interval for the EB subset. The methodology is applied to a randomized clinical trial comparing two drugs for treating HIV infection.
In medical studies, some therapeutic decisions could lead to dependent censoring for the survival outcome of interest. This is exemplified by a study of paediatric acute liver failure, where death was subject to dependent censoring due to liver transplantation. Existing methods for assessing the predictive performance of biomarkers often pose the independent censoring assumption and are thus not applicable. In this work, we propose to tackle the dependence between the failure event and dependent censoring event using auxiliary information in multiple longitudinal risk factors. We propose estimators of sensitivity, specificity and area under curve, to discern the predictive power of biomarkers for the failure event by removing the disturbance of dependent censoring. Point estimation and inferential procedures were developed by adopting the joint modelling framework. The proposed methods performed satisfactorily in extensive simulation studies. We applied them to examine the predictive value of various biomarkers and risk scores for mortality in the motivating example.
The advent of high-resolution imaging has made data on surface shape widespread. Methods for the analysis of shape based on landmarks are well established but high-resolution data require a functional approach. The starting point is a systematic and consistent description of each surface shape and a method for creating this is described. Three innovative forms of analysis are then introduced. The first uses surface integration to address issues of registration, principal component analysis and the measurement of asymmetry, all in functional form. Computational issues are handled through discrete approximations to integrals, based in this case on appropriate surface area weighted sums. The second innovation is to focus on sub-spaces where interesting behaviour such as group differences are exhibited, rather than on individual principal components. The third innovation concerns the comparison of individual shapes with a relevant control set, where the concept of a normal range is extended to the highly multivariate setting of surface shape. This has particularly strong applications to medical contexts where the assessment of individual patients is very important. All of these ideas are developed and illustrated in the important context of human facial shape, with a strong emphasis on the effective visual communication of effects of interest.
Length of stay (LOS) is an essential metric for the quality of hospital care. Published works on LOS analysis have primarily focused on skewed LOS distributions and the influences of patient diagnostic characteristics. Few authors have considered the events that terminate a hospital stay: Both successful discharge and death could end a hospital stay but with completely different implications. Modelling the time to the first occurrence of discharge or death obscures the true nature of LOS. In this research, we propose a structure that simultaneously models the probabilities of discharge and death. The model has a flexible formulation that accounts for both additive and multiplicative effects of factors influencing the occurrence of death and discharge. We present asymptotic properties of the parameter estimates so that valid inference can be performed for the parametric as well as nonparametric model components. Simulation studies confirmed the good finite-sample performance of the proposed method. As the research is motivated by practical issues encountered in LOS analysis, we analysed data from two real clinical studies to showcase the general applicability of the proposed model.
Causal mediation analysis aims to characterize an exposure's effect on an outcome and quantify the indirect effect that acts through a given mediator or a group of mediators of interest. With the increasing availability of measurements on a large number of potential mediators, like the epigenome or the microbiome, new statistical methods are needed to simultaneously accommodate high-dimensional mediators while directly target penalization of the natural indirect effect (NIE) for active mediator identification. Here, we develop two novel prior models for identification of active mediators in high-dimensional mediation analysis through penalizing NIEs in a Bayesian paradigm. Both methods specify a joint prior distribution on the exposure-mediator effect and mediator-outcome effect with either (a) a four-component Gaussian mixture prior or (b) a product threshold Gaussian prior. By jointly modeling the two parameters that contribute to the NIE, the proposed methods enable penalization on their product in a targeted way. Resultant inference can take into account the four-component composite structure underlying the NIE. We show through simulations that the proposed methods improve both selection and estimation accuracy compared to other competing methods. We applied our methods for an in-depth analysis of two ongoing epidemiologic studies: the Multi-Ethnic Study of Atherosclerosis (MESA) and the LIFECODES birth cohort. The identified active mediators in both studies reveal important biological pathways for understanding disease mechanisms.
Recurrent events such as hospitalisations are outcomes that can be used to monitor dialysis facilities' quality of care. However, current methods are not adequate to analyse data from many facilities with multiple hospitalisations, especially when adjustments are needed for multiple time scales. It is also controversial whether direct or indirect standardisation should be used in comparing facilities. This study is motivated by the need of the Centers for Medicare and Medicaid Services to evaluate US dialysis facilities using Medicare claims, which involve almost 8,000 facilities and over 500,000 dialysis patients. This scope is challenging for current statistical software's computational power. We propose a method that has a flexible baseline rate function and is computationally efficient. Additionally, the proposed method shares advantages of both indirect and direct standardisation. The method is evaluated under a range of simulation settings and demonstrates substantially improved computational efficiency over the existing R package survival. Finally, we illustrate the method with an important application to monitoring dialysis facilities in the U.S., while making time-dependent adjustments for the effects of COVID-19.
To draw real-world evidence about the comparative effectiveness of multiple time-varying treatments on patient survival, we develop a joint marginal structural survival model and a novel weighting strategy to account for time-varying confounding and censoring. Our methods formulate complex longitudinal treatments with multiple start/stop switches as the recurrent events with discontinuous intervals of treatment eligibility. We derive the weights in continuous time to handle a complex longitudinal data set without the need to discretise or artificially align the measurement times. We further use machine learning models designed for censored survival data with time-varying covariates and the kernel function estimator of the baseline intensity to efficiently estimate the continuous-time weights. Our simulations demonstrate that the proposed methods provide better bias reduction and nominal coverage probability when analysing observational longitudinal survival data with irregularly spaced time intervals, compared to conventional methods that require aligned measurement time points. We apply the proposed methods to a large-scale COVID-19 data set to estimate the causal effects of several COVID-19 treatments on the composite of in-hospital mortality and intensive care unit (ICU) admission relative to findings from randomised trials.
Screening is a powerful tool for infection control, allowing for infectious individuals, whether they be symptomatic or asymptomatic, to be identified and isolated. The resource burden of regular and comprehensive screening can often be prohibitive, however. One such measure to address this is pooled testing, whereby groups of individuals are each given a composite test; should a group receive a positive diagnostic test result, those comprising the group are then tested individually. Infectious disease is spread through a transmission network, and this paper shows how assigning individuals to pools based on this underlying network can improve the efficiency of the pooled testing strategy, thereby reducing the resource burden. We designed a simulated annealing algorithm to improve the pooled testing efficiency as measured by the ratio of the expected number of correct classifications to the expected number of tests performed. We then evaluated our approach using an agent-based model designed to simulate the spread of SARS-CoV-2 in a school setting. Our results suggest that our approach can decrease the number of tests required to regularly screen the student body, and that these reductions are quite robust to assigning pools based on partially observed or noisy versions of the network.
The COVID-19 pandemic created an unprecedented global health crisis. Recent studies suggest that socially vulnerable communities were disproportionately impacted, although findings are mixed. To quantify social vulnerability in the US, many studies rely on the Social Vulnerability Index (SVI), a county-level measure comprising 15 census variables. Typically, the SVI is modelled in an additive manner, which may obscure non-linear or interactive associations, further contributing to inconsistent findings. As a more robust alternative, we propose a negative binomial Bayesian kernel machine regression (BKMR) model to investigate dynamic associations between social vulnerability and COVID-19 death rates, thus extending BKMR to the count data setting. The model produces a 'vulnerability effect' that quantifies the impact of vulnerability on COVID-19 death rates in each county. The method can also identify the relative importance of various SVI variables and make future predictions as county vulnerability profiles evolve. To capture spatio-temporal heterogeneity, the model incorporates spatial effects, county-level covariates, and smooth temporal functions. For Bayesian computation, we propose a tractable data-augmented Gibbs sampler. We conduct a simulation study to highlight the approach and apply the method to a study of COVID-19 deaths in the US state of South Carolina during the 2021 calendar year.
We consider unsupervised classification by means of a latent multinomial variable which categorizes a scalar response into one of the L components of a mixture model which incorporates scalar and functional covariates. This process can be thought as a hierarchical model with the first level modelling a scalar response according to a mixture of parametric distributions and the second level modelling the mixture probabilities by means of a generalized linear model with functional and scalar covariates. The traditional approach of treating functional covariates as vectors not only suffers from the curse of dimensionality, since functional covariates can be measured at very small intervals leading to a highly parametrized model, but also does not take into account the nature of the data. We use basis expansions to reduce the dimensionality and a Bayesian approach for estimating the parameters while providing predictions of the latent classification vector. The method is motivated by two data examples that are not easily handled by existing methods. The first example concerns identifying placebo responders on a clinical trial (normal mixture model) and the other predicting illness for milking cows (zero-inflated mixture of the Poisson model).
In COVID-19 surveillance, detecting significant case increases within regions over specific periods is crucial. Classical methods, typically relying on strict parametric assumptions, struggle with the rare events characteristic of COVID-19's early spread. An alternative strategy is employing nonparametric approaches based on p-value combination methods. However, initial COVID-19 outbreaks across regions exhibit varying signal sparsity levels, while existing p-value combination methods demonstrate power in detecting either moderately sparse or ultra sparse signals in practice, but not both. We present a modified Fisher's method, utilizing a weakly geometric system-based search strategy to adapt across the entire spectrum of signal sparsity. Our method is theoretically and numerically powerful across the whole spectrum of sparsity. Under mild conditions, we examine the robustness of our method in combining approximated p-values, demonstrating its powerful performance even when the number of p-values far surpasses the sample sizes for their derivation, offering a novel nonparametric strategy for COVID-19 surveillance. An efficient algorithm is developed to calculate the p-value of our method. Focusing on the early COVID-19 surveillance in the United States, our method consistently detects outbreaks across regions with varying signal sparsity, uncovering diverse patterns of COVID-19's spread, while competing methods struggle with either ultra-sparse or moderately sparse signals.
The incremental cost-effectiveness ratio (ICER) and incremental net benefit (INB) are widely used for cost-effectiveness analysis. We develop methods for estimation and inference for the ICER and INB which use the semiparametric stratified Cox proportional hazard model, allowing for adjustment for risk factors. Since in public health settings, patients often begin treatment after they become eligible, we account for delay times in treatment initiation. Excellent finite sample properties of the proposed estimator are demonstrated in an extensive simulation study under different delay scenarios. We apply the proposed method to evaluate the cost-effectiveness of switching treatments among AIDS patients in Tanzania.
Network meta-analysis is a powerful tool to synthesize evidence from independent studies and compare multiple treatments simultaneously. A critical task of performing a network meta-analysis is to offer ranks of all available treatment options for a specific disease outcome. Frequently, the estimated treatment rankings are accompanied by a large amount of uncertainty, suffer from multiplicity issues, and rarely permit possible ties of treatments with similar performance. These issues make interpreting rankings problematic as they are often treated as absolute metrics. To address these shortcomings, we formulate a ranking strategy that adapts to scenarios with high-order uncertainty by producing more conservative results. This improves the interpretability while simultaneously accounting for multiple comparisons. To admit ties between treatment effects in cases where differences between treatment effects are negligible, we also develop a Bayesian non-parametric approach for network meta-analysis. The approach capitalizes on the induced clustering mechanism of Bayesian non-parametric methods, producing a positive probability that two treatment effects are equal. We demonstrate the utility of the procedure through numerical experiments and a network meta-analysis designed to study antidepressant treatments.
The motivation for this paper is to determine factors associated with time-to-fertility treatment (TTFT) among women currently attempting pregnancy in a cross-sectional sample. Challenges arise due to dependence between time-to-pregnancy (TTP) and TTFT. We propose appending a marginal accelerated failure time model to identify risk factors of TTFT with a model for TTP where fertility treatment is included as a time-varying treatment to account for their dependence. The latter requires extending backwards recurrence survival methods to incorporate time-varying covariates with time-varying coefficients. Since backwards recurrence survival methods are a function of mean survival, computational difficulties arise in formulating mean survival when fertility treatment is unobserved, i.e. when TTFT is censored. We address these challenges by developing computationally friendly forms for the double expectation of TTP and TTFT. The performance is validated via comprehensive simulation studies. We apply our approach to the National Survey of Family Growth and explore factors related to prolonged TTFT in the U.S.
Wildland fire smoke exposures are an increasing threat to public health, highlighting the need for studying the effects of protective behaviours on reducing health outcomes. Emerging smartphone applications provide unprecedented opportunities to deliver health risk communication messages to a large number of individuals in real-time and subsequently study the effectiveness, but also pose methodological challenges. Smoke Sense, a citizen science project, provides an interactive smartphone app platform for participants to engage with information about air quality, and ways to record their own health symptoms and actions taken to reduce smoke exposure. We propose a doubly robust estimator of the structural nested mean model that accounts for spatially and time-varying effects via a local estimating equation approach with geographical kernel weighting. Moreover, our analytical framework also handles informative missingness by inverse probability weighting of estimating functions. We evaluate the method using extensive simulation studies and apply it to Smoke Sense data to increase the knowledge base about the relationship between health preventive measures and health-related outcomes. Our results show that the protective behaviours' effects vary over space and time and find that protective behaviours have more significant effects on reducing health symptoms in the Southwest than the Northwest region of the U.S.