Seminar

Phylogenetic analysis: expectations, reality and bridging the gap between the two

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Phylogenetic tree reconstruction and multiple sequence alignment are fundamental tasks in phylogenetics that should ideally be performed jointly, since trees and alignments are inherently circularly dependent. In reality, however, the two structures are more often than not inferred sequentially as independent estimates. Moreover, aligners infer MSAs that contain gaps representing insertions and deletions, while most standard likelihood-based phylogenetic tree reconstruction tools take the MSA as input and proceed to filter out gappy columns due to the lack means to include them in the phylogenetic likelihood computation. A way to solve this disparity in modelling is to include evolutionary models that explicitly account for indels in the likelihood, allowing us to use all available data in inference without having to hack our way around unequal sequence lengths.

In this talk, we will discuss the status quo in phylogenetic modelling and talk about the current work in Maria Anisimova's group on an efficient joint inference method that will hopefully bring us one step closer to making our expectations a reality.

Two MSc student presentations (Liam Gilson & Giuliano Netto Flores Cruz)

To join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11:00am – 11:30am

Speaker: Liam Gilson, UBC Statistics MSc student

Title: Extending an invariant models method to domain generalization settings in ecological data

Abstract: Applied ecology and forest ecology in particular are challenged by changing environmental conditions, especially climate change, situations which often necessitate prediction under unobserved future conditions. This can be described as a domain generalization problem, a setting in which examples from the test task are not available at training time. Present, commonly used methods for model selection are often inappropriate in this setting, and do not emphasize reducing catastrophic error under domain shift. I propose an extension of a model selection method applicable to a limited domain generalization setting, invariant models, to situations commonly encountered in ecological data, specifically where linear mixed-effect models might be applied. In these situations, some predictor variables may be unobserved or otherwise unavailable at training time, complicating the domain generalization setting further. This method is investigated in a series of simulations, and on an example forest ecology dataset.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Giuliano Netto Flores Cruz, UBC Bioinformatics MSc student

Title: Evaluating omics-based tests with Bayesian Decision Curve Analysis

Abstract: Omics-based tests (OBTs) combine high-dimensional omics features into clinical prediction models that predict diagnosis, prognosis, or treatment effects. Past incidences of premature implementation of OBTs into clinical trials have demonstrated the need for increased rigour in their clinical evaluation. However, their performance assessment is often limited to classification metrics such as sensitivity and specificity, with little regard for formal analysis of clinical decision-making. Decision curve analysis (DCA) complements classification metrics by combining classical assessment of predictive performance with the consequences of using a test or model to guide clinical decisions. In DCA, the best clinical decision strategy, such as diagnosing or treating based on an OBT, is the one that maximizes the concept of net benefit: the net number of true positives (or negatives) provided by a given clinical decision strategy. Before reaching real patients, we must be sufficiently confident that new OBTs actually provide superior clinical decision strategies, as compared to default, standard-of-care strategies. Trained on hundreds to thousands of features, OBTs are particularly prone to chance results. In this context, the present work develops parametric Bayesian approaches to DCA that allow uncertainty quantification around four fundamental concerns when evaluating OBT-guided clinical decision strategies: (i) which strategies are clinically useful, (ii) what is the best available decision strategy, (iii) direct pairwise comparisons between strategies, and (iv) what is the consequence of the current level of uncertainty. We evaluate the methods using simulation studies and present a comprehensive case study. We also provide an application to a recently-developed OBT for multi-cancer early detection. Software implementation of the method is freely available in the bayesDCA R package. Ultimately, the Bayesian DCA workflow may help clinicians and health policymakers make better-informed decisions when choosing and implementing clinical decision strategies based on OBTs.

arXiv preprint: https://arxiv.org/abs/2308.02067/

Risk Prediction for Premature Menopause in Childhood Cancer Survivors

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Female childhood cancer survivors are at much higher risk of premature menopause (a.k.a. primary ovarian insufficiency - POI), which causes infertility and the accompanied wide range of post-menopausal symptoms. To counsel survivors, accurate risk estimates at different ages are needed. In this talk, I will discuss our recently completed study where we used data from 7891 participants in the childhood cancer survivors study to develop and 1349 survivors in the St Jude Life study (SJLIFE) to externally validate prediction algorithms based on age-specific logistic regression and XGBoost, as well as the added contributions of relevant polygenic risk scores (PRSs). Given cancer diagnosis and treatment information, these POI risk prediction algorithms perform excellently (AUROCs ranging from 0.91-0.96 in SJLIFE) and can inform decisions regarding fertility preservation interventions among female childhood cancer survivors.

Regularized Relative Risk Regression: A Non-GLM Approach with Emphasis on Large p, Small N Simulations

To join this seminar virtually: please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Binary regression models, such as Poisson and logistic regression, are commonly employed in clinical studies to estimate measures like relative risks (RR) or odds ratios (OR). While RR are preferred for their straightforward interpretation, logistic regression, which models OR, is the most widely used approach. However, it only provides a reliable estimate of RR when the outcome of interest is infrequent. Meanwhile, the Poisson regression can estimate RR directly but can produce fitted probabilities outside the range of zero and one. To address these challenges, Richardson et al. (2017) proposed a novel binary regression model that estimates RR directly via a log odds-product nuisance model. However, this method encountered challenges in high-dimensional and sparse data estimation (p > N). To address these issues, this study introduces an estimator founded on the binary regression model, which is further refined with an algorithm using Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) to solve the optimization problem. This algorithm encourages sparsity in the solution and enables variable selection, thereby improving the model’s utility for high-dimensional and sparse data. Finally, we examine the properties of our estimator in a simulation study that focuses on cases when p > N.

Two UBC Statistics MSc student presentations (Nirupama Tamvada & Maggie Liu)

To join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11:00am – 11:30am

Speaker: Nirupama Tamvada, UBC Statistics MSc student

Title: Penalized Competing Risks Analysis using Case-base sampling

Abstract: In biomedical studies, quantifying the association of prognostic genes/markers on the time-to-event is crucial for predicting a patient's risk of disease based on their specific covariate profile. Modelling competing risks is essential in such studies, as patients may be susceptible to multiple mutually exclusive events, such as death from alternative causes. Existing methods for competing risks often yield coefficient estimates that lack interpretability, as they cannot be associated with the event rate. Moreover, the high dimensionality of genomic data, where the number of variables exceeds the number of subjects, presents a significant challenge. In this work, we propose a novel approach that involves fitting an elastic-net penalized multinomial model using the case-base sampling framework developed by Hanley and Miettinen (2009) to model competing risks survival data. Furthermore, we develop a simple, two-step method known as the de-biased case-base to enhance the prediction performance of the risk of disease. Through a comprehensive simulation study that emulates biomedical data, we show that the case-base method is competent in terms of variable selection and survival prediction, particularly in scenarios such as non-proportional hazards. We additionally showcase the flexibility of this approach in providing smooth-in-time incidence curves, which improve the accuracy of patient risk estimation.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Yitong (Maggie) Liu, UBC Statistics MSc student

Title: Robust Sparse Covariance-Regularized Regression for High-Dimensional Data with Casewise and Cellwise Outliers

Abstract: Modern biomedical datasets, such as those found in genomic studies, often involve a large number of predictor variables relative to the number of observations, pointing to the need for statistical methods specifically designed to handle high-dimensional data. In particular, for a regression task, regularized methods are needed to select a subset of features among the large number of variables that are relevant for predicting a response. The presence of outliers in the data further complicates this task. Many existing robust and sparse regression methods are computationally expensive when the dimensionality of the data is high. Furthermore, most of these previously developed methods were developed under the assumption that outliers occur casewise, which is not always a realistic assumption in high-dimensional settings. We propose a sparse and robust regression method for high-dimensional data that is based on regularized precision matrix estimation. Our method can handle both casewise and cellwise outliers in low- and high-dimensional settings. Through simulation studies, we also compare our method to existing sparse and robust methods by evaluating computational efficiency, prediction performance, and variable selection capabilities.

Statistical Methods for Population Size Estimation on Trees

To Join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Populations of interest are often hidden from data for a variety of reasons, though their magnitude remains important in determining resource allocation and appropriate policy. Increasing data collection and linkage across diverse fields suggests accessible methods of estimating population size with synthesized data are needed. In public health and epidemiology, these linkages often admit a tree structure, with the target population represented by the root, and paths from root-to-leaf representing pathways of care after a health event. We propose an extension to the well-known multiplier method which is applicable to tree-structured data, where multiple subpopulations and corresponding proportions combine to generate a population size estimate via the minimum variance estimator. The methodology is compared a Bayesian hierarchical model, for both simulated and real world data, the latter provided by BC's opioid overdose cohort. Finally, two R packages developed to facilitate the use of these methods on similar applications and lower the technical barrier of implementation will be discussed.

Conditional Inference and Prediction of a Survival Response Based on Vine Copula Models

To join this seminar virtually: please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Copulas are multivariate models that allow the modeling of dependence between variables without restrictions on marginal distributions. Copulas are useful for capturing complex dependence properties such as tail dependence, asymmetry, and nonlinearity. The vine copula can be viewed as an extension of a Gaussian copula after the correlation matrix is reparameterized to a set of algebraically independent correlations and partial correlations. It can be used to construct high-dimensional copula models with flexible dependence structures.  

We propose extensions to conditional inference and prediction methods based on vine copula models. For conditional inference, an algorithm is developed to compute arbitrary conditional distributions of one variable given the other variables for cross-prediction from a single joint distribution fitted by vine copula models. An existing algorithm is also modified to simulate data from a vine copula given that one variable takes extreme values. To predict a right-censored time-to-event response, a vine copula regression model is fitted to the response variable with the explanatory variables. Existing vine copula regression algorithms are modified to provide point and interval predictions for the censored response.

On factor copula-based mixed regression models

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: A copula-based method for mixed regression models is introduced, where the conditional distribution of the response variable, given covariates, is modelled by a parametric family of continuous or discrete distributions, and the effect of a common latent variable pertaining to a cluster is modelled with a factor copula. We show how to estimate the parameters of the copula and the parameters of the margins, and we find the asymptotic behaviour of the estimation errors. Numerical experiments are performed to assess the precision of the estimators for finite samples. An example of an application is given using COVID-19 vaccination hesitancy from several countries. Computations are based on R package CopulaGAMM.

Two UBC Statistics MSc student presentations (Giuseppe Tomio & Jana Osea)

To join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11:00am – 11:30am

Speaker: Giuseppe Tomio, UBC Statistics MSc student

Title: A new data driven framework for simulating mendelian randomization data

Abstract: Mendelian randomization (MR) is a causal inference method that allows biostatisticians to leverage DNA measurements to study causal effects with only observed data. Recent advancements like two-sample summary-level MR (TS SL MR) and projects such as IEU GWAS database have lowered the barrier for conducting MR studies and opened the opportunity to mine causal effects in large-scale data sources. In the first part of my presentation, I show that there is a mismatch between how modern TS SL MR data is and how articles that propose popular TS SL MR models conduct their simulations. Next, I propose my solution: a data driven simulation framework for MR data that aims to be realistic, interpretable and easy to use thanks to a complementary R package implementation. As for the results, I show that models perform far better in literature-based simulations compared to more realistic simulations based on my proposed framework. Lastly, I warn that the mismatch between simulated and real data along with the obtained results may lead researchers to have over optimistic expectations about models performance in real applications.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Jana Osea, UBC Statistics MSc student

Title: Enhancing the Robustness of Instrumental Variable Estimation with Potentially Invalid Instruments and its Application to Mendelian Randomization

Abstract: Causal relationships between exposures and outcomes are vital in fields like epidemiology and medicine as they provide valuable insight into disease mechanisms, informing effective interventions. However, a common problem when attempting to extract the causal relationship between an exposure and an outcome in observational studies is the presence of unmeasured confounding. This leads to the exposure being correlated with the error term known as endogeneity. In order to obtain consistent estimates of the causal effect in the presence of endogeneity, we may use instrumetal variables (IVs) which are correlated with outcome only through its effect on the exposure. This captures the relationship between the exposure and outcome that is unaffected by the endogeneity. The most common use of IVs in the field of epidemiology and medicine is Mendelian Randomization (MR) which uses genetic variants as IVs. However, the validity of the IVs is often questionable due to the presence of pleiotropy and linkage disequilibrium, where the genetic variants affect the outcome through pathways other than the exposure. Furthermore, exposure and outcome data are often contaminated by outliers. In this thesis, we propose a novel algorithm, the robustified some valid some invalid instrumental variable estimator (rsisVIVE), that obtains estimates of the causal effect of an exposure on an outcome in the presence of invalid IVs and high levels of endogeneity while tolerating large proportions of contamination. The algorithm is based on the sisVIVE algorithm of Kang et. al. (2016) but we propose using robust estimators in both stages of the sisVIVE. Simulation results show that the rsisVIVE more accurately estimates the causal parameter than the sisVIVE when IVs are weak and outperforms competitor IV estimators in all cases when there is contamination.

Post hoc inference for genomics and neuroimaging

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Analyzing and interpreting high-dimensional data in applications such as genomics and neuroimaging often requires the simultaneous testing of a large number of hypotheses. In such multiple testing situations, the state-of-the-art approach selects one subset of hypotheses whose False Discovery Rate (FDR) in controlled. The FDR is the expected proportion of false positives (FDP) among selected hypotheses.

In contrast, post hoc inference aims to build confidence bounds on the FDP contained in arbitrary subsets of hypotheses, leading to improvements in interpretability and reproducibility of the results. We show how to construct and efficiently implement post hoc bounds that are adaptive to unknown dependence. We illustrate their application to differential gene expression studies in genomics, and fMRI studies in neuroimaging.