Seminar

Quantifying Uncertainty in Clustering Analysis: With Applications in Genetic Data and Phylogenetic Trees

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: Clustering analysis is an area of unsupervised learning that carries great potential in analyzing the increasingly immense amount of available data. Whereas it is straightforward to apply clustering on datasets, it is much less so when it comes to assessing the quality of the clustering results, since we often lack some notions of ground truth. This, coupled with the fact that two samples collected in di?erent settings may not be the same and hence the respective results may exhibit variations, makes this problem both important and challenging at the same time.

In his 1985 paper Confdence Limits on Phylogenies: An Approach Using the Bootstrap, Felsensteins proposes the idea of the bootstrap probability as a tool to quantify this uncertainty in the context of hierarchical clustering. It has since gained much appreciation and become one of the main tools for this problem, so much so that it is often referred to as a p-value.

We review the 1996 paper by Efron et al., Bootstrap confidence levels for phylogenetic trees, which argues that the bootstrap probability lacks a framework of a model and a null hypothesis for it to be formalized into a statistical p-value. We then explore an alternative interpretation based on the idea of bootstrap as a drop-in replacement for the sampling probability if we have access to the population. We run simulations to explore the performance of the bootstrap probability and how it tracks the target sampling probability.

We find that while the bootstrap probability is indeed a very elegant and simple-to-calculate metric, there are situations in which we can have great confidence in the results and others whereas we should be less so and further analyses may be necessary.

Advancements in Probabilistic Circuits for Tractable Probabilistic Inference

To Join via Zoom: To join this seminar virtually, please register here.

Abstract: Probabilistic circuits are a recent development in machine learning and act at the intersection of neural networks, traditional probabilistic models like mixture models, and propositional logic. Contrary to traditional representations of probabilistic models, circuits utilize a low-level representation through simple arithmetic operations and allow us to guarantee tractability (exact and efficient computation) of certain inferences based on the properties of the circuit itself. Henceforth, they have become a valuable tool to reason about classes of distributions for which inferences can tractably be represented. In this talk, I will briefly review probabilistic circuits and showcase recent advancements. First, I will introduce probabilistic circuits as a representational tool for tractable probabilistic inference focusing on a broader picture. Second, I will discuss recent endeavours to extend probabilistic circuits to model more complex probability distributions. For this, I will focus on work aimed at representing distributions that do not admit a closed-form density and a recent work focusing on circuits that can subtract density without violating non-negativity constraints aka mixtures with negative weights.

Coin Sampling: Gradient-Based Bayesian Inference without Learning Rates

To Join via Zoom: To join this seminar virtually, please register here.

Abstract: In recent years, particle-based variational inference (ParVI) methods such as Stein variational gradient descent (SVGD) have grown in popularity as scalable methods for Bayesian inference. Unfortunately, the properties of such methods invariably depend on hyperparameters such as the learning rate, which must be carefully tuned by the practitioner in order to ensure convergence to the target measure at a suitable rate. In this work, we introduce a suite of new particle-based methods for scalable Bayesian inference based on coin betting, which are entirely learning-rate free. We illustrate the performance of our approach on a range of numerical examples, including several high-dimensional models and datasets, demonstrating comparable performance to other ParVI algorithms with no need to tune a learning rate.

Inference with joint models under misspecified random effects distributions

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from ea@stat.ubc.ca

Abstract: Joint models are often used to analyze longitudinal and time-to-event data, where latent random effects are used to describe the association between the two outcome processes. It is typically assumed that the random effects follow a multivariate normal distribution. The likelihood analysis under a correctly specified random effects distribution may provide valid inferences. But if the distribution is misspecified, then the maximum likelihood (ML) estimators can be biased and hence may lead to invalid inferences. In this talk, I will discuss the joint analysis under various types of normal and nonnormal random effects. We propose a robust method of estimation that can address uncertainties in the distribution of random effects. I will discuss empirical properties of the proposed estimators based on a simulation study. I will also present an application using a large clinical dataset from the genetic and inflammatory markers of sepsis (GenIMS) study.

Approximate Marginal Likelihood Inference in Mixed Models for Grouped Data

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: I introduce a method for approximate marginal likelihood inference via adaptive Gaussian quadrature in mixed models with a single grouping factor. The core technical contributions are (a) an algorithm for computing the exact gradient of the approximate log marginal likelihood and (b) a useful parameterization of the multivariate Gaussian. The former leads to efficient quasi-Newton optimization of the marginal likelihood that is several times faster than established methods; the latter gives Wald confidence intervals for random effects variances that attain nominal coverage and low bias if enough quadrature points are used. The Laplace approximation is a special case of the method and is shown in simulations to perform exceptionally poorly for binary random slopes models, but this is mitigated by just adding more quadrature points.

Phylogenetic analysis: expectations, reality and bridging the gap between the two

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Phylogenetic tree reconstruction and multiple sequence alignment are fundamental tasks in phylogenetics that should ideally be performed jointly, since trees and alignments are inherently circularly dependent. In reality, however, the two structures are more often than not inferred sequentially as independent estimates. Moreover, aligners infer MSAs that contain gaps representing insertions and deletions, while most standard likelihood-based phylogenetic tree reconstruction tools take the MSA as input and proceed to filter out gappy columns due to the lack means to include them in the phylogenetic likelihood computation. A way to solve this disparity in modelling is to include evolutionary models that explicitly account for indels in the likelihood, allowing us to use all available data in inference without having to hack our way around unequal sequence lengths.

In this talk, we will discuss the status quo in phylogenetic modelling and talk about the current work in Maria Anisimova's group on an efficient joint inference method that will hopefully bring us one step closer to making our expectations a reality.

Two MSc student presentations (Liam Gilson & Giuliano Netto Flores Cruz)

To join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11:00am – 11:30am

Speaker: Liam Gilson, UBC Statistics MSc student

Title: Extending an invariant models method to domain generalization settings in ecological data

Abstract: Applied ecology and forest ecology in particular are challenged by changing environmental conditions, especially climate change, situations which often necessitate prediction under unobserved future conditions. This can be described as a domain generalization problem, a setting in which examples from the test task are not available at training time. Present, commonly used methods for model selection are often inappropriate in this setting, and do not emphasize reducing catastrophic error under domain shift. I propose an extension of a model selection method applicable to a limited domain generalization setting, invariant models, to situations commonly encountered in ecological data, specifically where linear mixed-effect models might be applied. In these situations, some predictor variables may be unobserved or otherwise unavailable at training time, complicating the domain generalization setting further. This method is investigated in a series of simulations, and on an example forest ecology dataset.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Giuliano Netto Flores Cruz, UBC Bioinformatics MSc student

Title: Evaluating omics-based tests with Bayesian Decision Curve Analysis

Abstract: Omics-based tests (OBTs) combine high-dimensional omics features into clinical prediction models that predict diagnosis, prognosis, or treatment effects. Past incidences of premature implementation of OBTs into clinical trials have demonstrated the need for increased rigour in their clinical evaluation. However, their performance assessment is often limited to classification metrics such as sensitivity and specificity, with little regard for formal analysis of clinical decision-making. Decision curve analysis (DCA) complements classification metrics by combining classical assessment of predictive performance with the consequences of using a test or model to guide clinical decisions. In DCA, the best clinical decision strategy, such as diagnosing or treating based on an OBT, is the one that maximizes the concept of net benefit: the net number of true positives (or negatives) provided by a given clinical decision strategy. Before reaching real patients, we must be sufficiently confident that new OBTs actually provide superior clinical decision strategies, as compared to default, standard-of-care strategies. Trained on hundreds to thousands of features, OBTs are particularly prone to chance results. In this context, the present work develops parametric Bayesian approaches to DCA that allow uncertainty quantification around four fundamental concerns when evaluating OBT-guided clinical decision strategies: (i) which strategies are clinically useful, (ii) what is the best available decision strategy, (iii) direct pairwise comparisons between strategies, and (iv) what is the consequence of the current level of uncertainty. We evaluate the methods using simulation studies and present a comprehensive case study. We also provide an application to a recently-developed OBT for multi-cancer early detection. Software implementation of the method is freely available in the bayesDCA R package. Ultimately, the Bayesian DCA workflow may help clinicians and health policymakers make better-informed decisions when choosing and implementing clinical decision strategies based on OBTs.

arXiv preprint: https://arxiv.org/abs/2308.02067/

Risk Prediction for Premature Menopause in Childhood Cancer Survivors

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Female childhood cancer survivors are at much higher risk of premature menopause (a.k.a. primary ovarian insufficiency - POI), which causes infertility and the accompanied wide range of post-menopausal symptoms. To counsel survivors, accurate risk estimates at different ages are needed. In this talk, I will discuss our recently completed study where we used data from 7891 participants in the childhood cancer survivors study to develop and 1349 survivors in the St Jude Life study (SJLIFE) to externally validate prediction algorithms based on age-specific logistic regression and XGBoost, as well as the added contributions of relevant polygenic risk scores (PRSs). Given cancer diagnosis and treatment information, these POI risk prediction algorithms perform excellently (AUROCs ranging from 0.91-0.96 in SJLIFE) and can inform decisions regarding fertility preservation interventions among female childhood cancer survivors.

Regularized Relative Risk Regression: A Non-GLM Approach with Emphasis on Large p, Small N Simulations

To join this seminar virtually: please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Binary regression models, such as Poisson and logistic regression, are commonly employed in clinical studies to estimate measures like relative risks (RR) or odds ratios (OR). While RR are preferred for their straightforward interpretation, logistic regression, which models OR, is the most widely used approach. However, it only provides a reliable estimate of RR when the outcome of interest is infrequent. Meanwhile, the Poisson regression can estimate RR directly but can produce fitted probabilities outside the range of zero and one. To address these challenges, Richardson et al. (2017) proposed a novel binary regression model that estimates RR directly via a log odds-product nuisance model. However, this method encountered challenges in high-dimensional and sparse data estimation (p > N). To address these issues, this study introduces an estimator founded on the binary regression model, which is further refined with an algorithm using Fast Iterative Shrinkage-Thresholding Algorithm (FISTA) to solve the optimization problem. This algorithm encourages sparsity in the solution and enables variable selection, thereby improving the model’s utility for high-dimensional and sparse data. Finally, we examine the properties of our estimator in a simulation study that focuses on cases when p > N.

Two UBC Statistics MSc student presentations (Nirupama Tamvada & Maggie Liu)

To join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11:00am – 11:30am

Speaker: Nirupama Tamvada, UBC Statistics MSc student

Title: Penalized Competing Risks Analysis using Case-base sampling

Abstract: In biomedical studies, quantifying the association of prognostic genes/markers on the time-to-event is crucial for predicting a patient's risk of disease based on their specific covariate profile. Modelling competing risks is essential in such studies, as patients may be susceptible to multiple mutually exclusive events, such as death from alternative causes. Existing methods for competing risks often yield coefficient estimates that lack interpretability, as they cannot be associated with the event rate. Moreover, the high dimensionality of genomic data, where the number of variables exceeds the number of subjects, presents a significant challenge. In this work, we propose a novel approach that involves fitting an elastic-net penalized multinomial model using the case-base sampling framework developed by Hanley and Miettinen (2009) to model competing risks survival data. Furthermore, we develop a simple, two-step method known as the de-biased case-base to enhance the prediction performance of the risk of disease. Through a comprehensive simulation study that emulates biomedical data, we show that the case-base method is competent in terms of variable selection and survival prediction, particularly in scenarios such as non-proportional hazards. We additionally showcase the flexibility of this approach in providing smooth-in-time incidence curves, which improve the accuracy of patient risk estimation.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Yitong (Maggie) Liu, UBC Statistics MSc student

Title: Robust Sparse Covariance-Regularized Regression for High-Dimensional Data with Casewise and Cellwise Outliers

Abstract: Modern biomedical datasets, such as those found in genomic studies, often involve a large number of predictor variables relative to the number of observations, pointing to the need for statistical methods specifically designed to handle high-dimensional data. In particular, for a regression task, regularized methods are needed to select a subset of features among the large number of variables that are relevant for predicting a response. The presence of outliers in the data further complicates this task. Many existing robust and sparse regression methods are computationally expensive when the dimensionality of the data is high. Furthermore, most of these previously developed methods were developed under the assumption that outliers occur casewise, which is not always a realistic assumption in high-dimensional settings. We propose a sparse and robust regression method for high-dimensional data that is based on regularized precision matrix estimation. Our method can handle both casewise and cellwise outliers in low- and high-dimensional settings. Through simulation studies, we also compare our method to existing sparse and robust methods by evaluating computational efficiency, prediction performance, and variable selection capabilities.