Seminar

Toward large-scale Bayesian inference with structured generative models

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: Conditional generative models are an increasingly popular machine learning approach for performing Bayesian Inference in large-scale scientific applications. Despite their success, the sample-complexity of these methods often scale poorly with the growing dimensions of model parameters and observations. These models require large training datasets, which are typically unavailable with computationally expensive forward models as in numerical weather prediction. In these settings, it becomes crucial to identify the low-dimensional structure in posterior distributions and encode it in generative models for accurate inference. In this presentation, I will introduce an information-theoretic perspective for reducing the dimensions of parameters and observations in inverse problems with guarantees on the posterior approximation error. I will show how to identify relevant subspaces for these variables from either gradient evaluations of the log-likelihood function or by using score-based generative models when gradients are not available. The benefit of the dimension reduction approach will be showcased on a turbulent flow estimation problem from aerodynamics where traditional methods are unstable in small sample regimes.

Advancing Biomedical Data Science: Leveraging ChatGPT for genomics data & Testing data-driven hypothesis post-clustering

Date/Time: Thursday, January 25, 2024, 10:25am to 11:25am

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: My research centers around bringing statistical insights and understanding to the practice of modern data science, and I will cover two projects related to this research vision in this talk.

In the first part of the talk, I will consider how to leverage large language models (LLMs) such as ChatGPT for biomedical discovery. While significant progress has been made in customizing large language models for biomedical data, these models often require extensive data curation and resource-intensive training. In the context of single-cell RNA-sequencing data, I will show that we can achieve surprisingly competitive results on many downstream tasks via a much simpler alternative: I input textual descriptions of genes into an off-the-shelf LLM, such as ChatGPT, to obtain low-dimensional representations of the genes, or “embeddings.” I then use these embeddings as features in downstream tasks. A similar approach enables LLM-derived embeddings of cells. This work highlights the potential of LLMs to provide meaningful and concise representations for biomedical data, and also raises a number of challenging statistical questions. Addressing these questions requires bringing principled statistical thinking to the practice of modern data science.

The second part of my talk is motivated by the practice of testing data-driven hypotheses. In biomedical sciences, it has become increasingly common to collect massive datasets without a pre-specified research question. In this setting, a data analyst might use the data both to generate a research question, and to test the associated null hypothesis. For example, in single-cell RNA-sequencing analyses, researchers often first cluster the cells, and then test for differences in the expected gene expression levels between the clusters to quantify up- or down-regulation of genes, annotate known cell types, and identify new cell types. However, this popular practice is invalid from a statistical perspective: once we have used the data to generate hypotheses, standard statistical inference tools are no longer valid. To tackle this problem, I developed a conditional selective approach to test for a difference in means between pairs of clusters obtained via k-means clustering. The proposed approach has appropriate statistical guarantees (e.g., selective Type 1 error control).

This talk features joint work with Lucy Gao (University of British Columbia), Daniela Witten (University of Washington), and James Zou (Stanford University).

Identifiable and interpretable nonparametric factor analysis

Date/Time: Thursday, January 18, 2024, 10:25am to 11:25am

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: Factor models have been widely used to summarize the variability of high-dimensional data through a set of factors with much lower dimensionality. Gaussian linear factor models have been particularly popular due to their interpretability and ease of computation. However, in practice, data often violate the multivariate Gaussian assumption. To characterize higher-order dependence and nonlinearity, models that include factors as predictors in flexible multivariate regression are popular, with GP-LVMs using Gaussian process (GP) priors for the regression function and VAEs using deep neural networks. Unfortunately, such approaches lack identifiability and interpretability and tend to produce brittle and non-reproducible results. To address these problems by simplifying the nonparametric factor model while maintaining flexibility, we propose the NIFTY framework, which parsimoniously transforms uniform latent variables using one-dimensional nonlinear mappings and then applies a linear generative model. The induced multivariate distribution falls into a flexible class while maintaining simple computation and interpretation. We prove that this model is identifiable and empirically study NIFTY using simulated data, observing good performance in density estimation and data visualization. We then apply NIFTY to bird song data in an environmental monitoring application.

Learning-Rate-Free Methods for Scalable Bayesian Inference

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: In recent years, particle-based variational inference (ParVI) methods such as Stein variational gradient descent have grown in popularity as scalable methods for Bayesian inference. Unfortunately, the properties of such methods invariably depend on hyperparameters such as the learning rate, which must be carefully tuned by practitioners in order to ensure convergence to the target measure at a suitable rate. In this work, we introduce a suite of new particle-based methods for scalable Bayesian inference based on coin betting, which are entirely learning-rate free. We illustrate the performance of our approach on a range of numerical examples, including several high-dimensional models and datasets, demonstrating comparable or superior performance to other ParVI algorithms with no need to tune a learning rate.

The design and analysis of annealing algorithms for sampling and generative modelling

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: Generating samples from complex probability distributions is a fundamental challenge in statistical modelling. This problem is called "sampling" when we can access only an un-normalised density and "generative modelling" when we can only access a dataset of existing samples. In practice, this is generally impossible, and we must introduce a simpler reference distribution, such as a Gaussian, and manipulate its density and samples to approximate the target. In general, direct inference is reliable when the reference is close to the target and fragile when it is not. Annealing is a popular technique motivated by this principle and introduces a sequence of distributions that interpolates between the reference and target, ensuring the neighbouring distributions are close. An annealing algorithm specifies how to traverse this bridge of distributions to incrementally transform samples from the reference into samples approximating the target.

In this talk, I will discuss how annealing can be applied to sampling and generative modelling. As a case study, I will introduce parallel tempering (PT) and denoising diffusion models (DDM), two annealing algorithms recently gaining popularity in the sampling and generative modelling literature. I will show how our analysis of PT and DDM can provide insights into understanding the design choices behind their success and motivate the design choices for other annealing algorithms in sampling and generative modelling. Finally, I will discuss the applications of this body of work to Bayesian inference and PDE surrogate modelling.

Quantifying Uncertainty in Clustering Analysis: With Applications in Genetic Data and Phylogenetic Trees

To join this seminar virtually: Please request Zoom connection details from ea@stat.ubc.ca

Abstract: Clustering analysis is an area of unsupervised learning that carries great potential in analyzing the increasingly immense amount of available data. Whereas it is straightforward to apply clustering on datasets, it is much less so when it comes to assessing the quality of the clustering results, since we often lack some notions of ground truth. This, coupled with the fact that two samples collected in di?erent settings may not be the same and hence the respective results may exhibit variations, makes this problem both important and challenging at the same time.

In his 1985 paper Confdence Limits on Phylogenies: An Approach Using the Bootstrap, Felsensteins proposes the idea of the bootstrap probability as a tool to quantify this uncertainty in the context of hierarchical clustering. It has since gained much appreciation and become one of the main tools for this problem, so much so that it is often referred to as a p-value.

We review the 1996 paper by Efron et al., Bootstrap confidence levels for phylogenetic trees, which argues that the bootstrap probability lacks a framework of a model and a null hypothesis for it to be formalized into a statistical p-value. We then explore an alternative interpretation based on the idea of bootstrap as a drop-in replacement for the sampling probability if we have access to the population. We run simulations to explore the performance of the bootstrap probability and how it tracks the target sampling probability.

We find that while the bootstrap probability is indeed a very elegant and simple-to-calculate metric, there are situations in which we can have great confidence in the results and others whereas we should be less so and further analyses may be necessary.

Advancements in Probabilistic Circuits for Tractable Probabilistic Inference

To Join via Zoom: To join this seminar virtually, please register here.

Abstract: Probabilistic circuits are a recent development in machine learning and act at the intersection of neural networks, traditional probabilistic models like mixture models, and propositional logic. Contrary to traditional representations of probabilistic models, circuits utilize a low-level representation through simple arithmetic operations and allow us to guarantee tractability (exact and efficient computation) of certain inferences based on the properties of the circuit itself. Henceforth, they have become a valuable tool to reason about classes of distributions for which inferences can tractably be represented. In this talk, I will briefly review probabilistic circuits and showcase recent advancements. First, I will introduce probabilistic circuits as a representational tool for tractable probabilistic inference focusing on a broader picture. Second, I will discuss recent endeavours to extend probabilistic circuits to model more complex probability distributions. For this, I will focus on work aimed at representing distributions that do not admit a closed-form density and a recent work focusing on circuits that can subtract density without violating non-negativity constraints aka mixtures with negative weights.

Coin Sampling: Gradient-Based Bayesian Inference without Learning Rates

To Join via Zoom: To join this seminar virtually, please register here.

Abstract: In recent years, particle-based variational inference (ParVI) methods such as Stein variational gradient descent (SVGD) have grown in popularity as scalable methods for Bayesian inference. Unfortunately, the properties of such methods invariably depend on hyperparameters such as the learning rate, which must be carefully tuned by the practitioner in order to ensure convergence to the target measure at a suitable rate. In this work, we introduce a suite of new particle-based methods for scalable Bayesian inference based on coin betting, which are entirely learning-rate free. We illustrate the performance of our approach on a range of numerical examples, including several high-dimensional models and datasets, demonstrating comparable performance to other ParVI algorithms with no need to tune a learning rate.

Inference with joint models under misspecified random effects distributions

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from ea@stat.ubc.ca

Abstract: Joint models are often used to analyze longitudinal and time-to-event data, where latent random effects are used to describe the association between the two outcome processes. It is typically assumed that the random effects follow a multivariate normal distribution. The likelihood analysis under a correctly specified random effects distribution may provide valid inferences. But if the distribution is misspecified, then the maximum likelihood (ML) estimators can be biased and hence may lead to invalid inferences. In this talk, I will discuss the joint analysis under various types of normal and nonnormal random effects. We propose a robust method of estimation that can address uncertainties in the distribution of random effects. I will discuss empirical properties of the proposed estimators based on a simulation study. I will also present an application using a large clinical dataset from the genetic and inflammatory markers of sepsis (GenIMS) study.

Approximate Marginal Likelihood Inference in Mixed Models for Grouped Data

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: I introduce a method for approximate marginal likelihood inference via adaptive Gaussian quadrature in mixed models with a single grouping factor. The core technical contributions are (a) an algorithm for computing the exact gradient of the approximate log marginal likelihood and (b) a useful parameterization of the multivariate Gaussian. The former leads to efficient quasi-Newton optimization of the marginal likelihood that is several times faster than established methods; the latter gives Wald confidence intervals for random effects variances that attain nominal coverage and low bias if enough quadrature points are used. The Laplace approximation is a special case of the method and is shown in simulations to perform exceptionally poorly for binary random slopes models, but this is mitigated by just adding more quadrature points.