Seminar

Epidemic models: can we make them behave better?

To Join this seminar: Please request Zoom connection details from headsec@stat.ubc.ca

Abstract: The COVID-19 pandemic has illustrated both the utility and limitation of using epidemic models for understanding and forecasting disease spread. One of the many difficulties in modelling epidemic spread is that caused by behavioural change in the underlying population. This can be a major issue in public health since, as we have seen during the COVID-19 pandemic, behaviour in the population can change drastically as infection levels vary, both due to government mandates and personal decisions. Such changes in the underlying population result in major changes in transmission dynamics of the disease, making the modelling challenges. However, these issues arise in agriculture and public health, as changes in farming practice are also often observed as disease prevalence changes. We propose a model formulation where time-varying transmission is captured by the level of alarm in the population and specified as a function of the past epidemic trajectory. The model is set in a data-augmented Bayesian framework as epidemic data are often only partially observed, and we can utilize prior information to help with parameter identifiability. We investigate the identifiability of the population alarm across a wide range of scenarios, using both parametric functions and non-parametric Gaussian process and splines. The benefit and utility of the proposed approach is illustrated through an application to COVID-19 data from New York City.

A corrected Clarke test for model selection

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: We introduce a large family of model selection tests based on the expectation of an arbitrary, possibly non-smooth, parametric criterion function of the data. It covers the case of strictly locally non-nested models and some overlapping models. The asymptotic theory of the proposed test statistic will be presented. A general exchangeable bootstrap scheme allows the evaluation of its limiting law as well as its asymptotic variance. In a simulation study, we empirically verify the distributional approximation of our test statistic in a finite sample and examine the empirical level and power of the corresponding model selection tests in various settings. Finally, an analysis of a financial dataset illustrates the proposed model selection procedure at work. The talk is based on a joint work with Florian Brueck and Jean-David Fermanian.

Markov Chain Monte Carlo and Langevin equations on a Stratification

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Title: Markov Chain Monte Carlo and Langevin equations on a Stratification

Abstract: Many sampling problems involve constraints — statistical models may involve relationships between parameters; noisy physical systems may involve stiff forces that constrain the system near a manifold, such as stiff bonds between particles. In some cases the constraints are not fixed, but rather can be added or removed (such as when bonds between particles form or break), so the probability measure of interest lives on sets of different dimensions. How can we sample from such a measure? I will introduce an MCMC algorithm to sample a probability measure supported on a stratification: a union of manifolds of different dimensions, glued together at their boundaries in a nice enough way. I will show this can accelerate simulations of interacting particles by up to several orders of magnitude. Then I will talk about our progress toward simulating Langevin equations on a stratification, which harnesses the theory of sticky diffusions. These algorithms are motivated by applications to systems of interacting particles, and I will be interested to learn about other areas of application. 

Co-op Report: Cardiovascular Network of Canada (CANet)

To Join this seminar: Please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: This report outlines my experience as a student biostatistician during an 8-month co-op; I highly recommend the experience to other students as a means to gain valuable work experience and to bridge the gap between the theoretical and the practical application of concepts taught in the MSc Statistics (Biostatistics) program.

CANet is an NCE-funded research network based in London, Ontario, focused on developing virtual care platforms for cardiovascular and other related health conditions. I was part of a clinical team performing statistical analyses and developing statistical protocols for projects including virtual care of atrial fibrillation and Post-MI management. I also lead the analysis of a clinical trial assessing a virtual care model for COVID-19 that became the basis of a research project undertaken with the supervision of Daniel McDonald).

Bayesian Models for Hierarchical Clustering of Network Data

To Join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Network data exist in many forms, like social networks, or interactions between cell proteins. Generally, they represent relational information between interacting entities. In many real-world examples, these entities tend to exhibit grouping structure. For example, the highly connected communities of people within a social network. Uncovering the underlying structure in networks is an important task for studying their composition and behaviour. Hierarchical clustering is a technique for discovering this structure across multiple scales, where a dendrogram represents the full hierarchy of clusters. This talk will explore Bayesian models for hierarchical clustering of network data, which aim to infer the posterior distribution over dendrograms.

The “Hierarchical Random Graph” is likely the most popular Bayesian approach to hierarchical clustering of network data. Yet, due to simplifications made in its inference scheme, we identify some potentially undesirable model behaviour. To rectify these issues, we introduce a general class of models that are characterized by a sampling construction, defining a generative process for simple graphs. We propose four Bayesian models from this class, and derive the marginalized posterior distribution over dendrograms, to isolate the problem of inferring a hierarchical clustering. We implement these models in a probabilistic programming language (Blang) that leverages state-of-the-art approximate inference methods (non-reversible Parallel Tempering). Finally, the empirical performance of our models is demonstrated on examples of real network data.

Two UBC Statistics MSc student presentations (William Laplante & Elvis Cai)

To join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11:00am – 11:30am

Speaker: William Laplante, UBC Statistics MSc student

Title: Improving Uncertainty Quantification of Epidemiological Models with Probabilistic Numerics

Abstract: Recent work in Probabilistic Numerics (PN) – a subfield of machine learning that aims to quantify uncertainty arising from intractable numerical computation – has developed a new class of numerical algorithms to solve ordinary differential equations (ODEs) with a latent force. These solvers pose the problem of numerically computing the solution of ODEs as one of statistical inference: each step of numerical integration is a task of prediction with uncertainty. By formulating a state-space model (SSM) with two likelihoods – one to fit the data, and one to ensure state alignment – and by using Kalman filters and smoothers to estimate the states of the SSM, the estimates (with uncertainty) of an ODE's solution and its latent force are obtained in a single, linear complexity pass. In this work, we demonstrate the practicality of these PN methods in epidemiology by fitting data from the COVID-19 pandemic with a compartmental model – a type of epidemiological model that divides a population in compartments and is expressed as ODEs. To facilitate the fitting process, we propose a "fix-all-vary-one" approach to calibrate the model's hyperparameters, and implement the EM algorithm to estimate the likelihood's covariance. From estimates of the compartmental model's states, we retrieve (1) the model's time-varying contact rate (the latent force), (2) an estimate of the time-varying instantaneous reproductive number, and (3) a prediction for daily and cumulative case counts. Overall, we show that the key feature of PN methods, "uncertainty-awareness", can greatly benefit quantitative epidemiologists that make extensive use of differential equations to describe epidemics.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Elvis (Zhenglun) Cai, UBC Statistics MSc student

Title: Modeling and Estimating the Effective Reproduction Number

Abstract: The effective Reproduction Number (Rt) of an infectious disease is a latent variable that measures the total number of secondary infections generated by an individual on average. It informs policymakers on the virulence of infectious diseases so that they can decide on the type of non-medical intervention that should be implemented. In this project, we model Rt with penalized Poisson regression using the Renewal Equation and provide a framework that can handle various smoothness assumptions of Rt. The penalty terms that are determined by smoothness assumptions yield a convex, separable, but non-differentiable objective function that is solved with the linearized Alternating Direction Method of Multiplier (ADMM). The corresponding algorithm is implemented in our R package, “RtEstim”, with cross validation. We compare the RMSE/RMAE of the estimated Rt and its corresponding case counts between “RtEstim” and “EpiEstim” – one of the most widely used Rt estimation packages in R – using various synthetic datasets. We find that “RtEstim” has a smaller prediction error than “EpiEstim” on most synthetic datasets.

Statistical Imaging of Black Holes using the Event Horizon Telescope

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Title: Statistical Imaging of Black Holes using the Event Horizon Telescope

Abstract: In 2019 and 2021, the Event Horizon Telescope produced the first-ever images of the black holes M87* and Sgr A*, respectively. However, the Event Horizon Telescope is not a regular camera. It does not directly measure the on-sky image, and the high computational cost is converting the estimated sparse data to an actual image. As a result, imaging requires high-performance computing and statistical modeling. In this presentation, I will present the different statistical techniques used by the EHT to analyze the data. The focus will be on recent advances in applying computational Bayesian inference to the imaging problem. I will introduce the statistical model we use, which requires modeling the instrument and the image using non-linear and often weakly non-identifiable models. To sample from this posterior required using novel statistical inference techniques, such as the recently developed non-reversible optimal parallel tempering algorithm developed at UBC. These results demonstrate the potential collaboration between computational statistics and radio imaging and how it can benefit both communities. This collaboration will become more critical shortly with the advent of more powerful telescopes, such as the next-generation Event Horizon telescope, that will increase the data volume and model complexity by 2-3 orders of magnitude.

Statistical implications of group invariance of distributions

Abstract: Consider a large random structure – a random graph, a stochastic process on the line, a random field on the grid – and a function that depends only on a small part of the structure. Now use a family of transformations to ‘move’ the domain of the function over the structure, collect each function value, and average. Under suitable conditions, the law of large numbers generalizes to such averages; that is one of the deep insights of modern ergodic theory. My own recent work with Morgane Austern (Harvard) shows that central limit theorems and other higher-order properties also hold. Loosely speaking, if the i.i.d. assumption of classical statistics is substituted by suitable properties formulated in terms of groups, the fundamental theorems of inference still hold.

VanBUG Bioinformatics Seminar: Yongjin Park

Registration & talk details

Date: Thursday, April 20, 2023

Time: 5:00 PM - 9:00 PM (Pacific Time)

Schedule:

5:00 - 5:45 PM : Meet-The-Speaker *

5:45 - 6:00 PM : Break

6:00 - 6:05 PM : Seminar begins / Announcements

6:05 - 6:25 PM : Presentation by Trainee Speaker

6:25 - 7:20 PM : Presentation by Featured Speaker

7:20 - 9:00 PM : Seminar ends / Networking + light refreshments

* Pending on the number of RSVPs

Please fill out RSVP form if you are interested in attending this seminar in-person (or the meet-the-speaker session before the seminar at 5PM).

Talk Title: Learning deep biology with shallow statistical models

Abstract: As deep learning approaches gained much popularity, typical statistical learning methods were replaced by deep generative models and black-box classification algorithms in many types of biological data analysis, including genetic variant calling, regulatory genomics, high-dimensional data embedding, missing value imputations, and risk predictions. Despite being readily accessible with general-purpose libraries, so-called deep methods demand long hours of training, specialized hardware resources, and a large amount of data; yet, we are often startled at seeing only marginally improved classification performance, model overfitting, or lack of generalizability. Not undermining important advancements made possible by deep models, this talk will seek to showcase that we can deepen our understanding of biological systems with shallow statistical models.

First, I will discuss our scalable algorithm for probabilistic topic modelling in single-cell genomics data. Based on probabilistic topic assignments in each cell, we identify the latent representation of cellular states and heterogeneity, and the latent topic vectors often yield much-improved clustering results than other types of dimensionality reduction methods. We were initially inspired by several empirical observations: (1) Data sets compressed by repeatedly applying random projection operations highlight cell type-specific signature genes. (2) We also noted that rare cell types are better characterized with lowly-expressed genes (in total data) that are typically removed in quality control steps, thus, not included in embedding model estimations. (3) Bulk sequencing data sets are generally less prone to zero inflation or measurement errors. Based on these key findings, we designed a statistical framework termed ASAP--short for Annotating Single-cell data matrix by Approximate Pseudo-bulk projection) to identify cell topics. ASAP seeks to reduce sample size by interactively collapsing cell-level expression vectors into pseudo-bulk vectors in order to accurately perform non-negative matrix factorization to learn topic-specific gene frequency patterns.

Another example will be a sparse regression model with a causal inference flavour that can effectively handle putative confounding issues in a genetic fine-mapping problem. Our idea is rooted in Rubin's causal inference framework (Rubin and Rosenbaum, 1983), with which non-genetic and indirect genetic effects can be cancelled out in implicit adjustment steps. In order to handle the non-binary nature of exposure variables (genetic dosage), our method first matches individuals with one another based on a covariate similarity matrix. If paired individuals share non-genetic factors, then any gene expression changes between them will be removed so that genetic dosage will become a causal factor in the expression divergence. We implemented the algorithm in the SuSiE framework (Wang et al. 2020).

Overall, this talk will reassure us that it is okay to work on a shallow model. In fact, if our scientific goal is straightforward enough to be coded in a simple model armed with intuitive algorithms, such a modelling approach can help uncover deeper layers of biological mechanisms underneath high-dimensional genomics data.

Two UBC Statistics MSc student presentations (Yuwei Yang & Marc Wettengel)

To join this seminar: Please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11:00am – 11:30am

Speaker: Yuwei Yang, UBC Statistics MSc student

Title: Statistical Consulting and Process

Abstract: Improvement in Surgical Research Projects The use of statistical analysis and data-driven approaches in healthcare is crucial for improving patient outcomes and optimizing resource allocation. During my co-op at UBC Province Wide Division of General Surgery at Vancouver General Hospital, I contributed to over 15 surgical research projects around the province, focusing on study design, data analysis, and statistical consulting. My work involved managing missing data, validating data integrity, collaborating on grant applications and abstracts, conduct statistical analysis, and supporting quality improvement initiatives through metrics monitoring and assessment. Through this co-op experience, I have gained valuable insights into the application of statistical methods in healthcare and the importance of data-driven decision-making in surgical research and quality improvement.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Marc Wettengel, UBC Statistics MSc student

Title: Analysis of the associations between environmental conditions and norovirus outbreaks in shellfish harvest zones on Vancouver Island

Abstract: Norovirus is a common cause of gastroenteritis with infection characterized by diarrhea, vomiting, and stomach pain. Bivalve molluscan shellfish (oysters, muscles, scallops, etc.) are common sources of community norovirus outbreaks in the Lower Mainland. This project uses longitudinal and time series methods to analyze the relationship between environmental conditions and norovirus outbreaks in shellfish harvesting zones on Vancouver Island. Weekly measurements on rainfall, ocean salinity, sea surface temperature and other environmental conditions were compiled from 2003 through 2019. Community norovirus case counts were also included as norovirus is not naturally present environment. This covariate is used as a proxy to determine if norovirus could potentially be present in the inter tidal zones which shellfish harvesting occurs. These covariates were compared to outbreak periods in shellfish harvesting zones during the same time period. Generalized linear mixed effect models, generalized estimating equations, and distributed lag non-linear models were fitted and compared with each other to determine the associations between environmental conditions and norovirus outbreaks in shellfish.