Seminar

Statistical Methods for Population Size Estimation on Trees

To Join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Populations of interest are often hidden from data for a variety of reasons, though their magnitude remains important in determining resource allocation and appropriate policy. Increasing data collection and linkage across diverse fields suggests accessible methods of estimating population size with synthesized data are needed. In public health and epidemiology, these linkages often admit a tree structure, with the target population represented by the root, and paths from root-to-leaf representing pathways of care after a health event. We propose an extension to the well-known multiplier method which is applicable to tree-structured data, where multiple subpopulations and corresponding proportions combine to generate a population size estimate via the minimum variance estimator. The methodology is compared a Bayesian hierarchical model, for both simulated and real world data, the latter provided by BC's opioid overdose cohort. Finally, two R packages developed to facilitate the use of these methods on similar applications and lower the technical barrier of implementation will be discussed.

Conditional Inference and Prediction of a Survival Response Based on Vine Copula Models

To join this seminar virtually: please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Copulas are multivariate models that allow the modeling of dependence between variables without restrictions on marginal distributions. Copulas are useful for capturing complex dependence properties such as tail dependence, asymmetry, and nonlinearity. The vine copula can be viewed as an extension of a Gaussian copula after the correlation matrix is reparameterized to a set of algebraically independent correlations and partial correlations. It can be used to construct high-dimensional copula models with flexible dependence structures.  

We propose extensions to conditional inference and prediction methods based on vine copula models. For conditional inference, an algorithm is developed to compute arbitrary conditional distributions of one variable given the other variables for cross-prediction from a single joint distribution fitted by vine copula models. An existing algorithm is also modified to simulate data from a vine copula given that one variable takes extreme values. To predict a right-censored time-to-event response, a vine copula regression model is fitted to the response variable with the explanatory variables. Existing vine copula regression algorithms are modified to provide point and interval predictions for the censored response.

On factor copula-based mixed regression models

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: A copula-based method for mixed regression models is introduced, where the conditional distribution of the response variable, given covariates, is modelled by a parametric family of continuous or discrete distributions, and the effect of a common latent variable pertaining to a cluster is modelled with a factor copula. We show how to estimate the parameters of the copula and the parameters of the margins, and we find the asymptotic behaviour of the estimation errors. Numerical experiments are performed to assess the precision of the estimators for finite samples. An example of an application is given using COVID-19 vaccination hesitancy from several countries. Computations are based on R package CopulaGAMM.

Two UBC Statistics MSc student presentations (Giuseppe Tomio & Jana Osea)

To join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Presentation 1

Time: 11:00am – 11:30am

Speaker: Giuseppe Tomio, UBC Statistics MSc student

Title: A new data driven framework for simulating mendelian randomization data

Abstract: Mendelian randomization (MR) is a causal inference method that allows biostatisticians to leverage DNA measurements to study causal effects with only observed data. Recent advancements like two-sample summary-level MR (TS SL MR) and projects such as IEU GWAS database have lowered the barrier for conducting MR studies and opened the opportunity to mine causal effects in large-scale data sources. In the first part of my presentation, I show that there is a mismatch between how modern TS SL MR data is and how articles that propose popular TS SL MR models conduct their simulations. Next, I propose my solution: a data driven simulation framework for MR data that aims to be realistic, interpretable and easy to use thanks to a complementary R package implementation. As for the results, I show that models perform far better in literature-based simulations compared to more realistic simulations based on my proposed framework. Lastly, I warn that the mismatch between simulated and real data along with the obtained results may lead researchers to have over optimistic expectations about models performance in real applications.

Presentation 2

Time: 11:30am – 12:00pm

Speaker: Jana Osea, UBC Statistics MSc student

Title: Enhancing the Robustness of Instrumental Variable Estimation with Potentially Invalid Instruments and its Application to Mendelian Randomization

Abstract: Causal relationships between exposures and outcomes are vital in fields like epidemiology and medicine as they provide valuable insight into disease mechanisms, informing effective interventions. However, a common problem when attempting to extract the causal relationship between an exposure and an outcome in observational studies is the presence of unmeasured confounding. This leads to the exposure being correlated with the error term known as endogeneity. In order to obtain consistent estimates of the causal effect in the presence of endogeneity, we may use instrumetal variables (IVs) which are correlated with outcome only through its effect on the exposure. This captures the relationship between the exposure and outcome that is unaffected by the endogeneity. The most common use of IVs in the field of epidemiology and medicine is Mendelian Randomization (MR) which uses genetic variants as IVs. However, the validity of the IVs is often questionable due to the presence of pleiotropy and linkage disequilibrium, where the genetic variants affect the outcome through pathways other than the exposure. Furthermore, exposure and outcome data are often contaminated by outliers. In this thesis, we propose a novel algorithm, the robustified some valid some invalid instrumental variable estimator (rsisVIVE), that obtains estimates of the causal effect of an exposure on an outcome in the presence of invalid IVs and high levels of endogeneity while tolerating large proportions of contamination. The algorithm is based on the sisVIVE algorithm of Kang et. al. (2016) but we propose using robust estimators in both stages of the sisVIVE. Simulation results show that the rsisVIVE more accurately estimates the causal parameter than the sisVIVE when IVs are weak and outperforms competitor IV estimators in all cases when there is contamination.

Post hoc inference for genomics and neuroimaging

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Analyzing and interpreting high-dimensional data in applications such as genomics and neuroimaging often requires the simultaneous testing of a large number of hypotheses. In such multiple testing situations, the state-of-the-art approach selects one subset of hypotheses whose False Discovery Rate (FDR) in controlled. The FDR is the expected proportion of false positives (FDP) among selected hypotheses.

In contrast, post hoc inference aims to build confidence bounds on the FDP contained in arbitrary subsets of hypotheses, leading to improvements in interpretability and reproducibility of the results. We show how to construct and efficiently implement post hoc bounds that are adaptive to unknown dependence. We illustrate their application to differential gene expression studies in genomics, and fMRI studies in neuroimaging.

Epidemic models: can we make them behave better?

To Join this seminar: Please request Zoom connection details from headsec@stat.ubc.ca

Abstract: The COVID-19 pandemic has illustrated both the utility and limitation of using epidemic models for understanding and forecasting disease spread. One of the many difficulties in modelling epidemic spread is that caused by behavioural change in the underlying population. This can be a major issue in public health since, as we have seen during the COVID-19 pandemic, behaviour in the population can change drastically as infection levels vary, both due to government mandates and personal decisions. Such changes in the underlying population result in major changes in transmission dynamics of the disease, making the modelling challenges. However, these issues arise in agriculture and public health, as changes in farming practice are also often observed as disease prevalence changes. We propose a model formulation where time-varying transmission is captured by the level of alarm in the population and specified as a function of the past epidemic trajectory. The model is set in a data-augmented Bayesian framework as epidemic data are often only partially observed, and we can utilize prior information to help with parameter identifiability. We investigate the identifiability of the population alarm across a wide range of scenarios, using both parametric functions and non-parametric Gaussian process and splines. The benefit and utility of the proposed approach is illustrated through an application to COVID-19 data from New York City.

A corrected Clarke test for model selection

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: We introduce a large family of model selection tests based on the expectation of an arbitrary, possibly non-smooth, parametric criterion function of the data. It covers the case of strictly locally non-nested models and some overlapping models. The asymptotic theory of the proposed test statistic will be presented. A general exchangeable bootstrap scheme allows the evaluation of its limiting law as well as its asymptotic variance. In a simulation study, we empirically verify the distributional approximation of our test statistic in a finite sample and examine the empirical level and power of the corresponding model selection tests in various settings. Finally, an analysis of a financial dataset illustrates the proposed model selection procedure at work. The talk is based on a joint work with Florian Brueck and Jean-David Fermanian.

Markov Chain Monte Carlo and Langevin equations on a Stratification

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Title: Markov Chain Monte Carlo and Langevin equations on a Stratification

Abstract: Many sampling problems involve constraints — statistical models may involve relationships between parameters; noisy physical systems may involve stiff forces that constrain the system near a manifold, such as stiff bonds between particles. In some cases the constraints are not fixed, but rather can be added or removed (such as when bonds between particles form or break), so the probability measure of interest lives on sets of different dimensions. How can we sample from such a measure? I will introduce an MCMC algorithm to sample a probability measure supported on a stratification: a union of manifolds of different dimensions, glued together at their boundaries in a nice enough way. I will show this can accelerate simulations of interacting particles by up to several orders of magnitude. Then I will talk about our progress toward simulating Langevin equations on a stratification, which harnesses the theory of sticky diffusions. These algorithms are motivated by applications to systems of interacting particles, and I will be interested to learn about other areas of application. 

Co-op Report: Cardiovascular Network of Canada (CANet)

To Join this seminar: Please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: This report outlines my experience as a student biostatistician during an 8-month co-op; I highly recommend the experience to other students as a means to gain valuable work experience and to bridge the gap between the theoretical and the practical application of concepts taught in the MSc Statistics (Biostatistics) program.

CANet is an NCE-funded research network based in London, Ontario, focused on developing virtual care platforms for cardiovascular and other related health conditions. I was part of a clinical team performing statistical analyses and developing statistical protocols for projects including virtual care of atrial fibrillation and Post-MI management. I also lead the analysis of a clinical trial assessing a virtual care model for COVID-19 that became the basis of a research project undertaken with the supervision of Daniel McDonald).

Bayesian Models for Hierarchical Clustering of Network Data

To Join this seminar virtually: Please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Network data exist in many forms, like social networks, or interactions between cell proteins. Generally, they represent relational information between interacting entities. In many real-world examples, these entities tend to exhibit grouping structure. For example, the highly connected communities of people within a social network. Uncovering the underlying structure in networks is an important task for studying their composition and behaviour. Hierarchical clustering is a technique for discovering this structure across multiple scales, where a dendrogram represents the full hierarchy of clusters. This talk will explore Bayesian models for hierarchical clustering of network data, which aim to infer the posterior distribution over dendrograms.

The “Hierarchical Random Graph” is likely the most popular Bayesian approach to hierarchical clustering of network data. Yet, due to simplifications made in its inference scheme, we identify some potentially undesirable model behaviour. To rectify these issues, we introduce a general class of models that are characterized by a sampling construction, defining a generative process for simple graphs. We propose four Bayesian models from this class, and derive the marginalized posterior distribution over dendrograms, to isolate the problem of inferring a hierarchical clustering. We implement these models in a probabilistic programming language (Blang) that leverages state-of-the-art approximate inference methods (non-reversible Parallel Tempering). Finally, the empirical performance of our models is demonstrated on examples of real network data.