Seminar

Meta-Analytic Inference for the COVID-19 Infection Fatality Rate

To join via Zoom: To join this seminar, please request Zoom connection details from pims@uvic.ca  

Title: Meta-Analytic Inference for the COVID-19 Infection Fatality Rate

Abstract: Estimating the COVID-19 infection fatality rate (IFR) has proven to be challenging, since data on deaths and data on the number of infections are subject to various biases. I will describe some joint work with Harlan Campbell and others on both methodological and applied aspects of meeting this challenge, in a meta-analytic framework of combining data from different populations. I will start with the easier case when the infection data are obtained via random sampling. Then I will discuss drawing in additional infection data obtained in decidedly non-random manner.

 

Iterated Block Particle Filter for High-dimensional Parameter Learning: Beating the Curse of Dimensionality

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

Title: Iterated Block Particle Filter for High-dimensional Parameter Learning: Beating the Curse of Dimensionality

Abstract: Parameter learning for high-dimensional, partially observed, and nonlinear stochastic processes is a methodological challenge. Spatiotemporal disease transmission systems provide examples of such processes giving rise to open inference problems. We propose the iterated block particle filter (IBPF) algorithm for learning high-dimensional parameters over graphical state space models with general state spaces, measures, transition densities and graph structure. Theoretical performance guarantees are obtained on beating the curse of dimensionality (COD), algorithm convergence, and likelihood maximization. Experiments on a highly nonlinear and non-Gaussian spatiotemporal model for measles transmission reveal that the iterated ensemble Kalman filter algorithm (Li et al. (2020), Science) is ineffective and the iterated filtering algorithm (Ionides et al. (2015), PNAS) suffers from the COD, while our IBPF algorithm beats COD consistently across various experiments with different metrics.

CANCELLED: Instance-dependent Reinforcement Learning: A statistical viewpoint

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

Title: Instance-dependent Reinforcement Learning: A statistical viewpoint

Abstract: In recent years, there has been tremendous progress in the field of reinforcement learning (RL), especially on the empirical side. But it is fair to say that there is a considerable gap between theory and practice: many RL methods behave far better than existing worst-case theory would suggest, and often they work in settings where the current worst-case guarantees are completely prohibitive. In this talk, we will discuss why worst-case guarantees can severely overestimate the difficulty of reinforcement learning problems in presence of favorable structure. This motivates us to consider an instance-dependent difficulty measure that is responsive to the problem structure. Next, we discuss how we can construct estimators that adapt to this instance-dependent difficulty. We show that for problems with favorable structures our proposed estimators and associated confidence regions are significantly better than those obtained from the worst-case theory. Finally, we show that the techniques that we developed for constructing instance-dependent estimators are not specific to RL problems, and can be applied to a broad class of other problems.

***

Dear STAT news subscribers:

Please be advised that this seminar has been cancelled. We will reschedule the seminar in the near future. We sincerely apologize for any inconvenience.

Best wishes,

UBC Statistics Department

CANSSI Saskatchewan HSCC Webinar Series January-April 2022: Jiahua Chen

Registration & talk details

This talk has been organized by the Canadian Statistical Sciences Institute (CANSSI) Saskatchewan Health Science Collaborating Centre (HSCC). Learn more and register for this talk here.

Talk Title: Gaussian Mixture Reduction based on Composite Transportation Divergence

Abstract: In many applications, researchers wish to approximate a finite Gaussian mixture distribution with a high order by one with a lower order. Examples include density estimation, recursive tracking in hidden Markov model, and belief propagation. A direct solution to such a Gaussian Mixture Reduction problem is computationally challenging due to the non-convexity of commonly employed optimality targets.

One popular line of approach is to employ some clustering-based iterative algorithms. Neither their convergence nor destination, however, are thoroughly discussed. In this paper, we propose a new GMR method by minimizing some novel composite transportation divergence (CTD). This divergence permits an easy to implement Majorization-Minimization (MM) algorithm. We prove that the MM algorithms converge under general conditions, and many existing clustering-based algorithms are special cases of our approach. We further investigate the property of this approach with various choices of cost functions and demonstrate its effectiveness and computational costs.

CANCELLED: Non-reversible parallel tempering on optimized paths

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: MCMC methods are a popular tool in computation science used to evaluate expectations with respect to complex probability distributions over general state spaces. They work by averaging over the trajectory of a Markov chain stationary with respect to the target distribution. In theory, the MCMC algorithms converge asymptotically, but, in practice, for challenging problems where the target distributions are high-dimensional with well-separated modes, MCMC algorithms can get trapped exploring local regions of high probability and suffer from poor mixing.

Physicists and statisticians independently introduced parallel tempering (PT) algorithms to tackle this issue. PT delegates the task of exploration to additional annealed chains running in parallel with better mixing properties. They then communicate with the target chain of interest and help discover new unexplored regions of the sample space. Since their introduction in the 90s, PT algorithms are still extensively used to improve mixing in challenging sampling problems arising in statistics, physics, computational chemistry, phylogenetics, and machine learning.

The classical approach to designing PT algorithms was developed using a reversible paradigm that is difficult to tune and deteriorates in performance when too many parallel chains are introduced. This talk will introduce a new non-reversible paradigm for PT that dominates its reversible counterpart while avoiding the performance collapse endemic to reversible methods. We will then establish near-optimal tuning guidelines and efficient black-box methodology scalable to GPUs. Our work out-performs state-of-the-art PT methods and has been used at scale by researchers to study the evolutionary structure of cancer and discover magnetic polarization in the photograph of the supermassive black hole M87.

***

Dear STAT news subscribers:

Please be advised that this seminar has been cancelled due to the speaker's personal reason. We will reschedule the seminar some time soon. We sincerely apologize for the last minute notice.

Best wishes,

UBC Statistics Department

Statistical Learning and Matching Markets

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

Title: Statistical Learning and Matching Markets

Abstract: We study the problem of decision-making in the setting of a scarcity of shared resources when the preferences of agents are unknown a priori and must be learned from data. Taking the two-sided matching market as a running example, we focus on the decentralized setting, where agents do not share their learned preferences with a central authority. Our approach is based on the representation of preferences in a reproducing kernel Hilbert space, and a learning algorithm for preferences that accounts for uncertainty due to the competition among the agents in the market. Under regularity conditions, we show that our estimator of preferences converges at a minimax optimal rate. Given this result, we derive optimal strategies that maximize agents' expected payoffs and we calibrate the uncertain state by taking opportunity costs into account. We also derive an incentive-compatibility property and show that the outcome from the learned strategies has a stability property. Finally, we prove a fairness property that asserts that there exists no justified envy according to the learned strategies.

This is a joint work with Michael I. Jordan.

Observational Data with a Continuous Exposure: Study Design and Data Analysis

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

Title: Observational Data with a Continuous Exposure: Study Design and Data Analysis

Abstract: Statisticians and empirical researchers encounter observational data with a continuous exposure, either a treatment exposure or an IV-defined exposure, on a daily basis. Simple, ad-hoc methods that dichotomize the continuous exposure typically render the potential outcomes ill-defined, and obliterate the rich information contained in the original, continuous exposure. The current state-of-the-art study design approach to observational data with a continuous exposure is a technique called nonbipartite pair match; the method has enjoyed much success in empirical studies but suffers from two major limitations. From a theoretical perspective, a pair match is not optimal among the class of all subclassifications. From a very practical perspective, the pair match design often discards certain study units to design two groups well-balanced on observed covariates while separate in the exposure dose. In this talk, we propose a novel nonbipartite full match design that successfully solves both limitations. There are two types of data analysis facilitated by the proposed study design: instrumental variable (IV) analysis and dose-response relationship analysis. We illustrate the IV analysis using our recent empirical work investigating the association between intraoperative TEE use in CABG surgery and 30-day mortality rate, and the dose-response relationship analysis using our recent work on the effect of social mobility during the first phased reopening last year on subsequent Covid-19-related public health outcomes.

Parts of the talk are based on the following papers:

Application: https://www.onlinejase.com/article/S0894-7317(21)00029-8/fulltext

Statistical Methodology:  https://arxiv.org/pdf/2012.07182.pdf and  https://arxiv.org/pdf/2011.06917.pdf

Nonparametric Empirical Bayes Inference

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Title: Nonparametric Empirical Bayes Inference

Abstract: In an empirical Bayes analysis, we use data from repeated sampling to imitate inferences made by an oracle Bayesian with extensive knowledge of the data-generating distribution. Existing results provide a comprehensive characterization of when and why empirical Bayes point estimates accurately recover oracle Bayes behavior. In the first part of this talk, we construct flexible and practical nonparametric confidence intervals that provide asymptotic frequentist coverage of empirical Bayes estimands, such as the posterior mean and the local false sign rate. From a methodological perspective, we build upon results on affine minimax estimation, and our coverage statements hold even when estimands are only partially identified or when empirical Bayes point estimates converge very slowly. In the second part of the talk, we apply these ideas to study randomization-based inference for treatment effects in the regression discontinuity design under a model where the running variable has exogenous measurement error.

The folded concave Laplacian spectral penalty learns block diagonal sparsity patterns with the strong oracle property

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

Title: The folded concave Laplacian spectral penalty learns block diagonal sparsity patterns with the strong oracle property

Abstract: Structured sparsity is an important part of the modern statistical toolkit. We say a set of model parameters has block diagonal sparsity up to permutations if its elements can be viewed as the edges of a graph that has multiple connected components. For example, a block diagonal correlation matrix with K blocks of variables corresponds to a graph with K connected components whose nodes are the variables and whose edges are the correlations. This type of sparsity captures clusters of model parameters. To learn block diagonal sparsity patterns we develop folded concave Laplacian spectral penalty and provide a majorization-minimization algorithm for the resulting non-convex problem. We show this algorithm has the appealing computational and statistical guarantee of converging to the oracle estimator after two steps with high probability, even in high-dimensional settings. The theory is then demonstrated in several classical problems including covariance estimation, linear regression, and logistic regression.

Valid Inference After Hierarchical Clustering

To join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

Title: Valid Inference After Hierarchical Clustering 

Abstract: Testing for a difference in means between two groups is fundamental to answering research questions across virtually every scientific area. Classical tests control the type I error rate when the groups are defined a priori. However, if the groups are instead defined using a clustering algorithm, then applying a classical test yields an extremely inflated type I error rate. Surprisingly, this problem persists even if two separate and independent data sets are used for clustering and for hypothesis testing.

In this talk, I will propose a test for a difference in means between two estimated clusters that accounts for the fact that the null hypothesis is a function of the data, using a selective inference framework. Then, I will describe how to efficiently compute exact p-values for clusters obtained using hierarchical clustering. I will also show an application in the context of single-cell RNA-sequencing data, where it is common for researchers to cluster the cells, then test for a difference in mean gene expression between the clusters.

This talk is based on joint work with Jacob Bien (University of Southern California) and Daniela Witten (University of Washington).