Seminar

False Discovery Rate Estimation for High-dimensional Regression Models

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: A genome-wide association study (GWAS) aims to determine genetic variants statistically associated with phenotypes. However, because of linkage disequilibrium (LD), a characteristic of large-scale genomic datasets referring to the strong local dependencies between single-nucleotide polymorphisms (SNPs), it is usually challenging to identify the actual causal variants among its associated proxies. In this work, we propose a Bayesian variable selection method called the sparse mixed Gaussian prior for generalized linear models (SMG-GLM). It is an efficient high-dimensional Bayesian variable selection approach designed for arbitrary relationships between variants and phenotypes. Besides, it calibrates the selection uncertainty by estimating posterior inclusion probabilities, which many popular variable selection methods do not address. We additionally combine the SMG-GLM with knockoffs, named SMG-knockoffs, to account for the collinearity problem caused by LD. The SMG-knockoffs method can make inferences on the variable selection result and control the false discovery rate at an expected level. Its competence in discovering causal variables while controlling a desired false discovery rate has been shown in simulation studies conducted on a GWAS dataset.

Estimating Global and Country-Specific Excess Mortality During the COVID-19 Pandemic

To Join Via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Estimating the true mortality burden of COVID-19 for every country in the world is a difficult, but crucial, public health endeavor. Attributing deaths, direct or indirect, to COVID-19 is problematic.  A more attainable target is the ``excess deaths'', the number of deaths in a particular period, relative to that expected during ``normal times'', and we develop a model for this endeavor. The excess mortality requires two numbers, the total deaths and the expected deaths, but the former is unavailable for many countries, and so modeling is required for such countries. The expected deaths are based on historic data and we develop a model for producing estimates of these deaths for all countries. We allow for uncertainty in the modeled expected numbers when calculating the excess. The methods we describe  were used to produce the World Health Organization (WHO) excess death estimates. To achieve both interpretability and transparency we developed a relatively simple overdispersed Poisson count framework, within which the various data types can be modeled. We use data from countries with national monthly data to build a predictive log-linear regression model with time-varying coefficients for countries without data. For a number of countries, subnational data only are available, and we construct a multinomial model for such data, based on the assumption that the fractions of deaths in sub-regions remain approximately constant over time.  Our inferential approach is Bayesian, with the covariate predictive model  being implemented in the fast and accurate INLA software. The subnational modeling was carried out using MCMC in Stan or in some non-standard data situations, using our own MCMC code. Based on our modeling, the point estimate for global excess mortality, over 2020--2021, is 14.8 million, with a 95% credible interval of (13.2, 16.6) million.

This is joint work with William Msemburi, Victoria Knutson, Serge Aleshin-Guendel and Ariel Karlinsky.

Causal epistemology: estimands

To Join Via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: How do we know that the vaccine works and that cancer treatment prolongs life? We see treatments and we see changes in health. How do we know that one causes the other? It follows from reasoning. We reason that randomization makes treatment groups alike. We reason that stratification blocks covariation. We then credit the treatment with changes in health. In this talk, I will share a breakthrough made by discovering the hierarchy of seeing, doing, and imagining in causal attribution. The main learning objective is to understand the different types of estimands as we move through the hierarchy.

CANSSI Data Science ARES: Daniel J. McDonald

Registration & talk details

This talk is one of Data Science Applied Research and Education Seminar (ARES) series. Learn more and register for this talk here.

Talk Title: Markov-Switching State Space Models for Uncovering Musical Interpretation

Abstract: For concertgoers, musical interpretation is the most important factor in determining whether or not we enjoy a classical performance. Every performance includes mistakes—intonation issues, a lost note, an unpleasant sound—but these are all easily forgotten (or unnoticed) when a performer engages her audience, imbuing a piece with novel emotional content beyond the vague instructions inscribed on the printed page. In this research, we use data from the CHARM Mazurka Project—forty-six professional recordings of Chopin’s Mazurka Op. 68 No. 3 by consummate artists—with the goal of elucidating musically interpretable performance decisions. We focus specifically on each performer’s use of musical tempo by examining the inter-onset intervals of the note attacks in the recording. To explain these tempo decisions, we develop a switching state space model and estimate it by maximum likelihood combined with prior information gained from music theory and performance practice. We use the estimated parameters to quantitatively describe individual performance decisions and compare recordings. These comparisons suggest methods for informing music instruction, discovering listening preferences, and analyzing performances.

Robust Statistical Learning and Generative Adversarial Networks

To Join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Robust learning under Huber's contamination model has become an important topic in statistics and theoretical computer science. Statistically optimal procedures such as Tukey's median and other estimators based on depth functions are impractical because of their computational intractability. In this talk, we present an intriguing connection between f-GANs and various depth functions through the lens of f-Learning. Similar to the derivation of f-GANs, we show that these depth functions that lead to statistically optimal robust estimators can all be viewed as variational lower bounds of the total variation distance in the framework of f-Learning. This connection opens the door of computing robust estimators using tools developed for training GANs. In particular, we show in both theory and experiments that some appropriate structures of discriminator networks with hidden layers in GANs lead to statistically optimal robust location estimators for both Gaussian distribution and general elliptical distributions where first moment may not exist. This is a joint work with Chao Gao, Jiyi Liu, and Weizhi Zhu.

Short Bio: Yuan Yao is currently a Professor of Mathematics in Hong Kong University of Science and Technology (HKUST). Dr. Yao received his Ph.D. in Mathematics from UC Berkeley with Professor Steve Smale and worked in Stanford University and Peking University before joining HKUST in 2016. His main research interests lie in mathematics of data science and machine learning, with applications in computational biology and information technology.

CANCELLED: Computationally Efficient Bootstrap Sampling Using Markov Chains

To Join Via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: The bootstrap is a popular tool in modern statistics; however, it relies on sub-optimal IID Monte Carlo for approximating the resampling distribution. In this presentation, I will show that perhaps surprisingly, a simple MCMC algorithm can outperform IID sampling in this context. We aim to keep the same valuable features in the original method while hoping to improve sampling efficiency.

***

Dear STAT news subscribers:

Please be advised that this seminar has been cancelled. We will reschedule the seminar in the near future. We sincerely apologize for any inconvenience.

Best wishes,

UBC Statistics Department

Penalized Casebase in Survival Analysis

To Join Via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: A quantity of interest in clinical studies is the absolute risk of an event given a patient's covariate profile. In order for the Cox proportional hazard method to recover these absolute risk curves, the baseline hazard needs to be estimated separately, resulting in stepwise estimates in absolute risk that make the curves difficult to interpret. Casebase is an approach that transforms the data into a logistic regression, allowing for survival estimates that vary smoothly over time. We simulate data and compare the prediction performance of penalized casebase methods with penalized and unpenalized Cox methods. Casebase methods perform better in measures of concordance and time-dependent brier scores.

A Data-Driven Ensemble Framework for Modeling High-Dimensional Data: Theory, Methods, Algorithms and Applications

To Join Via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Sparse and ensemble methods are the two main approaches in the statistical literature for modeling high-dimensional data. On the one hand, sparse methods yield a single predictive model that is generally interpretable and possesses desirable theoretical properties. On the other hand, multi-model ensemble methods can generally achieve superior prediction accuracy, but current ensemble methodology relies on randomization or boosting to generate diverse models which results in uninterpretable ensembles. The diverse models generated by these “black box” algorithms are not insightful on their own and are only useful when they are pooled together.

In this dissertation, we introduce a new data-driven ensemble framework that combines ideas from sparse modeling and ensemble modeling. We search for optimal ways to select and split the candidate predictors into subsets for the different models that will be combined in an ensemble. Each model in the ensemble provides an alternative explanation for the relationship between the predictor variables and the response variable of interest. The degrees of sparsity of the individual models and diversity among the models are both driven by the data. The task of optimally splitting the candidate predictors into subsets results in a computationally intractable combinatorial optimization problem when the number of predictors is large. To demonstrate the potential of an exhaustive search for the optimal split of the predictors into the different models of an ensemble, we test our new approach on specifically designed low-dimensional data which mimic the typical behavior of high-dimensional data such as low signal-to-noise ratio and the presence of spurious correlations.

In this dissertation, we propose different computational approaches to the optimal split selection problem. We first introduce a multiconvex relaxation in the regression case and develop efficient algorithms to compute solutions for any level of sparsity and diversity. We show that the resulting ensembles yield consistent predictions and consistent individual models, and provide empirical evidence that this method outperforms state-of-the-art sparse and ensemble methods for high-dimensional prediction tasks using simulated data and a chemometrics application. We then extend the methodology, theory and algorithms to classification ensembles, and investigate the performance of the method on simulated data and a large collection of gene expression datasets. We finally propose a direct computational approach to calculate approximate solutions to the optimal split selection problem in the regression case and benchmark the performance of the method against the multi-convex relaxation on simulated and gene expression data.

Efficient software libraries with the implementations of the new computational methods provide researchers with tools to (1) achieve state-of-the-art accuracy for high-dimensional prediction tasks, and (2) aid in the scientific discovery of different mechanisms underlying the relationship between a large number of candidate predictors and the response variable of interest.

Boosting for regression problems with complex data

To Join Via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Boosting is a highly flexible and powerful approach when it comes to making predictions in non-parametric settings. In spite of the popularity and practical success of boosting algorithms, there is a lack of focus on its generalizations to “complex data”, such as data with outliers or functional variables. For data contaminated with outliers, we propose a two-stage boosting algorithm similar to what is done for robust linear MM-regression: it first minimizes a robust residual scale estimator, and then improves it by optimizing a bounded loss function. For data containing functional predictors, we propose a tree-based boosting algorithm that uses “base-learners” constructed with multiple projections. Our proposal incorporates possible interactions between indices, making it capable of approximating complex regression functions. Finally, we extend our proposals to contaminated functional data and explore two variations that can be used to perform robust functional regression. 

Vine copula mixture models and clustering for non-Gaussian data

To Join Via Zoom: Please register here.

Abstract: The majority of finite mixture models suffer from not allowing asymmetric tail dependencies within components and not capturing non-elliptical clusters in clustering applications. Since vine copulas are very flexible in capturing these dependencies, a novel vine copula mixture model for continuous data is proposed. The model selection and parameter estimation problems are discussed, and further, a new model-based clustering algorithm is formulated. The use of vine copulas in clustering allows for a range of shapes and dependency structures for the clusters. The simulation experiments illustrate a significant gain in clustering accuracy when notably asymmetric tail dependencies or/and non-Gaussian margins within the components exist. The analysis of real data sets accompanies the proposed method. The model-based clustering algorithm with vine copula mixture models outperforms others, especially for the non-Gaussian multivariate data.