Seminar

Statistical and computational phenomena in deep learning

To join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Title: Statistical and computational phenomena in deep learning

Abstract: Deep learning's success has revealed a number of phenomena that appear to conflict with classical intuitions in the fields of optimization and statistics.  First, the objective functions formulated in deep learning are highly nonconvex but are typically amenable to minimization with first-order optimization methods like gradient descent.  And second, neural networks trained by gradient descent are capable of 'benign overfitting': they can achieve zero training error on noisy training data and simultaneously generalize well to unseen data.  In this talk we go over our recent work towards understanding these phenomena.  We show how the framework of proxy convexity allows for tractable optimization analysis despite nonconvexity, while the implicit regularization of gradient descent plays a key role in benign overfitting.   In closing, we discuss some of the questions that motivate our current work on understanding deep learning, and how we may use our insights to make deep learning more trustworthy, efficient, and powerful.

Adversarial Bayesian Simulation

To join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca

Title: Adversarial Bayesian Simulation

Abstract: In the absence of explicit or tractable likelihoods, Bayesians often resort to approximate Bayesian computation (ABC) for inference. Our work bridges ABC with deep neural implicit samplers based on generative adversarial networks (GANs) and adversarial variational Bayes. Both ABC and GANs compare aspects of observed and fake data to simulate from posteriors and likelihoods, respectively. We develop a Bayesian GAN (B-GAN) sampler that directly targets the posterior by solving an adversarial optimization problem. B-GAN is driven by a deterministic mapping learned on the ABC reference by conditional GANs. Once the mapping has been trained, iid posterior samples are obtained by filtering noise at a negligible additional cost. We propose two post-processing local refinements using (1) data-driven proposals with importance reweighting, and (2) variational Bayes. We support our findings with frequentist-Bayesian results, showing that the typical total variation distance between the true and approximate posteriors converges to zero for certain neural network generators and discriminators. Our findings on simulated data show highly competitive performance relative to some of the most recent likelihood-free posterior simulators.

van Eeden seminar: The four pillars of machine learning

Registration

To join this seminar, please register via Zoom. Once your registration is approved, you'll receive an email with details on how to join the meeting.

If you have any questions about your registration or the seminar, please contact headsec@stat.ubc.ca.

Title

The four pillars of machine learning

Abstract

I will present a unified perspective on the field of machine learning research, following the structure of my recent book, "Probabilistic Machine Learning: Advanced Topics" (https://probml.github.io/book2). In particular, I will discuss various models and algorithms for tackling the following four key tasks, which I call the "pillars of ML": prediction, control, discovery and generation. For each of these tasks, I will also briefly summarize a few of my own contributions, including methods for robust prediction under distribution shift, statistically efficient online decision making, discovering hidden regimes in high-dimensional time series data, and for generating high-resolution images.

van Eeden speakers

Dr. Kevin Patrick Murphy has been invited by our department's graduate students to be this year's van Eeden speaker. A van Eeden speaker is a prominent statistician who is chosen by our graduate students each year to give a lecture, supported by the Constance van Eeden Fund.

Stabilized COre gene and Pathway Election uncovers pan-cancer shared pathways and a cancer specific driver

To Join Via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Approaches systematically characterizing interactions via transcriptomic data usually follow two systems: (1) co-expression network analyses focusing on correlations between genes; (2) linear regressions (usually regularized) to select multiple genes jointly. Both suffer from the problem of stability: a slight change of parameterization or dataset could lead to dramatic alternations of outcomes. Here, we propose Stabilized Core gene and Pathway Election, or SCOPE, a tool integrating bootstrapped LASSO and co-expression analysis, leading to robust outcomes insensitive to variations in data. By applying SCOPE to six cancer expression datasets (BRCA, COAD, KIRC, LUAD, PRAD and THCA) in The Cancer Genome Atlas, we identified core genes capturing interaction effects in crucial pan-cancer pathways related to genome instability and DNA damage response. Moreover, we highlighted the pivotal role of CD63 as an oncogenic driver and a potential therapeutic target in kidney cancer. SCOPE enables stabilized investigations towards complex interactions using transcriptome data.

Understanding tumor heterogeneity through single-cell data

To Join Via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Cancer arises and evolves through the accumulation of somatic mutations which may provide a selective advantage. The interplay of mutations and their functional consequences shapes tumor progression and contributes to different clinical outcomes. Single-cell sequencing data enables a high-resolution characterization of this process, but requires powerful statistical models able to distinguish signal from noise. In this talk I discuss computational methods to analyze single-cell sequencing data from tumors to reconstruct the evolution of cancer cells, map genomic to transcriptional changes, and characterize the complex cell type composition of the tumor microenvironment. We present novel statistical models to integrate single-cell transcriptomes with copy number evolutionary trees, and to find hierarchical gene signatures from single-cell RNA-sequencing data. These methods provide rich descriptions of intra-tumor heterogeneity which are fundamental for the understanding of its complex dynamics and the development of targeted therapies.

False Discovery Rate Estimation for High-dimensional Regression Models

To Join via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: A genome-wide association study (GWAS) aims to determine genetic variants statistically associated with phenotypes. However, because of linkage disequilibrium (LD), a characteristic of large-scale genomic datasets referring to the strong local dependencies between single-nucleotide polymorphisms (SNPs), it is usually challenging to identify the actual causal variants among its associated proxies. In this work, we propose a Bayesian variable selection method called the sparse mixed Gaussian prior for generalized linear models (SMG-GLM). It is an efficient high-dimensional Bayesian variable selection approach designed for arbitrary relationships between variants and phenotypes. Besides, it calibrates the selection uncertainty by estimating posterior inclusion probabilities, which many popular variable selection methods do not address. We additionally combine the SMG-GLM with knockoffs, named SMG-knockoffs, to account for the collinearity problem caused by LD. The SMG-knockoffs method can make inferences on the variable selection result and control the false discovery rate at an expected level. Its competence in discovering causal variables while controlling a desired false discovery rate has been shown in simulation studies conducted on a GWAS dataset.

Estimating Global and Country-Specific Excess Mortality During the COVID-19 Pandemic

To Join Via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: Estimating the true mortality burden of COVID-19 for every country in the world is a difficult, but crucial, public health endeavor. Attributing deaths, direct or indirect, to COVID-19 is problematic.  A more attainable target is the ``excess deaths'', the number of deaths in a particular period, relative to that expected during ``normal times'', and we develop a model for this endeavor. The excess mortality requires two numbers, the total deaths and the expected deaths, but the former is unavailable for many countries, and so modeling is required for such countries. The expected deaths are based on historic data and we develop a model for producing estimates of these deaths for all countries. We allow for uncertainty in the modeled expected numbers when calculating the excess. The methods we describe  were used to produce the World Health Organization (WHO) excess death estimates. To achieve both interpretability and transparency we developed a relatively simple overdispersed Poisson count framework, within which the various data types can be modeled. We use data from countries with national monthly data to build a predictive log-linear regression model with time-varying coefficients for countries without data. For a number of countries, subnational data only are available, and we construct a multinomial model for such data, based on the assumption that the fractions of deaths in sub-regions remain approximately constant over time.  Our inferential approach is Bayesian, with the covariate predictive model  being implemented in the fast and accurate INLA software. The subnational modeling was carried out using MCMC in Stan or in some non-standard data situations, using our own MCMC code. Based on our modeling, the point estimate for global excess mortality, over 2020--2021, is 14.8 million, with a 95% credible interval of (13.2, 16.6) million.

This is joint work with William Msemburi, Victoria Knutson, Serge Aleshin-Guendel and Ariel Karlinsky.

Causal epistemology: estimands

To Join Via Zoom: To join this seminar virtually, please request Zoom connection details from headsec@stat.ubc.ca.

Abstract: How do we know that the vaccine works and that cancer treatment prolongs life? We see treatments and we see changes in health. How do we know that one causes the other? It follows from reasoning. We reason that randomization makes treatment groups alike. We reason that stratification blocks covariation. We then credit the treatment with changes in health. In this talk, I will share a breakthrough made by discovering the hierarchy of seeing, doing, and imagining in causal attribution. The main learning objective is to understand the different types of estimands as we move through the hierarchy.

CANSSI Data Science ARES: Daniel J. McDonald

Registration & talk details

This talk is one of Data Science Applied Research and Education Seminar (ARES) series. Learn more and register for this talk here.

Talk Title: Markov-Switching State Space Models for Uncovering Musical Interpretation

Abstract: For concertgoers, musical interpretation is the most important factor in determining whether or not we enjoy a classical performance. Every performance includes mistakes—intonation issues, a lost note, an unpleasant sound—but these are all easily forgotten (or unnoticed) when a performer engages her audience, imbuing a piece with novel emotional content beyond the vague instructions inscribed on the printed page. In this research, we use data from the CHARM Mazurka Project—forty-six professional recordings of Chopin’s Mazurka Op. 68 No. 3 by consummate artists—with the goal of elucidating musically interpretable performance decisions. We focus specifically on each performer’s use of musical tempo by examining the inter-onset intervals of the note attacks in the recording. To explain these tempo decisions, we develop a switching state space model and estimate it by maximum likelihood combined with prior information gained from music theory and performance practice. We use the estimated parameters to quantitatively describe individual performance decisions and compare recordings. These comparisons suggest methods for informing music instruction, discovering listening preferences, and analyzing performances.

Robust Statistical Learning and Generative Adversarial Networks

To Join via Zoom: To join this seminar, please request Zoom connection details from headsec@stat.ubc.ca

Abstract: Robust learning under Huber's contamination model has become an important topic in statistics and theoretical computer science. Statistically optimal procedures such as Tukey's median and other estimators based on depth functions are impractical because of their computational intractability. In this talk, we present an intriguing connection between f-GANs and various depth functions through the lens of f-Learning. Similar to the derivation of f-GANs, we show that these depth functions that lead to statistically optimal robust estimators can all be viewed as variational lower bounds of the total variation distance in the framework of f-Learning. This connection opens the door of computing robust estimators using tools developed for training GANs. In particular, we show in both theory and experiments that some appropriate structures of discriminator networks with hidden layers in GANs lead to statistically optimal robust location estimators for both Gaussian distribution and general elliptical distributions where first moment may not exist. This is a joint work with Chao Gao, Jiyi Liu, and Weizhi Zhu.

Short Bio: Yuan Yao is currently a Professor of Mathematics in Hong Kong University of Science and Technology (HKUST). Dr. Yao received his Ph.D. in Mathematics from UC Berkeley with Professor Steve Smale and worked in Stanford University and Peking University before joining HKUST in 2016. His main research interests lie in mathematics of data science and machine learning, with applications in computational biology and information technology.