Seminar

2018-19 Constance van Eeden Lecture: The Blessings of Multiple Causes

*A small pre-talk reception will be served just outside the venue (ESB 2012) at 5pm.*

Abstract: Causal inference from observational data is a vital problem, but it comes with strong assumptions. Most methods require that we observe all confounders, variables that correlate to both the causal variables (the treatment) and the effect of those variables (how well the treatment works). But whether we have observed all confounders is a famously untestable assumption. We describe the deconfounder, a way to do causal inference from observational data with weaker assumptions that the classical methods require.

How does the deconfounder work? While traditional causal methods measure the effect of a single cause on an outcome, many modern scientific studies involve multiple causes, different variables whose effects are simultaneously of interest. The deconfounder uses the multiple causes as a signal for unobserved confounders, combining unsupervised machine learning and predictive model checking to perform causal inference.

We describe the theoretical requirements for the deconfounder to provide unbiased causal estimates, and show that it requires weaker assumptions than classical causal inference. We analyze the deconfounder's performance in three types of studies: semi-simulated data around smoking and lung cancer, semi-simulated data around genomewide association studies, and a real dataset about actors and movie revenue. The deconfounder provides a checkable approach to estimating close-to-truth causal effects.

This is joint work with Yixin Wang.
 

Speaker’s Biography: David Blei is a Professor of Statistics and Computer Science at Columbia University, and a member of the Columbia Data Science Institute. He studies probabilistic machine learning, including its theory, algorithms, and application. David has received several awards for his research, including a Sloan Fellowship (2010), Office of Naval Research Young Investigator Award (2011), Presidential Early Career Award for Scientists and Engineers (2011), Blavatnik Faculty Award (2013), ACM-Infosys Foundation Award (2013), and a Guggenheim fellowship (2017). He is the co-editor-in-chief of the Journal of Machine Learning Research. He is a fellow of the ACM and the IMS.

---------

This talk is supported by the van Eeden fund, the Department of Statistics, and PIMS.

Prediction based on conditional distributions of vine copulas

Vine copula models are a flexible tool in multivariate non-Gaussian distributions. For data from an observational study where the explanatory variables and response variables are measured over sampling units, we propose a vine copula regression method that uses regular vines and handles mixed continuous and discrete variables. This method can efficiently compute the conditional distribution of the response variable given the explanatory variables. Furthermore, we provide a theoretical analysis of the asymptotic conditional cumulative distribution function and quantile functions arising from vine copulas. Assuming all variables have been transformed to standard normal, the conditional quantile function could be asymptotically linear, sublinear, or constant with respect to the explanatory variables in different extreme directions, depending on the dependence properties of bivariate copulas in joint tails. The performance of the proposed method is evaluated on simulated data sets and a real data set. The experiments demonstrate that the vine copula regression method is superior to linear regression in making inferences with conditional heteroscedasticity. 

CANCELLED: Geo-dependent spatial models of infectious disease transmission

UNFORTUNATELY, TODAY'S SEMINAR IS CANCELLED:

Numerous examples exist of infectious disease models that incorporate spatial distance and other covariates at the individual level. This has been most noticeable perhaps in agricultural case studies such as the UK 2001 foot and mouth disease epidemic. However, both in agriculture and public health, many salient covariates that display spatial structure are collected at a regional level. Here, we extend individual level infectious disease models of the type proposed by Deardon et al. (2010) to incorporate such spatially structured regional/aggregate level information. This is done primarily within the context of influenza data from Calgary, Alberta. We discuss issues of both inference and computation.

Deardon, R., Brooks, S., Grenfell, B., Keeling, M., Tildesley, M., Savill, N., Shaw, D., Woolhouse, M. (2010). Inference for individual-level models of infectious diseases in large populations. Statistica Sinica  20:239-261.

Neyman-Pearson Classification Algorithms and NP Receiver Operating Characteristics

In many binary classification application in biomedical sciences, such as disease diagnosis, practitioners commonly face the need to control type I error (i.e., the conditional probability of misclassifying a class 0 observation as class1) so that it remains below a desired threshold. To address this need, the Neyman-Pearson (NP) classification paradigm is a natural choice; it minimizes type II error (i.e., the conditional probability of misclassifying a class 1 observation as class 0) while enforcing an upper bound, alpha, on the type I error. Although the NP paradigm has a century-long history in hypothesis testing, it has not been well recognized and implemented in classification schemes. Common practices that directly control the empirical type I error to no more than alpha do not satisfy the type I error control objective because the resulting classifiers are still likely to have type I errors much larger than alpha. As a result, the NP paradigm has not been properly implemented for many classification scenarios in practice. In this work, we develop the first umbrella algorithm that implements the NP paradigm for all scoring-type classification methods, such as logistic regression, support vector machines and random forests. Powered by this umbrella algorithm, we propose a novel graphical tool for NP classification methods: NP receiver operating characteristic (NP-ROC) bands, motivated by the popular receiver operating characteristic (ROC) curves. NP-ROC bands will help choose alpha in a data adaptive way and compare different NP classifiers. We demonstrate the use and properties of the NP umbrella algorithm and NP-ROC bands, available in the R package nproc, through simulation and real data case studies.

Clustering and modelling of phase variation for functional data

Our work is motivated by an analysis of elephant seal dive profiles which we view as functional data, specifically, as depth as a function of time, with data recorded almost continuously by sensors attached to the animal. The objective is to group profiles by shape to better understand the corresponding behavioural states of the seals. Most existing approaches rely on multivariate clustering methods applied to ad hoc summaries of the dive profile. Instead, we view each profile as arising from a function that is a deformation of a base shape. The deformation is regarded as phase variation and is represented by a latent warping function with a finite mixture distribution.

We first propose a curve registration model to explicitly model amplitude and phase variations of functional data, with phase variation represented by smooth time transformations called warping functions. Inference is conducted via the stochastic approximation expectation-maximization (SAEM) algorithm. Our simulation study shows that the SAEM algorithm is computationally more stable and efficient than existing approaches in the literature for inference of this class of curve registration model with flexible warping.

We then propose two clustering approaches based on our curve registration model for functional data: 1) a simultaneous approach that smooths the noisy raw profiles and estimates the base shape, the warping functions and the cluster membership and via SAEM algorithms; and 2) a two-step approach that applies clustering algorithms on the estimated warping functions. In contrast to generic clustering algorithms in the literature, our methods treat the clustering structure as heterogeneity in phase variation. The proposed method is applied to the analysis of elephant seal dive profiles and an analysis of human growth curves. We are able to obtain more intuitive clusters by focusing the clustering effort on phase variation.

The Use of Probabilistic Graphical Models in Evaluating the Current Opioid Overdose Crisis

North America is currently experiencing an opioid overdose crisis, with estimates of life expectancy dropping for two consecutive years owing to the large number of deaths due to opioid overdose. This has been particularly pronounced in British Columbia, where illicit-drug overdose deaths have increased by 750% in the last ten years, primarily being driven by the potent synthetic opioid fentanyl. The province has responded by declaring a public health emergency in April 2016, and by introducing or scaling up a number of interventions such as take-home naloxone kits, overdose prevention sites, and by increasing prescriptions of treatment for opioid use disorder. Current metrics for evaluation have focussed on numbers of intervention received or distributed. However, there is a need to understand how these interventions have had an impact on the number of overdose-related deaths.

In this talk, I will present the use of probabilistic graphical models (a Bayesian analysis method) in order to estimate the impact of intervention. Using the methodology, I combined together into one model disparate data routinely collected by the province including ambulance-attended overdoses, deaths due to overdose, and intervention data. I was also able to incorporate other factors which may exacerbate the epidemic, including drug supply contaminant information collected through urinalysis, and the impact of weather. I will provide a few examples of the approach before providing a detailed description of the current analysis for British Columbia and its implications.

As this new public health crisis emerges, it is imperative that we have tools and methods to understand the impact the current interventions have had, as well as planning for the future. Large-scale public health data are increasingly being collected and centralised. Use of these Bayesian analysis methods enable these data to be fully utilised to understand how policy and intervention impact population health.

Profiling Hospital Performance Based on Racial and Socioeconomic Disparities in Readmissions

Over the last decade, the U.S. Department of Health and Human Services has been promoting the use of quality measures that tie government insurance reimbursements to hospitals' relative performances in patients care. Although the Hospital Readmission Reduction Program, established under the Affordable Care Act, witnessed great success of quality measures in reducing post-discharge unplanned readmissions and the associated expenses, the disparities in healthcare delivered to patients from different socioeconomic or racial groups remain unattended. As a first step to incentivize hospitals to reduce healthcare inequity, we extend the methodology underlying current quality measures to quantify disparities in unplanned readmissions attributable to hospitals between black and white patients and between low and high socioeconomic status patients. We propose the absolute rate difference as a more intuitive effect measure than odds ratio and derive its interval estimate using a bootstrap procedure. Alternative analyses based on Bayesian hierarchical models and propensity scores are also considered. Using recent national database for patients with three common medical conditions, we discovered that racial and socioeconomic disparities persist even after accounting for differences in patients' characteristics and that the statistically significant variations in racial and socioeconomic disparities across hospitals indicate room for some hospitals to improve their practices to achieve the long-standing goal of healthcare equity.
 

From single-cells to patient diagnosis: statistical machine learning approaches to biomedical discovery

Recent technological advances have enabled routine collection of large quantities of biomedical information, ranging from molecular data such as gene expression in single-cells to population level datasets such as critical care databases. In this talk I will outline three current projects that apply and adapt cutting edge statistical approaches to effectively leverage this information to provide biomedical insights. Firstly, I'll introduce methodology to integrate single-cell RNA and DNA sequencing data to assign gene expression states to mutational cancer clones, discussing how we implemented our model using stochastic gradient Variational Bayes in Tensorflow. Secondly, I will introduce cellassign, a probabilistic model that automates the annotation of single-cell RNA-sequencing data to known types, allowing effective quantification of the tumour microenvironment in human cancers. Finally, I will introduce a probabilistic framework for optimizing the order of diagnostic test acquisition in a critical care context, and demonstrate how this can be applied to black box prediction functions (such as deep neural nets) in the context of differential diagnosis.

Statistical Inference for Cataloging the Visible Universe

A key task in astronomy is to locate astronomical objects in images and to characterize them according to physical parameters such as brightness, color, and morphology. This task, known as cataloging, is challenging for several reasons: many astronomical objects are much dimmer than the sky background, labeled data is generally unavailable, overlapping astronomical objects must be resolved collectively, and the datasets are enormous -- terabytes now, petabytes soon. Existing approaches to cataloging are largely based on algorithmic software pipelines that lack an explicit inferential basis. In this talk, I present a new approach to cataloging based on inference in a fully specified probabilistic model. I consider two inference procedures: one based on variational inference (VI) and another based on MCMC. A distributed implementation of VI, written in Julia and run on a supercomputer, achieves petascale performance -- a first for any high-productivity programming language. The run is the largest-scale application of Bayesian inference reported to date. In an extension, using new ideas from variational autoencoders and deep learning, I avoid many of the traditional disadvantages of VI relative to MCMC, and improve model fit.

Accurate inference of DNA methylation data: statistical challenges lead to biological insights

DNA methylation is an epigenetic modification widely believed to act as a repressive signal of gene expression. Whether or not this signal is causal, however, is currently under debate. Recently, a groundbreaking experiment probed the influence of genome-wide promoter DNA methylation on transcription and concluded that it is generally insufficient to induce repression. However, the previous study did not make full use of statistical inference in identifying differentially methylated promoters. In this talk, I’ll detail the pressing statistical challenges in the area of DNA methylation sequencing analysis, as well as introduce a statistical method that overcomes these challenges to perform accurate inference. Using both Monte Carlo simulation and complementary experimental data, I’ll demonstrate that the inferential approach has improved sensitivity to detect regions enriched for downstream changes in gene expression while accurately controlling the False Discovery Rate. I will also highlight the utility of the method through a reanalysis of the landmark study of the causal role of DNA methylation. In contrast to the previous study, our results show that DNA methylation of thousands of promoters overwhelmingly represses gene expression.